What the work actually involves

You are both the author of the test and its judge. Each task starts with you designing a Spanish-language prompt (typically one to five turns) that can only be answered well if the model draws on something real about you — a trip you actually planned, a subscription in your inbox, a channel you actually watch. You then run it, and evaluate the responses across defined dimensions: Grounding (is every claim about you supported by an actual data source, or is the model inferring and hallucinating?), Integration (is the personal detail woven in naturally, or does the model overnarrate — "Since I see from your Gmail that you..."), and overall Helpfulness.

Much of the volume is side-by-side comparison: two responses, stack-ranked, with a written rationale in clear prose that points to the exact turn where the problem or the win occurred. You'll also extract Debug Info to confirm which data sources and chat summaries the model actually pulled from, and delete evaluation conversations afterward so your test chatter doesn't pollute your own future history. That data-hygiene step is part of the job, not an afterthought.

What the screen looks for

  • Genuine Spanish reading and writing fluency — rationales are assessed as writing, not just as ratings.
  • Whether you can name the failure modes of personalization precisely: forced connections, stale context, over-inference from thin evidence, privacy-uncomfortable disclosures.
  • Willingness to connect a real personal Google account with real data sources. Candidates who want to use a clean test account are not a fit for this project.
  • Realistic availability. Turing is staffing a 24-hour global operation and wants four hours of daily overlap with PST.

Logistics

Remote, contractor engagement, roughly three months, at least 4 hours per day and up to 40 hours per week. The observed rate on this posting is $15/hour; Turing states rates per project and they are not guaranteed to hold across engagements. The process runs from a Job Interest Form to a timed assessment that must be returned within 24 hours, then pre-onboarding. You need a desktop or laptop and a stable connection.