What the work actually involves

You write starting prompts — typically one to five turns — that create a legitimate reason for the model to reach into your personal context: past Gemini conversations, Gmail, Google Search history, YouTube activity. Then you evaluate what comes back along defined dimensions. Grounding: is every claim the model makes about you supported by real evidence, or is it a plausible-sounding inference it has no basis for? Integration: is the personal detail woven in naturally, or does the model overnarrate — announcing that it checked your email before answering a simple question? Helpfulness: setting personalization aside, is this a better response than the alternative?

Much of the day is side-by-side comparison. Two responses, stack-ranked, with a written rationale that points to where exactly the problem occurred — turn three, this sentence, this forced connection. You will also pull Debug Info from the model to verify which data sources and chat summaries were actually used, which sometimes contradicts what the response implies. And you delete your evaluation conversations afterward, because they are using your real account and stale test threads would pollute both your own future personalization and the next day's work.

What the screen is looking for

Turing's process starts with a Job Interest Form, then a profile review, then a timed assessment you must return within 24 hours. The assessment is the real gate: it is looking for whether you can distinguish a flawed inference from a hallucination, whether your rationales are specific enough to be defensible to someone who did not see the conversation, and whether you can spot subtle naturalness differences rather than defaulting to whichever response is longer. A BS/BA in an analytical field — policy, law, ethics, linguistics, journalism, CS — or equivalent experience is the stated baseline; prior annotation, AI evaluation, or content moderation work is strongly preferred.

  • The personal-account requirement is a genuine screen-out. Testing accounts do not produce assessable personalization, so a synthetic or freshly created Google account will not work.
  • Work is remote and independent, but not fully async: at least 4 hours a day, up to 40 a week, with 4 hours overlapping PST.
  • Contractor engagement, stated at 3 months. Pay is not disclosed in the listing.

Practical logistics

You need a desktop or laptop and reliable internet — this is not phone-compatible work. The team is staffed across a 24-hour global operation, so your local time zone matters less than your ability to hold the PST overlap window consistently. Expect the work to be repetitive in structure and highly variable in content: the quality of your evaluations depends heavily on the quality of the prompts you invent, and inventing genuinely varied personal scenarios day after day is the part most people underestimate.