What the work actually involves

You write short multi-turn conversations in Polish — usually one to five turns — designed to force the model to reach into your real personal context: past Gemini chats, Gmail, Google Search history, YouTube activity. Then you evaluate what comes back. The rubric has teeth: Grounding (is every claim the model makes about you actually supported, or is it a plausible-sounding inference it invented?), Integration (is the personal detail woven in naturally, or does the model overnarrate — "Since I can see from your Gmail that you recently booked a flight to Kraków..."), and Helpfulness overall. Much of the work is side-by-side: two responses, you stack-rank them and write a rationale that cites specific turn numbers and specific failures. You will also pull Debug Info to confirm which data sources the model actually used, and delete evaluation conversations afterward so they don't contaminate your own chat history.

The unusual condition here is that this runs on your primary personal Google account, not a sandbox. The evaluation is only meaningful if the personal data is real. Read that requirement carefully before applying; it is not negotiable and it is the most common reason candidates withdraw after being shortlisted.

What the screen is looking for

Turing's process is a Job Interest Form, then a profile review, then a timed assessment that must be returned within 24 hours. The assessment is the real filter. It is checking whether you can tell the difference between a model that genuinely used your data and one that produced a personalized-sounding response from nothing — and whether you can write that distinction down in Polish and English clearly enough that a reviewer who wasn't in the conversation can verify your judgment. Vague rationales ("Response A felt more natural") fail. Rationales that say "in turn 3, Response A claims I prefer morning workouts; nothing in the conversation or debug info supports that inference" pass.

Logistics

  • Fully remote, contractor engagement, roughly one month in length
  • Minimum 4 hours per day, up to 40 hours per week
  • Four hours of daily overlap with PST is required — the team runs 24-hour global operations, so your local schedule must stretch to meet that window
  • Desktop or laptop and a stable connection; no specialized hardware
  • Observed rate on this listing is $20/hr; Turing sets rates per project and they are not guaranteed to hold across renewals