What the work involves
You design short multi-turn conversations in Russian (typically 1–5 turns) that can only be answered well if the model draws on your actual personal context — past Gemini chats, Gmail, Search and YouTube activity. Then you grade what comes back. The scoring dimensions are specific: Grounding (is every claim about you actually supported, or is it a flawed inference or hallucination?), Integration (is the personal detail woven in naturally, or does the model overnarrate — "Since I see from your email that you..."), and overall Helpfulness. Much of the volume is side-by-side comparison: two responses, a stack-rank, and a written rationale that points to the exact turn where the problem occurred.
Two tasks sit around the evaluation itself. You extract and verify "Debug Info" to confirm which data sources and chat summaries the model actually pulled from — so your grounding calls rest on evidence rather than guesswork. And you delete evaluation conversations afterwards, so that synthetic prompts don't pollute the personal history the next task depends on. Data hygiene is part of the job, not an afterthought.
What the screen is looking for
Turing shortlists from a Job Interest Form, then sends a timed assessment that must be returned within 24 hours. The assessment is the real filter: expect to write prompts in Russian, rank paired responses, and justify the ranking in prose. Screeners weight three things — Russian writing that is genuinely native-level (rationales may be written in English, but prompt and response judgment is Russian-dependent), the ability to separate "wrong about me" from "clumsily phrased about me," and willingness to connect a real personal account. Candidates who hesitate on the account question are generally not placed, because the feature cannot be assessed on an empty test profile.
Logistics
- Remote contractor engagement, roughly 3 months, staffed as part of a 24-hour global operation.
- Minimum 4 hours per day, up to 40 hours per week, with 4 hours overlapping PST.
- Observed rate on this project is $15/hour; Turing states rates per project and they vary by track.
- Desktop or laptop and a stable connection required — this is not phone-workable.