What the work involves
You design short conversations — usually one to five turns — that force the model to reach into your actual personal context: past Gemini chats, Gmail, Google Search history, YouTube activity. Then you judge what comes back. Was the claim about you supported by real evidence, or an inference the model stitched together? Was the personal detail woven in naturally, or announced robotically in a way that reads as overnarrating? You rank two candidate responses side-by-side and write a rationale that cites the specific turn where the problem appeared. You also pull "Debug Info" to confirm which data sources the model actually drew on, and delete evaluation conversations afterward so they don't contaminate your own chat history.
The role has an unusual requirement worth being clear-eyed about: it uses your primary personal Google account, not a sandbox. The evaluation only works if the underlying data is real. If that arrangement doesn't sit right with you, this isn't the right quest.
What the screen is looking for
Turing's screen tests Vietnamese reading and writing at a high level of comprehension — rationales are expected to be clear, structured, and specific. Beyond language, it probes evaluation judgment: can you separate a response that is wrong about you from one that is merely awkward about you, and can you defend a ranking when both responses are mediocre in different ways? Prior data annotation, AI quality evaluation, or content moderation experience is strongly preferred. A BS/BA in an analytical field — linguistics, law, policy, journalism, CS — or equivalent experience is the stated baseline.
- Shortlisted candidates receive a Job Interest Form
- After profile review, a timed assessment must be completed within 24 hours
- Strong assessment outcomes lead to a pre-onboarding discussion
Logistics
Fully remote, contractor engagement, three months. The commitment is at least 4 hours per day and up to 40 hours per week, with 4 hours of overlap with PST — Turing is staffing a 24-hour global operations team, so your local-timezone availability matters as much as your total hours. You need a desktop or laptop and a reliable connection. The offered rate is $15/hour as posted; rates on Turing projects vary by project and are not guaranteed beyond what the listing states.