What the work actually is
You play two roles in the same session. First you are the author: you design conversational prompts in Chinese, typically one to five turns, that force the model to reach into your real personal context — past Gemini conversations, Gmail, Google Search history, YouTube activity. Then you switch to critic: you read two candidate responses side by side and decide which is more helpful, more natural, and better grounded, and you write a rationale that a reviewer who was not there can follow.
The failure modes you are hunting are specific. Grounding failures are claims about you that the data does not support — a plausible-sounding inference the model never earned. Forced connections are when the model drags in personalization that nobody asked for. Overnarrating is the model announcing what it knows about you instead of simply using it. You will also pull "Debug Info" from the model to confirm which chat summaries and data sources it actually drew on, rather than guessing. Data hygiene is part of the job: evaluation conversations get deleted so they do not pollute your own future chat history.
What the screen is looking for
- Genuine written Chinese fluency — rationales are written work product, not checkbox clicks
- Evidence you can distinguish a correct inference from a lucky-sounding one
- Rationales that cite turn numbers and quote the offending span, not summary verdicts like "Response A felt better"
- Prior exposure to annotation, AI evaluation, content moderation or similar rubric-driven work is strongly preferred, not strictly required
- A BS/BA or equivalent in an analytical field: policy, law, ethics, linguistics, journalism, computer science
Logistics
Contractor engagement, roughly three months, remote. Minimum four hours a day, up to forty a week, with four hours overlapping PST — that is a real constraint if you are in an Asian time zone, and worth thinking through before applying. The observed rate on this posting is $15/hour; Turing states rates per project and they vary. The process runs: Job Interest Form, profile review, then a timed assessment that must be returned within 24 hours, then pre-onboarding discussion.