What the work actually involves
You write short multi-turn conversations — usually one to five turns — that only your own life can answer well: a trip you actually planned, a subscription you actually have, a thread sitting in your Gmail. Then you judge what comes back. The evaluation dimensions are specific: Grounding (is every claim the model makes about you supported by real evidence, or is it an inference it invented?), Integration (does the personal detail sit naturally in the sentence, or does the model overnarrate — "Based on your recent YouTube activity, I noticed that you…"), and Helpfulness overall. Most tasks end in a side-by-side stack-ranking of two responses plus a written rationale that cites the specific turn where the problem appeared.
There is also housekeeping that matters more than it sounds: pulling Debug Info to confirm which data sources and chat summaries the model actually drew on, and deleting evaluation conversations afterward so they don't contaminate your own future chat history.
What the platform screens for
- Korean at working reading and writing level. The focus language is Korean, and naturalness judgments — whether a personalized sentence sounds like a person or a system announcement — depend on native-level ear.
- Willingness to connect a real personal Google account. This is a genuine gate, not a preference. Testing accounts have no history to personalize from, so the project cannot use them.
- Judgment on ambiguous cases. Screeners push on whether you can distinguish a correct-but-clumsy personalization from a plausible-sounding hallucination about you, and whether you can defend a close SxS call in writing.
- Real availability. Full-time-capable in your local time zone, minimum four hours a day, with four hours overlapping PST — a demanding ask from Korea time zones, and they will check the arithmetic.
Logistics
Remote contractor engagement, roughly three months, four to forty hours a week, staffed as part of a 24-hour global operation. The observed rate on this listing is $15/hour; Turing's posted rates vary by project and are not guaranteed beyond what the listing states. You need your own desktop or laptop and a stable connection. The process runs: Job Interest Form → profile review → a timed assessment to be completed within 24 hours → pre-onboarding discussion.