What the work actually involves

You design short multi-turn conversations (usually one to five turns) that force the model to reach into your real personal context — a trip you booked, a recipe you searched, a channel you watch — and then you judge what comes back. Each task ends in a side-by-side comparison of two responses, a stack ranking, and a written rationale in Indonesian or English that points to the specific turn where the personalization helped or failed. Three failure modes drive most of the scoring: grounding (the model asserts something about you that no data source supports), integration (the personal detail is bolted on or overnarrated — "Since you searched for flights to Bali last Tuesday…"), and helpfulness (whether the personalization made the answer better at all). You will also pull Debug Info to confirm which sources and chat summaries the model actually drew on, and delete your evaluation conversations afterwards so they don't contaminate your own chat history.

What the platform screens for

Turing's screen is a Job Interest Form, then a timed assessment that must be returned within 24 hours, then a conversation about pre-onboarding. The assessment is the real gate: expect to write starting prompts, rank two responses, and defend the ranking in prose. Screeners look for rationales that cite turn numbers and quote the offending phrase rather than summarising a feeling, and for the ability to separate a wrong inference from a merely awkward one. Indonesian reading and writing at a high level is non-negotiable, since the focus language is Indonesian and judgments about naturalness cannot be made in translation.

Logistics and the personal-account requirement

  • Fully remote, contractor engagement, roughly three months, 30 or 40 hours per week with at least four hours per day and four hours overlapping PST.
  • You must use your primary personal Google account with personal data sources enabled — a clean test account defeats the purpose of the project. Decide whether you are comfortable with that before applying.
  • Desktop or laptop with a stable connection; the work is browser-based and involves reading long response pairs side by side.
  • Background is flexible: a BS/BA in linguistics, law, policy, journalism, computer science or another analytical field, with prior annotation, AI evaluation or content moderation experience strongly preferred.