What the work actually involves

You design short multi-turn conversations — usually one to five turns — that force the model to reach into your real personal context: past Gemini chats, Gmail, Google Search history, YouTube activity. Then you switch hats and grade what came back. Two model responses are shown side-by-side; you stack-rank them and write a rationale in clear prose that points to the exact turn where personalization landed, misfired, or was invented outright. You also pull "Debug Info" to confirm which data sources and chat summaries the model genuinely used, and you delete your evaluation conversations afterwards so they don't contaminate your own future chat history.

The judgment calls are subtle. A response can be accurate and still fail: forced connections that shoehorn your hobby into an unrelated answer, poor inferences drawn from one stray email, or "overnarrating" where the model announces that it remembers you like hiking instead of simply answering like it knows. Turkish is the focus language, so you need to feel when a personalized reply reads naturally in Turkish versus when it reads like a translated template.

What the screen looks for

  • Real Turkish reading and writing competence — rationales are expected to be structured and precise, not summary-level.
  • Willingness to use your primary personal Google account with personal data sources enabled. A clean test account defeats the purpose of the project and is a genuine gate here.
  • Vocabulary and instinct around evaluation: can you distinguish a grounding failure from an integration failure, and explain which one you saw.
  • Prior annotation, AI quality evaluation, or content moderation experience is strongly preferred, alongside a BS/BA in an analytical field — policy, law, ethics, linguistics, journalism, computer science, or similar.

Logistics

Fully remote, contractor engagement, roughly three months. Minimum four hours per day and 30 hours per week, with either a 30 or 40 hour option, and four hours of daily overlap with PST — the team runs 24-hour global operations, so your local schedule matters less than that overlap window. Desktop or laptop and a stable connection required. The process runs: Job Interest Form, profile review, then a timed assessment you must complete within 24 hours, then a pre-onboarding conversation.