What the work involves

This is a personalization evaluation project, which means you are both the prompt author and the ground truth. You design short multi-turn conversations — typically one to five turns — that can only be answered well if the model draws on your actual Gmail, Google Search, YouTube activity and past Gemini chats. You then judge whether it did so correctly: was the claim about you supported by evidence, or was it a flawed inference? Was the personal detail woven in naturally, or did the model 'overnarrate' by announcing where it got the information? Most tasks end in a side-by-side ranking of two candidate responses plus a written rationale in Portuguese-language context that points to the specific turn where the problem occurred.

Two operational details matter more here than in most eval work. First, you must connect your primary personal Google account — not a clean test account — because the whole point is genuine personal context; if that is not something you are willing to do, this listing is not for you. Second, you are responsible for data hygiene: deleting evaluation conversations afterwards so they do not pollute your real chat history, and extracting 'Debug Info' to confirm which data sources the model actually touched.

What the screen looks for

  • Portuguese reading and writing at a level where you can describe subtle differences in tone and naturalness, not just correctness.
  • Evidence you can tell apart the failure modes this project names: wrong personalization, weak inference, forced connection, overnarration. Candidates who collapse all of these into 'the answer was bad' do not pass.
  • Rationales that are specific and defensible — turn numbers, quoted fragments, a stated reason one response wins rather than a preference.
  • Prior annotation, AI evaluation, or content moderation experience is strongly preferred rather than required; a BS/BA in an analytical field (policy, law, linguistics, journalism, CS) or equivalent experience is the stated baseline.

Logistics

Fully remote contractor engagement, roughly three months, with a minimum of four hours per day and up to 40 hours per week. The team runs 24-hour global operations, but this posting asks for four hours of daily overlap with PST — check that against your time zone before applying. You need your own desktop or laptop and a reliable connection. The observed rate on this listing is $15/hour. The process runs as a Job Interest Form, then a profile review, then a timed assessment that must be returned within 24 hours, then pre-onboarding discussion.