What the work involves
You design short multi-turn conversations in Dutch — usually one to five turns — that force the model to reach for personal context, then judge what comes back. Each response gets assessed on grounding (are claims about you actually supported, or inferred from thin air), integration (is the personal detail woven in naturally, or does the model overnarrate and recite your data back at you), and overall helpfulness. Most tasks end in a side-by-side stack ranking of two candidate responses with a written rationale that points to the exact turn where something went right or wrong.
There is a genuine privacy trade here and the listing is explicit about it: you use your primary personal Google account, not a sandbox, with personal data sources enabled. Part of the routine is extracting "Debug Info" to confirm which sources and chat summaries the model actually pulled from, and deleting evaluation conversations afterwards so they don't contaminate your own future chat history.
What the platform screens for
Turing's process starts with a Job Interest Form, then a profile review, then a timed assessment that must be returned within 24 hours. Expect the assessment to test three things: whether your written Dutch is genuinely native-level and idiomatic, whether you can tell a forced connection from a good inference, and whether your rationales are specific enough to be defensible — vague "response A felt better" rankings fail. Prior annotation, AI evaluation or content moderation experience is preferred rather than required; analytical degree backgrounds (linguistics, law, policy, journalism, CS) are named explicitly.
Logistics
- Remote contractor engagement, roughly one month in length
- Minimum 4 hours per day, up to 40 hours per week
- Four hours of daily overlap with PST is required — the team runs 24-hour global operations
- Desktop or laptop with a stable connection; observed rate for this project is $20/hour