What the work actually involves
You write prompts designed to expose what a model doesn't know about film and television — not trivia, but questions where the plausible answer is wrong. Release-date confusions between festival premiere and theatrical run, credit attributions that shifted in arbitration, awards categories that were renamed or retired, streaming rights that moved between platforms, franchise continuity that reboots quietly contradict. You then read multiple model responses to the same prompt, rank them, and write the justification: what specifically is wrong, and which authoritative source establishes it. AFI, BFI, Academy records, trade reporting from Variety or THR, studio press materials — the reference matters as much as the verdict.
A meaningful slice of the job is adversarial dataset construction. You're not just labelling; you're building benchmark items that will still discriminate between models six months from now. That means flagging your own prompts when they're ambiguous, rewriting questions that have two defensible answers, and documenting the reasoning so a researcher who has never seen a Cahiers du Cinéma poll can audit your call.
What the screen looks for
- Verifiable depth, not enthusiasm. Expect follow-ups that push past the first answer — if you cite a credit, you'll be asked how you'd confirm it and where the commonly repeated version goes wrong.
- Calibration. Knowing which entertainment claims are settled record and which are contested fan consensus is more valuable than knowing more facts.
- Written English under scrutiny. Your feedback is the product. Vague critique reads as unusable to a labelling pipeline.
- Honest availability. Forty hours weekly with four hours of PST overlap is a real constraint, not a negotiating position.
Logistics
Remote, US-only, with five states excluded. Contractor assignment for eight weeks, no benefits or paid leave. Pay is not disclosed in the listing; Turing's domain-expert bands vary by specialisation and are typically discussed after the assessment. The application begins with a take-home assessment — your submission is the main filter, and a Master's degree in any field is stated as a hard minimum.