What the work actually involves

You are handed model outputs — and in some project variants, prompts and source material to build from — and asked to decide whether the response holds up. That means reading long or tangled passages, decomposing them into claims, checking the claims against sources you find yourself, and marking where the reasoning breaks. The deliverable is rarely a score on its own: it is a score plus a written justification clear enough that a researcher tuning the model can see exactly which step failed. You will also be asked to author new material — scenarios, questions, worked examples with explanations — that becomes training data.

One honest caveat: the posting is titled Vision Annotators, but the role overview and responsibility list Turing published describe text-based analytical evaluation and LLM training data work. If image, video or document-image annotation matters to you, ask the recruiter directly which project queue you would be staffed on before signing. Turing runs many pools under overlapping titles and the actual assignment is decided at staffing, not at application.

What the screen looks for

  • Written English under load. Almost every task ends in prose. Screeners weigh whether your explanation is precise and ordered, not whether it is long.
  • Verification instinct. Can you tell a claim that needs a source from one that does not, and do you actually go look?
  • Calibration. Confident wrong answers are the expensive failure mode here. Saying "this part is unverifiable" is a valid, and often correct, output.
  • Independence. Queues are asynchronous with thin supervision; you are expected to resolve ambiguity against the guidelines rather than stall.

Logistics

Fully remote, contractor engagement with no leave or benefits. The commitment options are 20, 30 or 40 hours per week, with a floor of four hours per day and four hours of daily overlap with US Pacific time — the overlap is the real constraint for candidates outside the Americas. Initial contract is one month, extendable on performance and project demand. Pay is not disclosed in the posting; rates on Turing's annotation and evaluation pools vary by language, specialism and project, so treat any figure quoted to you as project-specific rather than a standing band. Required kit is modest: a reliable computer, stable internet, and working familiarity with Excel or Google Sheets.