What the work actually involves

You write problems the model is likely to get wrong, then write the correct answer and the reasoning that gets there. In practice that means three recurring shapes of task: reading a dense passage or dataset and decomposing it into logical blocks; taking a claim and verifying or falsifying it through online research with citable sources; and constructing constraint puzzles (ordering, scheduling, deduction) with a single defensible solution. A fair amount of output is in Russian, or involves checking that reasoning survives translation — idioms, number formatting, name declension and culturally specific references are all places models fail quietly.

The second half of the job is critique. You will read model outputs and annotate them: where the chain of reasoning broke, whether the final answer happens to be right for the wrong reason, whether a cited source actually supports the claim. Annotations are read by other reviewers and by lab-side researchers, so they need to be specific and self-contained rather than a verdict.

What the screen looks for

  • Genuine two-way Russian and English literacy — comprehension and production, not conversational fluency. Expect to be asked to reason in one language about material in the other.
  • Whether you can show your work. Turing cares less about your final answer than about whether the intermediate steps are legible and each is justified.
  • Basic quantitative comfort: percentage change versus absolute change, choosing a baseline period, noticing that early-month volatility distorts a growth comparison.
  • Research discipline: distinguishing a primary source from an aggregator that paraphrased it.
  • No specialised domain background is required, and the listing explicitly accepts candidates without a bachelor's degree who have relevant working experience.

Logistics

Fully remote, contractor terms with no medical or paid leave, maximum contract duration twelve months with extension possible on performance. The commitment is stated as 40 hours per week including 2–5 hours daily overlapping America/Los_Angeles (UTC-8). Compensation is not disclosed in this listing; Turing typically sets rates by experience and task tier, so treat any figure you see quoted elsewhere as observed rather than promised. You need your own desktop or laptop and a reliable connection.