What the work actually involves

You are producing the training material an LLM learns from, not using an LLM to do analysis. A typical shift means reading a block of content or a dataset, breaking it into logical components, and constructing questions a model is likely to get wrong — sales growth across locations where the choice of time window changes the answer, constraint-satisfaction puzzles with ordering rules, claims that sound plausible but fall apart under a web search. For each one you supply the correct answer plus a written explanation of the reasoning path, and where the project requires it, the equivalent in Malay. Much of the day is also spent grading model output: marking where the chain of reasoning broke, annotating why, and writing feedback specific enough that the failure is reproducible.

What the screen looks for

Turing's screening is less interested in your job titles than in whether you can show your reasoning in writing. Expect to be asked to walk through a quantitative judgement out loud — why one baseline period rather than another, how you handled an outlier, what you checked before trusting a figure. Malay fluency is tested as working fluency in both directions, including whether you can render analytical and technical phrasing naturally rather than translating word by word. No prior domain specialisation is required, but professional writing experience — analyst, journalist, editor, translator, technical writer — is what the platform reads as evidence you can produce clean annotations at volume.

Logistics

  • Contractor engagement: no medical or paid leave, contract extension tied to performance and project need.
  • 40 hours per week for the contract duration, with 2–5 hours per day overlapping UTC−8 (America/Los_Angeles).
  • Fully remote; you supply your own desktop or laptop and a reliable connection.
  • Pay is undisclosed on this listing and stated by Turing as based on experience and expertise — treat any figure you see quoted elsewhere as observed, not guaranteed.

The work suits someone who enjoys being adversarial with a problem: finding the ambiguity, the misleading average, the constraint everyone skips. It suits less well anyone who wants stable subject matter, since task types rotate as the labs' priorities shift.