What the work actually involves

You will spend most of your time writing board game scenarios that are hard for a model in a specific, diagnosable way — an edge case in a rules interaction, a decision point where the intuitive move is probabilistically wrong, a state that requires tracking hidden information across several turns. Then you grade what the model produces: not just whether the final answer is right, but whether the chain of reasoning that reached it was legitimate. A model that names the correct move for the wrong reason is a failure case worth documenting, and writing up why it is a failure is a large part of the job.

The rest of the time goes to infrastructure around the tasks: rubrics that another annotator could apply and reach the same score, evaluation guidelines, gold-standard reference solutions, and QA passes on other contributors' items. Projects in this category typically involve reviewing disputed labels and defending your own scoring when a second reviewer disagrees.

What the screen is looking for

  • Rules depth under follow-up. Expect probing on a specific game you claim to know — corner-case interactions, priority and timing, how ambiguous rules text is adjudicated in practice.
  • Genuine probabilistic reasoning. Not just knowing a move is good, but decomposing expected value, variance, and information state.
  • Ability to formalize. Turning a fuzzy "good play" intuition into a written criterion a stranger could apply consistently.
  • Writing. Rubrics and failure annotations are the deliverable; unclear prose makes the data unusable.

Prior annotation, RLHF, or benchmark experience is preferred but the listing does not gate on it. Python/SQL is listed as nice-to-have, not required.

Logistics

Fully remote contractor assignment with no benefits or paid leave. Minimum four hours per day and twenty hours per week, with at least four hours overlapping Pacific time — this is a real constraint, not a soft preference, and candidates outside compatible time zones are commonly filtered here. Contract runs two months with a start date roughly a week out. Pay is not disclosed in the listing; Turing sets rates per project and per candidate, so treat any figure you hear secondhand as unverified.