What the work actually involves

This is a hybrid engineering and data-authoring role. Part of the week looks like conventional product work: responsive React front-ends, Node back-end services, relational schema design against SQLite and PostgreSQL, and shipping features end to end without a separate QA layer to catch you. The other part is producing training and evaluation material — writing coding tasks with unambiguous specifications, building the reference solutions and test suites that grade them, preparing data for RL training runs, and iterating on experiments when a task turns out to be trivially solvable or unsolvable for the wrong reasons.

Task authoring is where most engineers underestimate the difficulty. A good task has a single defensible correct behaviour, tests that fail for the intended reason rather than on formatting, and a difficulty level calibrated so current models neither pass on the first attempt nor fail for want of context. You will also be asked to assess other people's tasks against benchmark criteria, which means articulating why a task is weak — leaky tests, underspecified requirements, hidden environment dependencies — rather than just flagging it.

What the screen looks for

Turing's process is AI-led early on and leans hard on specifics. Expect follow-ups on concrete React and Node decisions you made, not framework trivia: how you handled state that crossed several components, why you chose a particular Postgres index or transaction boundary, how you debugged something that only reproduced in production. The listing explicitly asks for RL and model-training exposure plus benchmark task experience, so be ready to describe an actual training or evaluation loop you touched, even a small one. Fluency with Codex, Cursor, and Claude is treated as a working requirement — screeners probe whether you review generated code critically or paste it forward.

Logistics

  • Remote and largely asynchronous, with Turing's centre of gravity in US Pacific time; some overlap is typically expected for experiment iteration and review cycles.
  • Engagements on Turing are usually contract-based and scoped by project, with hours negotiated per engagement rather than fixed.
  • Pay is undisclosed for this listing. Turing's coding engagements are commonly hourly; treat any figure you see quoted elsewhere as observed, not guaranteed.
  • Expect a code-based assessment or paid trial task in addition to the conversational screen.