What the work actually involves
You are writing code that becomes training signal, not code that ships to production. A typical day mixes several task types: authoring clean, well-annotated Python or JS/TS solutions to prompts a model failed at; comparing two or more model responses to the same coding query and ranking them with a written rationale that explains why one is better; constructing supervised fine-tuning examples where the problem, the reference solution, and the reasoning are all yours; and reviewing other contributors' submissions for correctness and style. Some pods also work on RLHF reward-model refinement alongside researchers and annotators.
The rationale is as important as the code. A ranking with no explanation is close to worthless to the lab consuming it, so the habit that matters most is articulating a defensible technical judgment — an off-by-one in an edge case, an unhandled promise rejection, a solution that passes the happy path but leaks state — in a few clear sentences.
What the screen is looking for
Turing runs roughly 75 minutes: a 60-minute technical interview and a 15-minute cultural and offer conversation. The technical round probes real fluency across both stacks — ES6 semantics, Node or Nest, React/Angular/Vue, Python idiom — plus modular architecture, testing discipline, and Docker, which the listing marks as mandatory rather than nice-to-have. Expect follow-ups that go a layer deeper than your first answer; depth under probing is the actual measure. Prior QA, test-planning, or TDD experience and hands-on LLM prompting both help but are listed as optional.
Logistics
- Fully remote, contractor assignment — no medical or paid leave.
- Minimum 4 hours per day, 20 hours per week; 30 and 40 hr/week options exist.
- Four hours of daily overlap with PST is a hard scheduling constraint.
- Initial contract is one month, with a start date typically the following week. Pay band is undisclosed on this listing.