What the work actually involves

Turing is staffing a frontier-lab data initiative split into two tracks: Connectors and Tasks. On the Connectors side you write Python backend applications that faithfully reproduce the behaviour of real SaaS products — Slack threads, Linear issue state transitions, Notion page trees, Gmail threading — so that an AI agent interacting with them cannot tell it is in a sandbox. On the Tasks side you mine how knowledge work is really done in those tools, then author multi-step tasks that force an agent to reason across several calls and surfaces, and write rubrics that state exactly what counts as complete and correct. Most days involve both: build the environment, break it, then write the task that exposes whether an agent can survive it.

QA is not a separate phase here. You are expected to validate tasks end to end for realism, achievability and gradability, hunt for ambiguity and grading gaps before anything ships, reproduce other engineers' failures, and give feedback specific enough to act on. Heavy daily use of AI coding tools — Claude Code, Cursor, Copilot or equivalent — is stated as a hard requirement, not a preference, and it applies to debugging and QA workflows as much as to writing code.

What the screen looks for

Expect probing on Python backend depth rather than framework trivia: how you test and debug services, how you reason about state and idempotency, how you'd model an API you've only read docs for. Working familiarity with GCP, Docker, VMs and Harbor comes up, as does concrete evidence you can drop into an unfamiliar codebase without hand-holding. The judgment half of the screen is about evaluation: given a task an agent could plausibly solve three different ways, can you write a rubric that still grades it fairly? Written communication is weighted unusually high because rubrics and QA notes are the deliverable, not a side artefact.

Logistics

  • Remote, contractor assignment — no medical or paid leave.
  • Minimum 6 hours per day, minimum 40 hours per week, with 6 hours overlapping PST.
  • Contract stated as 5 weeks, with a start date roughly a week out; extensions on programmes like this happen but are not promised.
  • Hiring restricted to India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey and Mexico.
  • Pay band is not disclosed in the listing; rates on Turing coding pods vary by track and seniority and are set in the recruiter conversation.