What the work actually is

The initiative produces long-horizon agent tasks — multi-step workflows an AI agent must complete inside replica environments of tools like Slack, Linear, Jira, Notion, Gmail and internal wikis — plus the rubrics that grade them. A mining team authors the tasks; connector engineers build the environments. QC sits at the end of that pipeline and decides what ships. Day to day that means reading someone else's Python, spinning up the connector environment, running the task through, and establishing whether the stated objective is actually achievable, whether the expected outcome holds, and whether the rubric would grade two reasonable attempts the same way.

A large part of the job is adversarial reading. Rubrics fail in specific ways: they accept a shortcut that skips the work, they hinge on a string match that a correct answer phrases differently, they leave a completion condition implicit, or they assume state that the environment doesn't actually produce. Tasks fail in their own ways — ambiguous scope, an edge case nobody enumerated, or a workflow that no real operator would perform in that order. When something breaks you are expected to reproduce it, localise it, and hand back a finding that the author or connector engineer can act on without a second round of questions.

What the screen looks for

  • Real backend Python depth — reading, running and debugging unfamiliar code without hand-holding, not just writing new code.
  • Comfort with the infrastructure the environments run on: GCP, Docker, virtual machines, Harbor.
  • Concrete evidence you use AI coding tools (Claude Code, Cursor, Copilot) heavily in daily work. The listing calls this a hard requirement, so expect to be asked how you use them and where you don't trust them.
  • Judgment about real-world workflows: can you tell whether a task faithfully represents how work is actually done in Jira or Slack, and argue it with specifics?
  • Written precision. Feedback is the deliverable, and screens tend to probe how you phrase a rejection.

Logistics

Fully remote contractor assignment, approximately 35 weeks, with an expected start roughly a week after selection. Minimum 4 hours a day and 20 hours a week, including a 4-hour overlap with PST — plan on late-evening or early-morning blocks depending on your timezone. No medical or paid leave. Hiring is limited to India, Pakistan, Nigeria, Egypt, Ghana, Bangladesh, Turkey and Mexico. Pay was not disclosed in the listing; Turing typically sets an hourly rate per assignment during the offer conversation, and candidates report it varies with seniority and region.