What the work actually involves

This is task authoring and quality control for agent evaluation, done by someone who codes. A typical day mixes investigative work — digging through real tools, APIs, and data models to understand how a piece of knowledge work genuinely gets done — with authoring: writing a multi-step objective an AI agent could plausibly be asked to complete across several SaaS applications, then writing the rubric that decides whether it succeeded. The rubric is the deliverable people underestimate. It has to define completion precisely enough that two graders agree, and tightly enough that an agent can't satisfy it by doing something superficially similar.

The other half is QA on tasks written by others. You reproduce them against the connector environments — Python backends built by a partner team to mimic Slack, Linear, Jira, Notion, Gmail, and wikis — and confirm the stated objective is actually achievable, the expected end state holds, and nothing in the setup is ambiguous or accidentally unsolvable. When something breaks, you debug far enough to hand the task author or connector engineer a specific finding, not a bug report that says "didn't work." As the initiative scales you're expected to harden the checklists and QC standards rather than just execute them.

What the screen looks for

  • Real backend Python depth, not scripting familiarity — plus working comfort with GCP, Docker, VMs, and Harbor.
  • Daily, fluent use of AI coding tools (Claude Code, Cursor, Copilot). Turing states this as a hard requirement, and expect it to be probed with specifics about your actual workflow.
  • Evaluation judgment: whether you can spot the ambiguity, the ungradeable step, or the gameable rubric before it ships.
  • Written clarity, because rubrics and QA notes are the artifacts you're paid for.
  • API/data-model familiarity with the named SaaS tools and any prior agentic/LLM evaluation work are advantages, not gates.

Logistics

Remote contractor assignment with no medical or paid leave. Roughly five weeks, with a start date stated as next week — so availability is a genuine filter, not a formality. Commitment is at least 6 hours per day and a minimum of 40 hours per week, with six hours of overlap with PST; this is not an async role. Hiring is restricted to India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, and Mexico. Pay is undisclosed on the listing.