What the work actually involves

Turing is assembling LLM evaluation and training datasets from public repository histories — synthetic task construction with a human in the loop. Your job is to turn a real commit-and-issue pair into a task that can be verified automatically: pick a suitable Go repository, reconstruct the pre-fix state, get the environment reproducible in Docker, confirm the test suite actually fails before the fix and passes after, and document the task so a model's attempt can be graded without a human re-reading the diff. A meaningful share of the day is unglamorous environment work — pinned dependency versions, flaky tests, build tooling that assumes a developer's laptop.

The other half is judgment. You triage issues across trending open-source libraries and decide which ones are genuinely hard for an LLM rather than merely long, you assess whether a repo's unit tests are strong enough to serve as a grader, and you work with researchers to widen coverage across language, difficulty, and task type. There are opportunities to lead a small group of junior engineers on the same pipeline.

What the screen looks for

The stated process is roughly 75 minutes: a 60-minute technical round plus a 30-minute technical and cultural discussion. Expect concrete probing on Go tooling, Docker, and Git internals — not algorithm puzzles. Screeners are testing whether you have actually worked inside unfamiliar large codebases, whether you can tell a strong test suite from a high-coverage-but-weak one, and whether you can reason about what makes a task verifiable. Vague familiarity with open source reads badly here; name repositories, name failures.

Logistics

  • Fully remote contractor assignment — no medical or paid leave.
  • Minimum 4 hours per day and 20 hours per week, with 4 hours overlapping PST; 20, 30, and 40 hr/week options exist.
  • Minimum 3+ years total engineering experience, with strong Go.
  • Hiring is restricted to India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, and Mexico.
  • Pay band undisclosed on this listing; Turing rates for comparable coding-evaluation work are typically hourly and vary by seniority and region.