What the work actually involves
Turing is assembling LLM evaluation and training datasets from public repository histories — synthetic task construction with a human in the loop. Your job is to turn a real commit-and-issue pair into a task that can be verified automatically: pick a suitable Go repository, reconstruct the pre-fix state, get the environment reproducible in Docker, confirm the test suite actually fails before the fix and passes after, and document the task so a model's attempt can be graded without a human re-reading the diff. A meaningful share of the day is unglamorous environment work — pinned dependency versions, flaky tests, build tooling that assumes a developer's laptop.
The other half is judgment. You triage issues across trending open-source libraries and decide which ones are genuinely hard for an LLM rather than merely long, you assess whether a repo's unit tests are strong enough to serve as a grader, and you work with researchers to widen coverage across language, difficulty, and task type. There are opportunities to lead a small group of junior engineers on the same pipeline.
What the screen looks for
The stated process is roughly 75 minutes: a 60-minute technical round plus a 30-minute technical and cultural discussion. Expect concrete probing on Go tooling, Docker, and Git internals — not algorithm puzzles. Screeners are testing whether you have actually worked inside unfamiliar large codebases, whether you can tell a strong test suite from a high-coverage-but-weak one, and whether you can reason about what makes a task verifiable. Vague familiarity with open source reads badly here; name repositories, name failures.
Logistics
- Fully remote contractor assignment — no medical or paid leave.
- Minimum 4 hours per day and 20 hours per week, with 4 hours overlapping PST; 20, 30, and 40 hr/week options exist.
- Minimum 3+ years total engineering experience, with strong Go.
- Hiring is restricted to India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, and Mexico.
- Pay band undisclosed on this listing; Turing rates for comparable coding-evaluation work are typically hourly and vary by seniority and region.