What the work actually is

This is a dataset-construction role dressed in normal engineering clothes. You pick through public C# repositories — ideally well-maintained, actively developed ones — and find issues and commits that can be reconstructed as a verifiable task: a broken state, a fix, and tests that unambiguously distinguish the two. That means reading unfamiliar codebases quickly, getting them to build and test reproducibly inside Docker, checking whether the existing unit tests actually cover the behaviour in question, and writing or tightening tests when they don't. You then run models or agent traces against the task and assess where and why they fail, working with researchers who are hunting for repositories and issues that genuinely defeat current LLMs.

What the screen looks for

Turing runs roughly 75 minutes across two rounds — a 60-minute technical interview plus a 30-minute technical and cultural conversation. Expect concrete C# and .NET questions (project vs. solution structure, SDK versioning, xunit/NUnit, mocking, flaky async tests), plus practical Docker and CI questions: how you'd containerise a .NET repo you've never seen, how you'd pin a build so it reproduces in six months. The judgment half matters as much as the technical half — they want to hear you distinguish a good candidate issue from a bad one, spot a test that passes for the wrong reason, and explain what makes a task hard for a model rather than merely hard for a human.

Logistics

  • Fully remote contractor assignment; no medical or paid leave.
  • Hiring is restricted to India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, and Mexico.
  • Minimum 4 hours per day and 20 hours per week, with 4 hours of overlap with PST. Commitment tiers of 20, 30, or 40 hours per week.
  • Pay is undisclosed on this listing; Turing typically quotes an hourly rate at offer stage that varies by tier and region.
  • Some contributors are given the option to lead a small group of junior engineers on the same pipeline.