What the work actually is
SWE Bench-style evaluation means taking real, production-like repositories and turning the work inside them into tasks an AI model can be scored on. For this variant the subject matter is data: ingestion and transformation pipelines, feature preparation, analysis notebooks turned into modules, validation and data-quality logic. You will write Python that processes structured and unstructured data, run it locally, and then construct the task around it — a clear problem statement, a reproducible environment, and tests or checks that pass only when the intended behaviour is present. A large share of the effort goes into the parts that are easy to skip: pinning dependencies, removing hidden state, making sure a fix cannot be reward-hacked by hardcoding an expected output.
The second half of the job is review. You will read other contributors' tasks and judge whether the failure is genuine, whether the tests are discriminating, and whether a competent engineer could solve the task from the statement alone. Collaboration with Turing's researchers shapes difficulty calibration — tasks that every model solves and tasks no model can attempt are both low value.
What the screen looks for
- Three or more years of hands-on work as a data engineer, data scientist, or data-focused software engineer — the screen probes for specifics, not titles.
- Python depth in a data context: pandas/PySpark or equivalent, plus the ability to navigate and modify a codebase you did not write.
- ML and data science fundamentals solid enough to judge whether a model's data-handling answer is correct or merely plausible.
- Reproducibility instincts. Candidates who talk about environment pinning, deterministic outputs, and non-leaky tests do noticeably better than those who only discuss the happy path.
- A 60-minute technical interview with a live coding challenge is the gate. Expect data-intensive or algorithmic problems, spoken through in English.
Logistics
Fully remote contractor assignment, initially three months and adjustable by engagement. Minimum four hours per day and twenty hours per week, with at least four hours overlapping Pacific time — so contributors in South Asia or Africa should plan on late-evening or very early-morning blocks. No medical or paid leave. Hiring is limited to India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, Brazil, and Mexico. Pay was not disclosed in this listing; Turing typically sets an hourly rate during the offer stage.