What the work actually is

This is code review as a measurement discipline. You are handed software engineering tasks and model-generated attempts at them, and you decide whether the solution is actually correct — not whether it looks plausible. That means reading the diff, reproducing the issue in a real environment, running the code, probing edge cases the model didn't consider, and separating "passes the happy path" from "meets the requirement." Alongside grading, you help build the instrument: writing evaluation datasets, refining grading rubrics so two reviewers reach the same verdict, and documenting failure modes where frontier models reliably break down on software problems.

The languages in scope are Python, Java, C/C++, Go, Swift, Objective-C, PHP and SQL — depth in one or two matters far more than surface familiarity with all of them. The stated screening is an automated coding challenge in Python plus a Docker test, so comfort working inside containers and reproducing builds is part of the gate, not an afterthought.

What the platform screens for

  • Verifiable engineering history — a CS or related degree and 3+ years professional experience, with specifics: codebases, scale, languages, what you owned.
  • Code review experience in production or large-scale systems, not just personal projects. Screeners push on how you handled a review where you and the author disagreed.
  • Debugging methodology under follow-up. Expect to be asked how you'd reproduce and isolate a failure, then asked again with a constraint removed.
  • Rubric judgment. Can you articulate why a working-but-unmaintainable solution scores differently from a clean one, and hold that line consistently across a hundred samples?
  • Genuine availability. Four hours a day, twenty hours a week minimum, four hours overlapping PST, starting roughly next week.

Logistics

Fully remote, US candidates only, contractor assignment with no medical or paid leave. The contract is written for one month; renewal on Turing engagements depends on project volume and reviewer agreement scores, and nothing in the listing promises extension. Pay is undisclosed here — ask for the hourly rate and the expected weekly volume before you accept, and confirm whether rubric-authoring time is billable at the same rate as grading. Work is largely asynchronous within the overlap window, with calibration discussions and quality standards set collaboratively with project teams.