What the work actually is

You pick problems from your own working life — the ones where a competent generalist would produce something plausible and wrong — and turn them into evaluation items. That means writing the scenario with enough context to be solvable, authoring a reference solution at the level you'd defend to a peer, and building a rubric that separates a correct answer from a merely fluent one. Then you grade model outputs against it, and your written justification for each score matters as much as the score. Tracks span software and systems engineering, formal methods and computational science, ML inference and GPU kernels, enterprise operations, security, hardware design, and creative technology; you apply to one where you meet the experience bar, not to all seven.

What the screen looks for

  • Currency. Hands-on and recent. Someone who ran GPU kernels in 2019 and has managed since then is a weaker fit than someone shipping now.
  • Track depth under follow-up. Expect the screener to pick a claim from your background and push two or three layers past the summary — tooling versions, failure modes, why you chose one approach over the alternative.
  • Rubric instinct. Can you articulate what distinguishes an excellent answer from an adequate one in your domain, in writing, specifically enough that a second grader would agree with you?
  • Writing. A large share of the deliverable is prose: problem statements, solution walkthroughs, structured feedback. Terse answers that assume the reader shares your context read as a risk here.
  • Realistic capacity. Async and flexible, but the cohort needs people who actually deliver batches.

Logistics

Fully remote and asynchronous, designed to sit alongside a primary job. Volume is typically negotiated in batches rather than fixed weekly hours, so state a number of hours you can hold for several consecutive weeks rather than an optimistic ceiling. Pay is banded $100–170/hr as observed on the platform, varying by track and seniority — not a guaranteed rate, and not uniform across the seven tracks. Prior peer review, code review, grading, or editing experience is the single most useful non-domain signal, and published or open-source work in your track helps verification move faster.