What the work actually is
You pick problems from your own research territory — the ones where a plausible-looking answer and a correct one diverge in ways only a practitioner catches. Then you write the problem statement so it is unambiguous, author the reference solution at full expert depth, and build a rubric that another qualified scientist could apply and land on the same score you did. After that, you grade model output against it: correctness first, then whether the reasoning path was sound or arrived at the right number by a wrong route.
Typical artifacts include a Lean or Coq formalization with the informal statement it corresponds to, a condensed-matter derivation with the approximations named and justified, a genomics pipeline question where the trap is a silent assumption about reference build or read orientation, or a synthesis problem where the model proposes a route that is chemically legal and practically absurd. Rubrics are expected to be discriminating — a rubric that gives full marks to three visibly different-quality answers gets sent back.
What the screen looks for
- Currency. Recently or currently hands-on in the field, not adjacent to it. Screeners probe for what you were computing, proving, or measuring in the last year or two.
- Depth under follow-up. Expect a claim you make in answer one to be pressed in answer two. Shallow familiarity surfaces quickly.
- Written precision. Much of the deliverable is prose. If your explanation of a subtlety is muddy, the artifact is unusable regardless of whether you were right.
- Grading judgment. Prior peer review, code review, TA grading, or editorial work is a strong signal — you already know the difference between an error and a stylistic disagreement.
- Track bonus signal. Theorem proving (Lean/Coq), computational biology and genomics, condensed-matter or quantum physics carry extra weight.
Logistics
Fully remote, fully asynchronous, no standing meetings. Work arrives in batches tied to project cycles, so weekly volume fluctuates; most contributors treat this as a 5–15 hour commitment alongside existing research or industry work. Deliverables carry deadlines even though hours are self-scheduled. Rate placement within the $100–170 band reflects experience and track scarcity, and is set at onboarding rather than negotiated per task.