What the work involves

You write single-turn prompts in your own technical area and label each one benign, dual-use, or adversarial. The interesting cases sit in the middle: the enrichment-cascade question a graduate student would genuinely ask, the material-balance query that reads as accountancy practice until you notice what it would let someone infer. You then read the model's response, judge it against a written policy standard, and produce a reference answer — what a correct response should have said, what it should have withheld, and the technical reasoning behind that line.

The rationale is the deliverable as much as the label. A reviewer with no nuclear background has to be able to follow why a response that looked helpful was actually over-disclosing, or why a refusal was a false positive that failed a legitimate researcher. Expect roughly half your time in prose.

What the platform screens for

  • Fuel-cycle depth, end to end — enrichment, reprocessing, criticality — plus real exposure to how material is controlled, accounted for, or diverted. Safeguards, MPC&A, and nonproliferation backgrounds are the strongest signal.
  • Calibrated judgment, not blanket caution. Screeners probe whether you can defend an answer as correct, not just refuse everything nuclear.
  • Writing evidence. Published research, technical reports, or expert witness work; have a sample or link ready.
  • Clean provenance. Nothing you contribute may draw on classified or export-controlled information, NDA-covered material, or anything under prepublication review. Existing obligations are not automatically disqualifying — disclose them and the work gets scoped around them.

Logistics

Remote and asynchronous, paid per task in an observed band of $65–75, with volume varying by campaign rather than a fixed weekly commitment. The material is misuse scenarios in your own field, read and written about for sustained stretches; Mercor briefs experts on this in advance and states you can pause or step away without penalty.