What the work actually is
Each task is a small, self-contained package: a single-turn prompt in your area of nonproliferation practice, a label placing it on a three-level scale (benign, dual-use, adversarial), an assessment of the model's response against a written policy standard, and a reference answer explaining what a correct response looks like and why. The reference answer is where most of the effort goes — it has to carry the technical reasoning for a reader who does not know the Nuclear Suppliers Group trigger list or what a catch-all control does, and it has to hold up when another reviewer disagrees with your label.
The interesting prompts are the ones that sit on the line. An exporter asking how a specific maraging steel spec maps to a control entry is doing their job. Someone asking which specifications sit just below the threshold is doing something else, and the surface text can look nearly identical. Mercor's stated view is that over-refusal is a failure mode with equal weight to over-compliance — a model that stonewalls a licensing officer is broken in a way that matters.
What the screen is looking for
- Professional work on proliferation questions: export control licensing, classification or enforcement; treaty and safeguards implementation; pathway or programme analysis; sanctions, procurement network or illicit trade analysis; open-source technical analysis of nuclear programmes.
- Evidence you can tell a compliance question from a circumvention question, with worked examples rather than assertion.
- Writing that a non-specialist can follow. Published research, technical reports, or expert witness experience are strong signals, and you should attach a sample or link.
- Clean information hygiene. The work must not draw on classified or export-controlled material, NDA-covered content, or anything subject to prepublication review. Holding such obligations does not disqualify you — declaring them lets the work be scoped around them.
Logistics
Fully remote and asynchronous, paid per completed task, with observed rates in the $65–75 range per task (not guaranteed; volume and rates vary by project phase). Task batches arrive with a policy standard and calibration examples, and throughput expectations are set per batch rather than by weekly hours. The work involves sustained reading and writing about misuse scenarios in your own field; Mercor briefs experts on this in advance and states that you can pause or step away without penalty.