What the work actually is

You write single-turn prompts in your own engineering domain and label each one benign, dual-use or adversarial. The model answers; you then judge whether it handled the request correctly against a defined policy standard, and write the reference answer — what a correct response should have contained and the technical reasoning behind it. A model that stonewalls a routine motor performance or qualification question counts as a failure in the same ledger as one that supplies something it shouldn't. Most of your effort goes into the middle band: prompts that a competent engineer would recognise as ordinary professional practice, and prompts that wear that costume while fishing for something else.

The domain sits on the dual-use line more than most. Internal ballistics, grain geometry, blast scaling, fragment velocity distributions, safe-and-arm interrupt logic — the analysis that qualifies hardware for flight is the same analysis that describes a capability, and the logic that keeps a system inert is the part an adversary wants defeated. Mercor's stated reason for recruiting practitioners rather than generalists is that only someone who has built, tested or qualified this hardware can place that line reliably and then defend the placement in writing.

What the screen is looking for

  • Hands-on design, test or qualification history with real hardware and real test data — not literature familiarity.
  • Ability to explain a judgment to a non-specialist in writing. Prior technical writing, published research or expert witness work is an explicit signal; bring a sample or link.
  • Calibration in both directions: candidates who refuse everything screen out as fast as candidates who answer everything.
  • Clean separation from classified, export-controlled, NDA-bound or prepublication-review material. Holding such obligations does not disqualify you — the listing says to disclose them so the work can be scoped around them.

Logistics and conditions

Remote and asynchronous, paid per task in an observed $65–75 band rather than a guaranteed rate; volume depends on batch availability and how your calibration holds up across review. Expect sustained reading and writing about misuse scenarios in your own field — the platform briefs experts on this in advance and states you can pause or step away at any point without penalty. Budget real time per task: the rationale, not the label, is the deliverable.