What the work involves
You write single-turn prompts in your own area of chemistry and label each one across three tiers: benign, dual-use, and adversarial. The model answers; you then score the response against a written policy standard and decide whether it was handled correctly. Finally you author the reference answer — what a correct response should have said, where it should have stopped, and the technical reasoning for why that line sits where it does.
The difficulty is calibration rather than obscurity. An over-refusal on a routine undergraduate mechanism question is as valuable a finding as a model that hands over a workable procedure. In chemistry the boundary usually runs between explaining mechanism and supplying actionable procedure — stoichiometry, conditions, workup, purification, sourcing — and Mercor's position is that only someone who has actually run the chemistry can locate that boundary reliably.
What the platform screens for
- Hands-on background in synthesis, process chemistry, analytical chemistry, chemical safety, industrial hygiene, or forensic/toxicological work — bench or plant experience, not a reading acquaintance.
- Evidence you can write. Publications, SOPs, process hazard analyses, expert reports, or safety documentation all count; the application asks for a sample or link, and it matters more here than in most eval roles because every score needs a rationale a non-specialist can audit.
- Judgment under follow-up: expect the screener to push on a case where your benign/dual-use call is arguable and see whether your reasoning holds or collapses into a rule of thumb.
- A clean compliance position. The work must not touch classified, export-controlled, NDA-bound, or prepublication-review material. Holding such obligations is not automatically disqualifying — Mercor asks you to disclose so scope can be adjusted.
Logistics
Remote and asynchronous, paid per task in an observed band of $65–75. Volume varies with campaign demand rather than a fixed weekly commitment, and tasks are self-scheduled. You will spend sustained stretches reading and writing about misuse scenarios in your own field; Mercor briefs experts on this beforehand and states you can pause or step away without penalty. Prior red-teaming or model evaluation experience is preferred but explicitly not required.