What the work actually looks like
You will spend most of your hours reading model output and deciding whether it would survive scrutiny inside a regulatory body. Typical items include a drafted eligibility determination, a permit or licence review with a recommendation, a violation report, or a regulatory analysis of a fact pattern. Some tasks ask for a single graded critique against a rubric; others are head-to-head comparisons where two model responses must be ranked on regulatory rigour, evidentiary soundness, and the quality of enforcement reasoning — then justified in writing. The written justification is the deliverable that matters most: reviewers are calibrated against each other, and a rating without a traceable reason is treated as noise.
What the screen is looking for
Mercor's screening is AI-led and follow-up heavy. It probes whether you have personally worked a file end to end — reviewed evidence, applied a specific statute or rule, written a finding someone else acted on — rather than supervised people who did. Expect to be asked which regulatory regimes you know by name, what standard of proof or sufficiency you applied, and how you handled a case where the facts were incomplete. Vague answers get one clarifying question, then a lower score. Strong candidates cite the rule, the jurisdiction, the decision they made, and what happened next.
Logistics
- Fully remote and largely asynchronous; work is pulled from a task queue rather than scheduled in meetings.
- 20+ hours per week is a real threshold, not aspirational — throughput is monitored.
- Duration is roughly 4–6 weeks with an immediate start. Extensions happen but are not promised.
- Payment begins as task-based, approved per task; experts who clear their first task inside the stated window move to hourly for the rest of the project. Rates in the $70–110/hr band are what contributors have reported, not a guarantee.
- Strong written English is load-bearing here, since every rating carries a prose rationale.