What the work actually looks like

You will sit with a patient chart and an AI-generated ED summary side by side and answer a narrow, demanding question: would an emergency physician picking this up at arrival be adequately informed, or misled? That means tracing every assertion in the output back to source documentation, marking fabricated findings and invented medication histories, and catching the quieter failure — the anticoagulant that was dropped, the prior ED visit for the same complaint that never made the summary, the triage vitals rendered in a way that flattens an unstable patient into a stable one. Annotations are written against a detailed rubric, and your free-text rationale matters as much as the label: engineering teams act on what you write, so "clinically wrong" without a specific, reproducible reason is a dead end.

Alongside per-item review you will be asked to contribute to the guidelines themselves. Edge cases in ED documentation are relentless — undifferentiated chest pain, the altered patient with no history available, hand-off notes that contradict the nursing record — and the rubric will not anticipate all of them. Physicians who flag ambiguity early, propose a defensible rule, and apply it consistently are more valuable here than physicians who simply grade faster.

What the screen is looking for

Mercor's screening is AI-led and verification-heavy. Expect confirmation of your MD, completed EM residency, and ABEM/AOBEM board certification, plus specifics on years in practice post-certification and your current clinical setting — community, academic, freestanding, volume, acuity mix. Depth probes tend to be practical rather than board-style: what belongs in an arrival summary, what you would consider a safety-critical omission versus a stylistic one, how you reason when the chart itself is incomplete. The platform also tests calibration — whether you can hold a consistent standard across items instead of drifting harsher or more lenient as fatigue sets in.

Logistics

  • Remote, part-time, US-based requirement is firm for this engagement
  • Asynchronous work through an annotation platform; no fixed shifts, though batches often carry turnaround windows
  • Typical commitments cluster around 5–15 hours per week, scheduled around clinical work
  • Expect a calibration round and guideline review before volume work opens up