What the work actually is
You will not see patients. You will decide how clinical AI gets graded. On a grading-criteria task, you take a clinical question and decompose the ideal answer into discrete checkable items — what must be said, what must not be said, what counts as adequate hedging — so that a reviewer who is not a specialist can apply your rubric consistently. On a dialogue-evaluation task, you read a multi-turn clinical conversation and judge accuracy, safety, completeness, and whether escalation advice was right. Other streams include clinical reasoning annotation (recording the differential you rejected, not just the one you kept), output review for hallucinated findings and unsupported certainty, specialty guideline authoring, and writing difficult cases that probe the edges of model reasoning.
This is a shared expert pool. After onboarding you may be matched to any of several concurrent workstreams based on specialty, availability, and interest, and you may move between streams as priorities shift. Task length runs from roughly 45 minutes for a dialogue evaluation to an hour or more for grading-criteria authoring, and each stream comes with its own throughput target.
What the screen looks for
- Verifiable credentials — MD or DO, completed residency in any specialty, active unrestricted licence, 2+ years post-residency.
- Written clinical reasoning a non-specialist can follow. Expect to be asked to justify a judgement in prose, not to pick a rating. Vague gestures at "clinical gestalt" score poorly; the platform wants criteria someone else could reapply.
- Calibration about danger. The most valued instinct here is spotting answers that are technically correct but clinically unsafe, and confidence that outruns the evidence.
- Honest availability. The 20-hour minimum is real, and some streams are time-boxed and expect hours to be concentrated.
Logistics
Fully remote and largely asynchronous, scheduled around your own clinical or post-clinical commitments. Rates are set per workstream; $150/hr is the band observed for general clinician work on this platform and is not a guarantee. U.S. licensure and familiarity with U.S. standards of care are preferred but not required; primary care, internal medicine, emergency medicine, and hospitalist backgrounds are favoured for breadth of presentation. If you have published research, board-exam item writing, resident assessment tools, or guideline work, link a sample — it is the single strongest signal in this application.