What the work actually is

You are not seeing patients and you are not signing off on live diagnosis. You are deciding how clinical AI gets measured. On a grading-criteria task, you take a clinical question and decompose the ideal answer into discrete checkable items — what must be named, what must be ruled out, what warrants a hedge — so that a dozen reviewers grade the same response the same way. On a dialogue evaluation, you read a multi-turn clinical conversation and judge accuracy, safety, completeness, appropriateness of hedging, and whether the escalation advice would have held up. Other streams involve annotating your own reasoning path including the differentials you rejected, flagging hallucinated findings and dangerous omissions in model output, authoring specialty guidelines that keep a large annotator group consistent, and writing hard cases designed to break current model reasoning.

Because this is a shared pool rather than a single project, you onboard once and get matched to whichever concurrent stream fits your specialty, availability, and interest. Streams shift; you may move between them. Each one comes with its own throughput target. Task lengths run roughly 45 minutes for a dialogue evaluation up to an hour or more for grading-criteria authoring, so the throughput expectation is measured in completed items, not clicks.

What the screen looks for

The platform verifies MD or DO with completed residency, an active unrestricted licence in your country of practice, and at least two years post-residency. Beyond credentials, the screen probes whether you can articulate clinical reasoning in writing that a non-specialist reviewer can follow and audit — the most common failure among strong clinicians is compressed expert shorthand that a reviewer cannot check. Expect follow-ups that push on how you'd handle an answer that is technically correct but clinically unsafe, how you set a threshold between a defensible hedge and evasion, and where your specialty's edge cases actually sit. Board certification, U.S. licensure and familiarity with U.S. standards of care, breadth-heavy backgrounds (primary care, IM, EM, hospitalist), and prior annotation, medical education or question-writing experience are all preferred, not required. If you have published research or sustained technical writing, link a sample.

Logistics

  • Fully remote and largely asynchronous, with a hard floor of 20 hours per week.
  • Some streams are time-boxed, so you need the ability to concentrate hours in a compressed window rather than spreading 20 hours thinly across seven days.
  • Pay is observed at $150/hr for this pool; rates and stream availability vary and are not guaranteed.
  • No patient care, no on-call, no clinical liability for model output.