What the work actually involves
This is evaluation design work, not annotation piecework. You will write clinical scenarios that put a model under genuine diagnostic pressure — ambiguous presentations, competing differentials, incomplete histories, situations where the guideline-correct answer and the patient-correct answer diverge — and then build rubrics that let a non-clinician reviewer tell a sound reasoning chain from a fluent but wrong one. Expect to document why a model's answer fails: whether it anchored early, ignored a red flag, cited a superseded guideline, or produced a defensible plan with a hallucinated justification.
The second half of the job is collaboration. AI researchers on the other side of the table generally do not have medical training, so your written failure analyses need to be legible to them: what the model missed, why it matters clinically, and what a corrected trajectory looks like. Physicians who can only say "this is wrong" without decomposing the reasoning error tend not to be extended.
What the screen looks for
- Active licensure and current practice. Any specialty. Turing states active clinical practice as a requirement, not merely a licence in good standing.
- Evidence-based medicine fluency under follow-up. Expect probes on where a specific recommendation comes from, what the evidence quality is, and what you do when guidance conflicts.
- Judgment about model output. Scenario questions where a model's answer is superficially reasonable and you must locate the flaw.
- Written clarity. Much of the deliverable is prose — scenarios, rubrics, failure write-ups.
Logistics
Fully remote and asynchronous, scheduled around clinical commitments. The ceiling is 30 hours per week; most physicians on engagements like this work well below it. Initial duration is one month with extension contingent on performance and fit. Pay is undisclosed on this listing — ask directly during the screen rather than assuming a band from comparable medical evaluation work.