What the work actually involves
You are not answering medical trivia for a chatbot. The task is building the instruments that decide whether a model's clinical reasoning holds up: authoring scenarios where the diagnosis hinges on a detail a generalist model would skip, writing rubrics that separate a confident wrong answer from a correctly hedged one, and grading model transcripts against those rubrics. Expect to construct cases with realistic ambiguity — incomplete histories, conflicting labs, comorbidities that change first-line management — and then to document precisely why one response fails and another passes. Researchers on the other side will ask you to convert your clinical intuition into criteria another reviewer could apply and reach the same verdict.
What the screen is looking for
Turing's screening is AI-led and pushes on depth. Expect follow-ups on any specialty claim you make: if you say you manage decompensated heart failure, you should be ready to talk about where guideline-directed therapy gets contested in practice. The screen also probes evaluation judgment — whether you can distinguish a model output that is wrong from one that is merely unconventional, and whether you can state a failure mode in terms a non-clinician engineer can act on. Active clinical practice matters here; the value of the role rests on current bedside familiarity, not on remembered training.
Logistics
- Remote, asynchronous, scheduled around clinical commitments
- Up to 30 hrs/week, but genuinely flexible — many contributors work far fewer
- Initial duration of one month, extensions offered on performance and fit
- Pay band undisclosed by the platform; ask directly during screening rather than assuming a rate
Any specialty is eligible. What differentiates strong candidates is the ability to write down the reasoning that usually stays tacit — the step between reading a presentation and knowing what to do next.