What the work involves

You build clinical test cases the way you'd build a case for a student you were precepting — but with the answer key written out in full. A typical unit of work is a scenario authored from chief complaint through history, exam findings, relevant labs or imaging, and disposition, paired with a reference answer that spells out why each decision follows from the data. From there you move to grading: reading model-generated workups and plans, scoring them for diagnostic accuracy, and flagging the ones that read fluently but would hurt someone — the missed red flag, the drug interaction nobody caught, the discharge that should have been an admission.

That last category is the point of the job. Frontier labs already have models that produce confident, well-formatted clinical prose. What they lack is a reliable signal for when that prose is wrong in a way a patient would feel. Your written rationale matters as much as your score, because it's the rationale that trains the grader.

What the platform screens for

  • Verifiable licensure and certification — active NP license plus AANP or ANCC board certification, and at least two years of practice after certification.
  • Genuine specialty depth — family, acute care, psych, or adult-gerontology. Screeners follow up inside your stated specialty, and vague answers get probed rather than accepted.
  • Written clinical reasoning — you'll be asked to explain a decision in prose, not just state it. Documentation habits from chart review, precepting, or QA work read well here.
  • Calibrated judgment on model output — the ability to distinguish a plan that is merely unconventional from one that is unsafe, and to say which is which without hedging.

Logistics

100% remote, fully async, no call and no clinic hours. You set your own hours week to week with a floor of about 10; contributors who take on more scenario authoring generally log more. Pay is hourly via Stripe on a weekly cycle, in the $80–125 range as observed on the platform — the band typically tracks specialty scarcity and whether you're authoring or grading. Project work is ongoing rather than a fixed engagement, so volume fluctuates with lab demand.