The work

You build the test material frontier labs use to find out whether a model can reason like a clinician. That means authoring case scenarios that read like real patients — chief complaint, medical history, intraoral findings, periapical or bitewing imaging described in enough detail to be diagnosable — and then writing the reference answer, including why one treatment path is defensible and the alternatives are not. The reasoning is the deliverable; a correct answer with no articulated pathway is close to useless to a lab.

The other half is grading. You read model output on cases like yours and score it for diagnostic accuracy, adherence to standard of care, and whether the recommended sequencing makes sense. A large part of the job is catching output that is fluent and wrong: fabricated guideline citations, extraction recommendations where a tooth is restorable, missed medical contraindications, endodontic diagnoses that skip pulp testing. You flag it and state precisely what makes it unsafe or out of scope — vague dissatisfaction does not train anything.

What the screen looks for

AfterQuery's screen is AI-led and follow-up heavy. It verifies the DDS or DMD and an active, unencumbered US license, then pushes on clinical breadth: restorative, endodontics, periodontics, and treatment planning are all in scope, and reviewers are checking whether you can reason across them rather than only in your comfort zone. Expect probing on radiographic interpretation and on how you'd distinguish an acceptable-but-different treatment plan from a wrong one. Written clarity is assessed directly from your typed answers.

Logistics

  • 100% remote, fully async, no scheduled calls or chair time
  • Minimum 10 hours per week; you set the schedule
  • Paid hourly via Stripe, weekly, in the observed $100–155/hr band — placement typically reflects specialty training and radiographic depth
  • Ongoing project work rather than a fixed-term engagement; volume varies by lab demand