What the work involves
You build the test cases frontier labs use to find out whether a model actually reasons about clinical documentation or just pattern-matches to familiar code strings. A typical task starts from chart documentation — an office note, an op report, a discharge summary — and runs through to a final code set: ICD-10-CM diagnoses, CPT or HCPCS procedures, modifiers, sequencing. Then you write the reference answer, and the reference answer is the real deliverable: not just the codes, but the guideline citation behind each choice. ICD-10-CM Official Guidelines section and paragraph, CPT parenthetical, NCCI edit, CMS policy. If you can't cite it, the scenario isn't finished.
The second half of the work is grading. You read model output and judge it on two separate axes: did it get the codes right, and would the code set survive an audit. Those come apart constantly. A model will produce something that looks clean — plausible codes, confident rationale — that unbundles, upcodes an E/M level the documentation doesn't support, or misses a required Z code. Flagging that gap is the highest-value thing you do here.
What the platform screens for
- A live credential. CPC (AAPC) or CCS (AHIMA) in good standing, plus at least two years of production coding, not classroom or externship hours.
- Guideline recall under follow-up. Screeners ask you to justify a code choice and then push on the justification. Naming the guideline and its logic matters more than reciting the code.
- Audit instinct. Whether you can look at a defensible-looking code set and name the specific exposure — medical necessity, unbundling, level-of-service support, query-worthy documentation.
- Written explanation. Your rationale is what a model gets trained against, so it has to be unambiguous to another coder who never saw the chart.
Logistics
Fully remote, fully async, no production quotas and no shift blocks. You set your hours week to week with a floor of ten. Pay is hourly via Stripe, disbursed weekly; the $65–95 range is what contributors on this listing have been observed at and varies with specialty depth, additional credentials such as CIC, CRC or CPMA, and task complexity — it is not a guaranteed rate. Work is ongoing rather than a fixed engagement, with volume rising and falling as labs open new coding domains.