What the work involves

You build the test material frontier labs use to find out where a model's medical interpretation actually breaks. That means authoring realistic encounters — a triage intake, an informed consent conversation, a discharge instruction delivered under time pressure — and then writing the reference interpretation alongside a written rationale for each terminology choice you made. The rationale matters as much as the rendering: annotators and model trainers need to know why derrame was resolved as pleural effusion and not stroke in this context, or why a patient's colloquial description of pain was rendered at register rather than upgraded into clinical vocabulary.

The grading half of the job is where the credential earns its band. You score model output on accuracy, register, and clinical meaning, and the highest-value flags are the fluent ones — output that reads smooth and professional but has quietly shifted dosage, negation, hedging, or the speaker's certainty. A model that renders "we should probably rule out" as a definitive diagnosis has produced good prose and a clinical error, and the platform is paying for people who catch that consistently and can explain the harm in writing.

What the platform screens for

  • A verifiable credential — CMI, CHI, or a recognised equivalent, plus at least two years actually interpreting in clinical settings rather than general community or legal work.
  • Terminology command in both directions. Screens probe register control, false cognates, and how you handle terms with no clean equivalent in the other language.
  • Written English good enough to document a decision. Your rationale is the deliverable; interpreters who can do the work but can't explain it in writing don't pass.
  • Judgment about error severity — distinguishing a stylistic difference from a meaning-changing one, and ranking which errors are clinically dangerous.

Logistics

Fully remote, fully async, no live sessions or scheduled shifts. You set your hours each week against a 10-hour minimum; payment is weekly via Stripe. The $55–95 range is as observed on the platform and typically tracks language pair scarcity, specialty depth, and how much grading versus authoring a given batch requires — treat it as a band, not a guarantee. Work is ongoing and batch-based, so volume fluctuates with what labs are currently testing.