What the work actually involves

You receive model outputs — a proposed test method for a Class II device, a biocompatibility rationale, a risk analysis, a signal-processing approach for a physiological monitor — and judge whether they would survive contact with a real design history file. That means checking not just whether the engineering is correct but whether the reasoning is traceable: does the verification protocol map to a design input, does the acceptance criterion have a basis, does the cited standard actually say what the model claims it says. A second stream of work is authoring: writing the expert demonstration yourself, at the level of detail a competent engineer would put in a DHF, so the lab has a reference trajectory to train against.

Ranking tasks pair two or more model responses and ask you to order them with written justification. The justification is the product. A ranking with no reasoning is close to worthless to the lab; a ranking that says response B invents a 60601-1 clause number and misstates the leakage current limit for a Type BF applied part is exactly what they pay for.

What the platform screens for

  • Verifiable regulatory exposure. 510(k), PMA, ISO 13485, IEC 60601, ISO 10993 — screeners probe whether you did the work or read about it. Expect follow-ups on specifics: which clause, which test, who signed.
  • Domain depth in one area. Biomechanics, biomaterials, imaging, instrumentation — breadth is fine, but a screen usually drills into whichever one you claim.
  • Willingness to say "I don't know." Fabricating a standard number in a screen is close to fatal, since detecting exactly that behaviour in models is the job.
  • Written clarity under a word budget. Most rubrics cap justification length. Dense, specific prose beats long prose.

Logistics

Fully async — no standing meetings, no fixed hours. Task batches appear and are claimed; throughput and consistency with the rubric drive how much work you're offered. Most contributors run 5–20 hours a week alongside a day job. Pay is hourly against logged, reviewed task time; the $75–140 band is what has been observed on this platform and moves with degree level, regulatory depth, and calibration scores, not a guarantee.