What the work actually involves

Each task is a small, self-contained package: a prompt you author in your own domain, a risk label (benign, dual-use, or adversarial), an evaluation of what the model returned, and a reference answer stating what a correct response looks like and why. The prompts that matter are the ones that sit on the line — a render-safe sequencing question a technician would genuinely ask, a magazine compatibility question a range officer needs answered, a post-blast reconstruction question phrased so that the underlying construction detail is the real payload. You are testing two failure modes at once: the model that answers something it shouldn't, and the model that refuses a routine professional question and is therefore useless to the people it exists to serve.

The written rationale is not an afterthought. Policy reviewers who have never approached a device have to follow your reasoning and reach your conclusion. "A technician would recognise this as fishing" is not a rationale; naming the specific detail the request is reaching for, and what a legitimate version of the same question would have asked instead, is.

What the platform screens for

  • Hands-on device assessment under certification. Hazardous Device School or equivalent, military EOD, or a documented post-blast investigation record. Reading on the topic does not qualify.
  • Calibration in both directions. Screeners probe whether you over-refuse as readily as whether you under-refuse. Candidates who treat every question in the domain as dangerous screen out.
  • Writing that stands on its own. Prior technical writing, published research, SOP or training-curriculum authorship, or expert witness reports. Include a sample or link.
  • Clean disclosure. No classified, export-controlled, NDA-bound, or prepublication-review material may enter this work. Holding such obligations does not disqualify you — concealing them does. Say so in the application and the work gets scoped around it.

Logistics and conditions

Fully remote and asynchronous, paid per accepted task in the observed $65–75 range rather than hourly; rates and volume vary by cohort and are never guaranteed. Most contributors work in blocks of a few hours and batch several tasks per sitting. You will be reading and writing about misuse scenarios in your field for sustained periods — the platform briefs experts on this before onboarding, and stepping away or pausing carries no penalty. Expect a short calibration round against the policy standard before your output is accepted at volume.