What the work actually is
You build the test material, then you grade against it. A typical task starts with authoring a clinical vignette — an intake presentation, a mid-treatment rupture, a disclosure of passive ideation during a routine session, a custody-evaluation boundary question — with enough detail that a competent clinician would recognize the decision point. You then write the reference response: not just the right call, but the reasoning chain and the safety framing behind it, at a level of explicitness that a grader with a different training background could apply consistently.
The second half is evaluation. You read model output against your own or another expert's reference and score it for clinical accuracy, ethical handling, and risk management. The failure mode this work exists to catch is the one that reads well: a response that is warm, validating, fluent, and quietly wrong about lethality, means restriction, duty-to-warn, or when to escalate. Flagging it is not enough — you write the specific clinical reason it fails, in terms an ML researcher who is not a clinician can act on.
What the screening looks for
- A verifiable license. Doctorate in psychology plus an active, unencumbered US license, and at least two years of post-licensure clinical or research practice. Expect the license number and jurisdiction to be checked.
- Real risk-assessment reps. Not familiarity with a screening instrument, but experience conducting suicide risk assessments and building safety plans, and the ability to explain your decision rule when the evidence is mixed.
- Depth in a named modality. CBT, DBT, assessment, or psychological testing — specific enough that follow-up questions about protocol adherence, measurement, or case conceptualization land.
- Written precision. Most of the deliverable is prose. Vague grading notes are the most common reason a strong clinician does not last in this work.
Logistics
Fully remote, fully async — no live sessions, no caseload, no on-call. You claim tasks and set your own hours, with a floor of about 10 hours per week to stay on active projects. Pay is hourly via Stripe on a weekly cycle; the $90–155 band is what has been observed on this platform and varies with task complexity and specialization, not guaranteed. Board certification (ABPP), academic appointment, crisis-line or suicide-prevention protocol experience, publications, or supervision history all strengthen an application but are not gates.