What the work actually involves
You read model output — therapeutic dialogue turns, case formulations, assessment interpretations, research vignettes — and judge whether it is clinically sound, ethically defensible, and linguistically natural in your language. Most tasks are a mix of rating against a rubric and writing the justification: what the model got wrong, what a competent clinician would have said instead, and what evidence supports that. Some batches ask you to author source material rather than grade it: case studies, patient-presentation dialogues, or culturally-situated prompts that probe whether the model handles distress idioms, family structures, and help-seeking norms that don't map onto a North American clinic.
The bilingual element is the point. A model can produce fluent Tagalog or Bulgarian and still be clinically useless — flattening somatic presentations of depression, mistranslating diagnostic terms that carry different weight in local practice, or applying a safety script that no one in that context would recognise. You are being paid to catch the failures that a translator without a clinical background, and a clinician without the language, both miss.
What the screen looks for
micro1 runs an AI-led interview before any human contact. Expect it to verify the doctorate or residency, confirm native-level command of a specific target language on the list, and then push on clinical depth with follow-ups — a claim about CBT for a presentation you named will get a second and third question. It also checks whether you can write the reasoning, not just the verdict: vague criticism of an AI answer scores badly, a specific correction with a named framework or guideline behind it scores well. Cultural-ethical judgment is probed directly through scenarios involving suicide risk, disclosure, and family involvement.
Logistics
- Fully remote, contractor engagement, no AI or ML background expected.
- Largely asynchronous task queues; some projects add recorded verbal-response tasks or brief calibration calls.
- Volume fluctuates by language — commitments in the 10–20 hr/week range are common, but a queue can go quiet between batches.
- Pay of $100–200/hr reflects rates observed on this listing and varies by language scarcity and project; it is not a guarantee.