What the work involves

You will spend most of your time reading model output and deciding whether it holds up clinically. That means judging whether an AI's response to a described presentation is diagnostically defensible, whether it respects boundaries around risk and self-harm disclosure, and whether it reproduces assessment language accurately in Lithuanian. Alongside review, you will author material: case vignettes, intake dialogue, assessment items and clinician-facing reasoning, written natively in Lithuanian rather than translated from an English draft. Expect to write a short justification for every rating — the annotation rationale is often more valuable to the customer than the score itself.

The bilingual element is the real constraint. Lithuanian clinical vocabulary does not map cleanly onto English DSM/ICD phrasing, and much of the work is deciding when a model's Lithuanian output is technically translated but clinically wrong — or when an English-derived framing imports assumptions about family structure, help-seeking, or stigma that do not fit Lithuanian practice.

What the platform screens for

micro1 runs an AI-led interview before any human contact. It probes for:

  • A verifiable PhD in psychology and real clinical or academic practice, not coursework
  • Genuine Lithuanian fluency at professional register — expect to be asked to explain terminology in Lithuanian and account for how you would render it in English
  • Evaluation judgment: whether you can separate "fluent and confident" from "correct," and whether you apply consistent criteria across items
  • Ethical reasoning about crisis content, scope-of-practice limits, and what an AI system should refuse

No prior AI or annotation experience is expected. Candidates are far more often screened out for vague clinical answers than for unfamiliarity with model evaluation.

Logistics

Fully remote contractor engagement, invoiced hourly. Work is largely asynchronous through a web annotation tool, with occasional calibration calls where reviewers reconcile disagreements against a rubric. Hours are flexible and self-scheduled; most contributors treat this as part-time alongside practice. Volume is project-driven — batches can arrive in bursts and then pause, so treat the band as observed rates on delivered work rather than a guaranteed weekly income.