What the work actually involves
You receive AI assistant responses — often long, often plausible-sounding, sometimes subtly wrong — and score them against a rubric that specifies dimensions like factual accuracy, instruction adherence, relevance, and reasoning quality. Alongside the score you write a short written justification: what the model got right, where it went wrong, and what a correct response would have looked like. A single task might take five minutes or forty, depending on how much verification the claim requires.
The harder part is consistency. You will see dozens of near-identical examples in a sitting, and the project cares more that your scores are reproducible and defensible than that any single score is elegant. Expect to flag ambiguous cases, participate in rubric-interpretation threads, and occasionally have your batch audited against a gold-standard set or against other reviewers' scores.
What the screen is looking for
- Genuine domain depth. micro1 recruits across many fields; the interview will push on your specific expertise to confirm you can catch errors a non-expert would miss.
- Evaluation habits, not AI credentials. Grading, editorial review, QA, peer review, clinical audit, code review, annotation — any history of applying a standard repeatedly and defending your calls.
- Real, daily use of AI assistants. They want people who already know how these models fail: confident fabrication, dropped constraints, broken tool calls, reasoning that looks like a chain but isn't.
- Written clarity under compression. Feedback is read by people tuning models, not by your peers. Vague criticism is unusable.
Logistics
Contractor engagement, fully remote, restricted to the US, Canada, UK, Ireland, Australia, and New Zealand. Work is asynchronous and drawn from a queue — there are no fixed shifts, but projects have throughput expectations and can surge or pause without much notice. Most contributors treat it as part-time, 10–25 hours a week; rate placement within the band depends on domain scarcity and on how you perform on calibration batches. Pay figures here are as observed on the listing and are not guaranteed.