What the work involves

You will be handed model outputs, prompts, and scenario drafts that touch on children — how a child of a given age might reason, what a caregiver should be told, whether a described behavior is typical or a red flag. Your job is to judge whether the content is developmentally accurate, culturally appropriate for the population it claims to describe, and safe. Tasks vary by batch: some days it is comparative grading of two model responses about, say, a four-year-old's theory-of-mind capacity; other days it is writing test cases that a model should fail, or drafting rubric language that a non-specialist annotator can apply consistently.

A recurring theme is age-banding. Models routinely flatten developmental stages — attributing metacognitive ability to toddlers, or treating a 14-year-old as interchangeable with an 8-year-old. Catching that, and writing the correction so an engineer without psychology training understands why it matters, is the core skill. Cross-cultural work matters too: much of the literature is WEIRD-sampled, and micro1's client is explicitly looking for contributors who can flag when a norm generalizes and when it does not. The named-language preference list is tied to this.

What the screen looks for

micro1's screening is AI-led and conversational, with follow-ups that go deeper when you make a specific claim. Expect to be asked to name your dissertation area, the age ranges and methods you actually worked with, and to defend a developmental judgment under pushback. Broad answers get broad follow-ups; precise ones get harder, more interesting ones. No AI or annotation experience is required, but you should be able to articulate how you would reason about an ambiguous case rather than simply asserting an answer.

Logistics

  • Fully remote contractor engagement, work delivered through a browser-based annotation platform
  • Largely asynchronous; occasional sync calls with interdisciplinary reviewers
  • Hours flexible and project-dependent — some contributors take 5–10 hrs/week, others far more when a batch is live
  • Pay observed in the $100–200/hr band; not guaranteed, and it varies with task type and language coverage