What the work actually looks like
You receive a technical prompt and two or more AI-generated responses to it. Your job is to decide which is better and say precisely why, in writing. The prompts skew toward technical communication rather than raw code generation: explaining a stack trace, commenting a function, drafting an API reference, describing why a service was split, walking a junior engineer through a migration. A typical task is scoring on rubric dimensions — correctness, clarity, tone, audience fit, usefulness — then composing a rationale that a reviewer who has never seen the prompt can follow.
The hard cases are the interesting ones. A response can be factually flawless and still useless because it answers a staff engineer's question in the register of a beginner tutorial. Another can read beautifully, cite a plausible flag on a real CLI tool, and be quietly wrong. Catching misleading fluency — output that carries the cadence of expertise without the substance — is the skill the project is buying.
What the screen is looking for
- Real production engineering experience in any modern stack; the specific language matters less than having shipped and maintained things.
- Evidence you have reviewed other people's work — code review, design review, doc review — and can articulate a standard rather than a reaction.
- Native-level written English. Your rationales are the deliverable, and they are read closely.
- Consistency: can you apply the same rubric the same way on task 200 as on task 3, and notice when your own calibration has drifted.
Prior AI, annotation, or RLHF experience is welcome but explicitly not required. micro1's screening is AI-led and conversational, with follow-ups that push on any claim you make, so answer with cases you can actually describe in detail.
Logistics
Fully remote, contractor engagement, asynchronous. You pull tasks and work against agreed deliverables and turnaround windows rather than fixed shifts; coordination with project managers and reviewers happens in writing. Volume varies by project phase, and many contributors run this alongside a primary role. Expect calibration rounds early on where your scores are compared against reviewers', and expect that feedback to change how you score.