What the work involves
micro1 recruits software engineers for two overlapping streams of work: model-training tasks (writing reference solutions, authoring test cases, ranking model-generated code, and writing critiques that explain why one output is better) and longer engagement placements where you are matched to a client team. On the evaluation side, a typical unit of work is a prompt plus two or more candidate completions — you decide which is correct, note compile or runtime failures, flag subtle bugs the model glossed over, and leave a written rationale a reviewer can audit. Other tasks ask you to originate the problem yourself: a realistic bug report, a repo-level refactor, a failing test the model must diagnose.
The hard part is rarely the coding. It is holding a consistent standard across hundreds of items, distinguishing wrong from stylistically not what I'd write, and resisting the pull toward the longer, more confident-sounding answer when the shorter one is actually correct.
What the platform screens for
micro1's intake is AI-led: an asynchronous interview with a conversational agent, usually paired with a live or recorded coding exercise. The agent asks follow-ups, so shallow answers unravel quickly. It is checking for:
- Real shipped experience in a named primary language, described with specifics rather than résumé phrasing
- Ability to reason out loud while coding, including about complexity and edge cases
- Judgment about code correctness that survives probing — can you defend a call when pushed?
- Clear written English, since the critique text is itself the deliverable
- Honest, stable availability numbers
Logistics
Fully remote and largely asynchronous, with your own machine and a stable connection. Most contributors work part-time in self-selected blocks; some projects ask for a minimum weekly commitment or overlap with US business hours for calibration calls. Onboarding usually includes a paid or unpaid calibration set that is reviewed before you get steady volume. Rates in the $50–100/hr range have been observed and vary by language, project, and tier — they are not guaranteed, and volume can be uneven between projects.