What the work involves
This is repository-grounded data work, not tutorial writing. A typical unit starts with a real issue or pull request in a public codebase: you reproduce the failure locally or in a container, pin the commit and environment so it fails deterministically, write or adapt tests that fail before the fix and pass after, and document the reasoning a competent engineer would follow. On evaluation shifts the direction reverses — you receive a model- or agent-generated patch and decide whether it genuinely resolves the issue, passes for the wrong reason, breaks something adjacent, or games the test. Written justification matters as much as the verdict; your notes are what a reviewer and, downstream, a training pipeline actually consume.
micro1's screening is AI-led and runs as a live interview against your stated background. Expect it to pull specifics from your public contribution history and press on them: which repos, which PRs, what the review conversation looked like, why a maintainer rejected an earlier version. It also probes environment fluency — dependency pinning, flaky test triage, containerised reproduction — because tasks that don't reproduce reliably are worthless. The strongest signal is being able to talk through a non-trivial debugging session end to end without hand-waving the middle.
Practical logistics
- Remote, contractor basis, invoiced hourly or per accepted task depending on the project.
- Largely asynchronous; some projects add a weekly sync or calibration session where graders reconcile disagreements.
- Volume fluctuates with client demand — treat it as supplementary rather than a guaranteed weekly load.
- You supply your own machine; work often involves Docker, multiple language toolchains, and cloning sizeable repositories.
- Rates in the $50–150/hr band are observed across postings, not promised. Scarcer stacks, systems-level work and sustained low rework rates sit at the upper end.