What the work involves
You spend most of your time reading AI-generated technical content and deciding whether it holds up. That means grading model answers about test strategy, defect triage, or automation code against a written rubric, and flagging the answers that read fluently but are wrong — a test case that misses the boundary condition, a bug report with reproduction steps that don't actually reproduce, a severity rating that doesn't match the described impact. Alongside evaluation, you author reference material: comprehensive test cases covering functional, regression, negative, and edge-case paths, and defect write-ups with reproduction steps precise enough that a developer never has to ask a follow-up question.
The written feedback is the deliverable as much as the score. A rating with no justification is close to worthless to the customer; a rating with a specific citation of what the model got wrong and why it matters is the product. Expect to work inside project guidelines that change, and to be asked for input when the rubric doesn't cleanly resolve a case.
What the platform screens for
micro1's screen is AI-led and follow-up heavy. Two gates dominate. First, real QA depth — expect probes on test design technique, regression scope decisions, severity versus priority, and the automation and test management tooling you name (Selenium, Playwright, Cypress, Appium, Postman, Jira, TestRail, Zephyr, BrowserStack). Second, and stated bluntly in the listing: prior paid human-data experience supporting AI training. Annotation, labeling, RLHF preference ranking, model response evaluation, or rubric-based grading. Software QA experience without a human-in-the-loop AI component is explicitly disqualifying, so be ready to name the platform or client, the task type, and roughly how long you did it.
Written English at B2 or above is a requirement because the annotations ship to the customer as-is. No degree is required; demonstrable testing work takes precedence.
Logistics and pay
- Fully remote contractor engagement, described as a high-volume customer project
- Largely asynchronous task work; hours are typically flexible, with throughput and quality expectations rather than fixed shifts
- Reliable internet and readiness to start promptly are stated requirements
- Observed pay band is $90–175/hr; actual rate depends on the project, seniority band, and screen outcome and is not guaranteed