What the work involves

You are not writing code. You are judging artifacts an AI produced for an engineering audience — an architecture deck, a sprint capacity spreadsheet, a postmortem write-up, an API migration plan, a technical spec — and deciding whether a working engineer could actually use them. A typical task gives you a prompt, one or more model outputs, and a rubric. You grade against the rubric, then write the justification: what is wrong, where, why it matters, and what a correct version looks like. Errors cluster in three bands — factual (a benchmark figure invented, a concurrency model described incorrectly, a Terraform snippet that would not apply), structural (the deck answers a different question than the prompt asked), and presentation (broken slide hierarchy, a spreadsheet with hardcoded values where formulas belong).

What the platform screens for

Mercor's screen is AI-led and leans on verifiable specifics. Expect it to test whether your five-plus years are real by asking about systems you owned, tradeoffs you made, and things that went wrong. It also probes tooling literacy directly — Excel formula competence, PowerPoint structure, Google Workspace collaboration — because a surprising share of the errors you will flag are document-craft errors, not engineering errors. Finally it checks evaluation judgment: can you separate a stylistic preference from a defect, can you rank two flawed outputs consistently, can you write feedback another reviewer would reach the same conclusion from.

Logistics

  • Fully remote, hourly, contractor engagement.
  • Asynchronous — tasks are pulled from a queue, not scheduled meetings. Some projects impose turnaround windows measured in hours.
  • Volume fluctuates by project. Many reviewers treat this as 10–20 hours a week alongside a primary role; commit to what you can reliably deliver.
  • Pay band reflects rates observed on this platform for this category, not a guaranteed offer. Rates vary by project and calibration performance.