What the work involves

You'll spend most of your time reading AI output that looks plausible and deciding whether it would survive a design review. Typical tasks include scoring a model's response to a gearbox sizing problem, comparing two AI-written failure analyses head-to-head, checking whether a generated GD&T callout or tolerance stack-up is internally consistent, and flagging where a model has invented a material property, misapplied a safety factor, or cited a standard clause that doesn't say what it claims. Some tasks ask you to interpret a drawing or specification and rewrite the model's answer as a reference response. Written justification matters as much as the score — a rating with no defensible reasoning behind it is usually rejected.

What the platform screens for

Mercor's screen is AI-led and follow-up heavy. It will ask you to describe specific projects — what you designed, what tools you used, what the failure mode was, how it was resolved — and then press on the details. Vague ownership claims ("I led the mechanical side") get probed until they produce specifics or collapse. Expect at least one live judgment exercise: a short engineering answer to critique, where the graders care about whether you catch the substantive error rather than the formatting. Written English is assessed directly from your typed responses, since the deliverable is written critique.

Logistics

  • Fully remote and largely asynchronous; you pick up tasks from a queue rather than attending meetings.
  • 20+ hours per week is a real threshold, not a preference — throughput is tracked.
  • Duration is roughly 4–6 weeks, with an immediate start.
  • Payment is task-based initially: you're paid on approval of your first task, and completing it within the stated window is what unlocks the hourly rate for the rest of the project.
  • Pay in the $70–110/hr band has been observed for this role; actual rate depends on assessed depth and project tier and is not guaranteed.