What the work involves

Day to day, you are handed technical material — derivations, datasets, plots, code, spec sheets, papers — and asked to judge whether a model's handling of it is correct. That can mean writing a hard problem in your subfield with a fully worked solution, grading two candidate answers against each other, tracing exactly which step of a derivation went wrong, or verifying that a figure the model described actually says what it claims. Tasks arrive through Mercor's workflow tooling rather than as open-ended research, and the unit of value is a defensible verdict with a written rationale, not a vibe.

Because the project centres on technical and scientific data, a lot of items touch numbers, units, and visualisation: reading a chart correctly, noticing an axis or unit error, checking that a computed value follows from the inputs given. Experience wrangling and plotting data is listed as preferred for exactly this reason. The listing's phrasing — "the judgment to verify a result against a source rather than eyeball it" — is the actual screening criterion, and it shows up in how reviewers read your written justifications.

What the platform screens for

  • Graduate-level training or substantial industry depth in engineering, physics, mathematics, chemistry, or software engineering / data science — and a specific subfield you can be pushed on, not a survey-level tour.
  • Whether you can state a claim, then say where it comes from: a standard result, a textbook relation, a computation you ran.
  • Written English clear enough that another expert can follow your reasoning without asking you what you meant.
  • Calibration: knowing when a model answer is wrong versus merely unconventional, and saying "I'd need to check" instead of guessing.

Mercor's first pass is typically an AI-led video interview that asks about your background and then follows up on the specifics you volunteer. Concrete beats broad here — one problem you solved in depth, with the method named, outperforms a list of fields you've touched.

Logistics

Fully remote and asynchronous, with no fixed shifts. Minimum five hours per week; many contributors do more when batches are open, and work can be uneven since flow depends on the lab's task queue. Applications are reviewed on a rolling basis, and pay within the $20–80/hr band is observed to track domain, credential level, and task difficulty rather than being a single flat rate.