What the work involves
You receive an artifact an AI model produced in response to a data science brief — a stakeholder deck summarizing an A/B test, an Excel workbook with a cohort retention model, a written memo recommending a feature set for a churn classifier — and you judge whether it would survive contact with a real team. That means checking the statistics (is the significance claim defensible at that sample size? is the baseline right?), the mechanics (do the formulas trace, are the pivot ranges correct, does the chart axis start where it should?), and the framing (would a VP of Product understand the recommendation, or does slide 4 bury the result in a 12-series line chart?). You then write structured feedback explaining what is wrong, how badly, and what a correct version looks like.
Most tasks come with a rubric covering dimensions like factual accuracy, formatting quality, instruction-following, and usability. The hard part is not spotting obvious errors — it is calibration: distinguishing a cosmetic flaw from a defect that would mislead a decision-maker, and scoring consistently across dozens of artifacts so your ratings mean something in aggregate.
What the platform screens for
Mercor's screen is AI-led and leans on verifiable specifics. Expect questions about where you worked, what you actually built, and which tools you used day to day, with follow-ups that go a layer deeper than your first answer. Claims of "5+ years at a top firm" get tested by asking what your modeling stack looked like and how you presented results to non-technical stakeholders. Office and Workspace proficiency is treated as a real requirement, not a checkbox — you will be evaluating spreadsheets and decks, so you need to know what a well-built workbook looks like.
Logistics
- Fully remote, asynchronous, hourly.
- Work volume varies by project; contributors commonly take 10–20 hours per week, sometimes more during active batches.
- No fixed schedule, but tasks often carry turnaround windows measured in days.
- Pay in the $100–150/hr range has been observed for this category; actual rates depend on project, seniority, and screening outcome, and are not guaranteed.