What the work actually involves
You are supplying the judgment that a frontier model is graded against. In practice that means three recurring task shapes. First, authoring modeling problems drawn from systems you personally shipped — the data, the objective, the constraint, and a reference solution that a model's attempt gets scored on. Second, grading model-written training and evaluation code: does the loss actually encode the stated objective, does the eval split leak through a shared user ID or a time boundary, does the ablation support the causal claim drawn from it. Third, red-teaming plausible-looking proposals — writing down why an approach would converge nicely on the offline benchmark and still degrade under distribution shift, label noise, or a metric that rewards the wrong behaviour.
Most of the output is prose. You are not asked to build a model; you are asked to explain, in writing a reviewer can audit, why a piece of ML work is right or wrong and what the specific failure mode is.
What the platform screens for
- Shipped experience over pedigree. The listing says it plainly: where you trained matters less than whether models you built went in front of real users and you dealt with what came back. Screens probe for incidents, regressions, and rollbacks, not coursework.
- Depth under follow-up. The AI interview asks a general question, then drills. Vague answers about "improving accuracy" collapse on the second question; concrete numbers, baselines, and what you tried that failed hold up.
- Written reasoning. Because grading work is explanation work, clarity and calibration — saying what you are unsure of — count more than confident assertion.
- Willingness to flag ambiguity. Projects arrive with imperfect instructions; the platform wants people who raise the problem rather than guess silently.
Logistics
Fully remote and asynchronous. Apply with a resume, confirm your work location, then take an AI interview of roughly 20 minutes. Mercor verifies your background and adds you to the machine learning pool. When a project needs ML expertise, matching people are invited to a separate listing that names the rate, hours, and client — hiring happens there. This listing itself never returns a decision, and the wait between joining the pool and a match can run from a week to several months. Most project work is part-time and self-scheduled, often 10–20 hours a week alongside a main job, though scope varies per engagement. Rates are set per project by scope and depth; $70–120/hr reflects what comparable ML projects have posted, not a guarantee.