What the work involves
You will spend most of your time reading model output and deciding whether it would survive contact with a real operation. A prompt might ask the model to build a two-week schedule for a 40-person site under a callout wave, draft a corrective-action plan for an underperforming district, or reconcile a departmental budget variance. Your job is to say whether the answer is operationally sound, financially literate, and feasible across the departments it touches — and then write down precisely why, in terms another manager would recognize.
A large share of tasks are head-to-head comparisons: two responses to the same scenario, and you pick the better one and justify the pick. The justification is the deliverable. "Response A is better" carries no signal; "Response A staffs to demand but ignores that the overtime it creates blows the labor line by roughly 6% — Response B is slower but stays inside budget and flags the tradeoff" is what the project is buying.
What the platform screens for
- Verifiable operational scope. Headcount, number of locations or departments, budget size, whether you owned the P&L or reported into someone who did. Vague titles get probed hard.
- Financial rigor under follow-up. Expect questions on labor cost as a percentage of revenue, variance analysis, or how you'd defend a staffing decision to a CFO. Screeners follow up on the second and third layer.
- Evaluation judgment. Can you separate an answer that is wrong from one that is merely written in a style you dislike? Can you name the specific failure rather than gesturing at "not practical"?
- Written English. Rubric feedback is prose, and unclear prose is unusable regardless of how good the underlying call was.
Logistics
Fully remote and asynchronous — you claim tasks from a queue and work when you want, though throughput expectations are real at 20+ hours per week. Start is immediate and the project window is short, roughly 4–6 weeks, so availability in the next several days matters more than usual. Payment structure is task-based initially: your first task must be approved, and completing it within the stated timeline is what qualifies you for hourly pay across the remainder. Pay in the $70–110/hr range is what has been observed on comparable Mercor operations projects, not a guarantee of your rate.