The work

Mercor is building gold-standard benchmark sets for economics and finance, and this quest puts you on one of two tracks. On the authoring track, you design original questions in your subdomain — Algorithmic Trading & Market Microstructure, Macroprudential Policy, Behavioral Finance & Experimental Economics, Urban Economics, or Tokenomics & DeFi — that test conceptual reasoning rather than definitional recall. Each item needs a self-contained problem statement, one defensible correct answer, nine plausible distractors that a strong solver could actually fall for, a Chain-of-Thought solution in markdown, a difficulty rating (Medium / Hard / Expert), and one to five references from peer-reviewed journals or university repositories.

On the verification track, you review items written by other experts. The judgment work is narrower but sharper: is the question solvable as written, is exactly one option correct, is any needed assumption missing, and does the stated difficulty match reality? You edit where you can, flag where you can't, and document your reasoning. Reviewers who simply approve everything are usually filtered out quickly.

What the screen looks for

The intake is AI-led and follow-up heavy. Expect it to verify your credential (PhD or doctoral candidate in economics, finance, or adjacent; strong master's-level candidates with real subdomain depth are considered), then push on a named area of expertise until the questions get genuinely technical. It also probes evaluation judgment — whether you can articulate why a distractor is pedagogically useful versus merely wrong, and whether you can spot ambiguity in a prompt you didn't write. Clear, compact written English matters here as much as domain knowledge; the deliverable is prose.

Logistics

  • Fully remote and asynchronous; no fixed meetings or shift windows.
  • Expected commitment of 10+ hours/week, self-scheduled.
  • Pay band observed at $77–98/hr; actual rate depends on track, subdomain, and calibration outcomes, and is not guaranteed.
  • Output is reviewed, and continued work typically depends on acceptance rates on your first batches.