The work

Mercor is building evaluation sets for agents operating inside a seeded FP&A environment — a multidimensional model with scenario trees, account hierarchies, CSV-imported actuals, note surfaces, and an accompanying document set (credit agreement, compliance certificates, borrowing base certificates, lender packages, board pre-reads). Your day is spent opening a completed agent rollout and deciding whether the answer holds: did it read the cube at the exact coordinate or at whatever grain the UI landed on, did it use the forecast of record or a later unlocked scenario, did it fix the figure by editing an input account when the situation called for overriding an assumption cell, did it state a correction in prose without ever writing it to the model.

Beyond pass/fail, you write the reasoning. That means naming the gating failure precisely — wrong vintage, wrong accounting basis, an EBITDA add-back the agreement caps or excludes, a conclusion transcribed from a document that already contained it — and then editing the grading rubric so the next reviewer catches it without you. You will also flag tasks that are too easy, solvable from the document pack alone, or answerable without touching the model at all, and propose harder seeded facts that force genuine retrieval.

What the screen is actually testing

  • Tool fluency, not tool awareness. The interview probes scenario inheritance, locking, parent/child behavior, and import unit-and-scale discipline with follow-ups. Describing the concept in general planning-software terms rather than in terms of what you did last week reads as secondhand.
  • Willingness to fail a plausible answer. A large share of the value here is marking wrong a number that looks reasonable, ties to something, and is stated at the wrong grain.
  • Covenant and measure-type depth. Frozen GAAP provisions, Test Periods, fixed charge coverage, total net leverage, Eligible Accounts and borrowing base redetermination; GAAP vs non-GAAP vs covenant measures; AICPA forecast-versus-projection distinctions; ASC 205-40, ASC 280, SAB 99.
  • Writing. Rubric language has to be unambiguous to a grader who did not see the model.

Logistics

Fully remote and asynchronous, contract, paid hourly. Observed range on this listing is $70–110/hr, positioned by depth of covenant and technical-accounting experience — stated as observed, not guaranteed. Most contributors commit 10–20 hours a week in blocks of their choosing; work arrives as batches of rollouts with turnaround windows rather than fixed shifts. Expect a calibration period where your judgments are compared against other reviewers before volume opens up. A CPA, CFA or MBA is welcome but not required.