What the work actually is
You pick a track you already work in — supply chain and ERP operations, FP&A and corporate finance, or regulatory compliance and audit — and build evaluation material inside it. A typical unit of work looks like this: construct a scenario with enough real-world texture that it can't be answered from a textbook (a demand plan with a distorted history, a three-statement model with an embedded covenant question, a control gap surfaced mid-audit), write the reference solution at the level a senior practitioner would defend it, then write a rubric that tells another reviewer how to distinguish a correct answer from a merely fluent one. Afterwards you grade model outputs against that rubric and write structured feedback explaining where the reasoning broke.
The hard part is rarely the domain answer. It's articulating the judgment you normally apply silently — why one allocation policy is defensible and another isn't, which assumption a model quietly swapped, where a plausible-sounding answer would fail in an actual close, audit, or S&OP cycle.
What the screen looks for
- Currency of practice. Recent hands-on work, not a decade-old role described in general terms. Expect follow-ups on specifics: systems used, cadence, what you personally owned.
- Writing under scrutiny. Rubrics and feedback are the deliverable. Vague prose fails here even when the domain reasoning is sound.
- Evaluation instinct. Prior peer review, audit review, model review, or grading experience is a strong signal — the ability to judge someone else's work is different from doing the work.
- Track depth over breadth. One deep track beats three shallow ones. Claiming supply chain, finance, and compliance equally invites probing on all three.
Logistics
Fully remote and asynchronous, no fixed hours or standups. Most contributors treat this as evenings-and-weekends supplementary work; a realistic commitment is roughly 5–15 hours per week, and the screen will ask you to state a number and defend it. Work is typically batched, so availability tends to matter in blocks rather than daily. Expect a paid or unpaid calibration task before steady volume begins — your first few rubrics get reviewed closely.