What the work actually looks like

You will spend most sessions reading model output and deciding whether it would survive contact with a real project. A task might hand you a scenario — a 14-week integration project slipping two weeks, a vendor invoice that exceeds the approved PO, a resource conflict between two workstreams — and two AI-generated responses. You judge which one plans better, and then write out why: the WBS is missing a dependency, the mitigation is a restatement of the risk, the status report buries a red item under three paragraphs of green, the budget reforecast never touches the contingency line. Other tasks are single-response rubric grading of artifacts: charters, RAID logs, schedule narratives, escalation emails to a steering committee.

The hard part is not spotting errors. It is being specific about them in writing, at speed, in a way another PM would agree with. Models produce project plans that read fluently and fail on feasibility — 40-hour weeks assumed for a shared resource, procurement lead times ignored, a critical path that is not actually the critical path. Your value is naming that concretely rather than marking it "unrealistic."

What the screen is looking for

  • Verifiable delivery history. Four-plus years actually running projects — schedules, budgets, staffing, procurement — not adjacent coordination. Expect follow-ups on project size, budget authority, and how you handled a specific slip.
  • Judgment that survives probing. The AI interviewer will push on your reasoning: why is that plan worse, what would you have changed, what does the rubric say versus what your instinct says.
  • Written English under time pressure. Rationales are the deliverable. Vague ones are rejected in QA.
  • Honest availability. 20+ hours weekly for the duration, not a hopeful estimate.

Logistics

Fully remote and asynchronous — you claim tasks from a queue and work whenever suits you, though throughput expectations are weekly rather than daily. Duration is roughly 4–6 weeks with an immediate start. Payment structure matters here: the first task is compensated on approval, and clearing it within the stated window is what unlocks hourly pay for the rest of the project. PMP or PRINCE2 is not required; delivery experience is.