What the work actually involves
You receive a prompt and a set of supplied artifacts — operational notes, a fragmentary order, a spreadsheet of unit activity — and produce the document a competent action officer would produce from them. Sometimes that is a one-page SITREP; sometimes a weekly activity roll-up for an O-6; sometimes a decision brief with a recommendation and a coordination line. These are the "golden responses" a model is trained against, so they have to be right in the way a real staff product is right: correct format, correct level of detail for the audience, BLUF where BLUF belongs, no invented specificity.
The second half of the job is evaluation. You read model-generated reports and score them against rubrics — often rubrics you helped draft. That means articulating why a paragraph fails: it buried the decision, it used a term of art from the wrong branch, it padded a SITREP with narrative that an S-3 would strike, it produced a confident timeline the source artifacts never supported. Feedback is written and, at times, recorded verbally. You will also be asked to describe stylistic variance — how the same AAR looks coming out of a Marine infantry battalion versus an Air Force wing staff — because the project wants breadth, not one house style.
What the screen is looking for
- Real, verifiable time inside DoD documentation workflows — roughly three years or more — not adjacent familiarity from a contractor reading room.
- Named artifact fluency: you can talk about WARs, SITREPs, AARs, memoranda, decision briefs, and executive summaries with the specificity of someone who has been sent back to rewrite them.
- Judgment about quality, not just format. Can you say what separates an adequate executive summary from one that survives the chief of staff?
- Awareness of classification and OPSEC boundaries. Everything you write here is unclassified and synthetic; screens probe whether you understand that line instinctively.
Logistics
Contractor engagement, fully remote, US-based contributors only. Work is largely asynchronous with task batches and turnaround windows, plus occasional live calls with other SMEs or reviewers to reconcile rubric interpretations. Hours are flexible and typically part-time — contributors commonly report a handful of hours a week to roughly fifteen, scaling with project phase. No AI or machine-learning background is expected; the platform trains you on the tooling.