What the work involves
You receive a procurement task prompt — draft an RFQ for a facilities-services category, build a weighted vendor-scoring matrix for three bidders, flag risk in a supplier SLA, or turn twelve months of spend data into a savings-pipeline deck — and you either produce the gold-standard artifact yourself or evaluate what the model produced against it. Most tasks land as a real deliverable: a formatted Word document, a working spreadsheet with live formulas, or a slide deck a category manager could take into a stage gate.
Evaluation work is where the judgment shows. A model-drafted RFP often reads fluently and still fails: scoring weights that don't sum, evaluation criteria that can't actually be scored from what the RFP asked bidders to submit, an SLA review that misses the absence of a service-credit mechanism, a spend cube that double-counts tail spend. You write the rationale explaining what broke and why it matters commercially — that written rationale is the training signal, not the score.
What the platform screens for
- Direct ownership, not adjacency: RFPs you ran, contracts you negotiated or redlined, categories you owned — with the specifics attached.
- Spreadsheet and deck craft at a level that survives inspection. Weighted scoring models, spend segmentation, savings-tracking logic.
- The ability to say precisely why a plausible-looking procurement document is wrong, in writing, without hedging.
- Ethos includes an AI voice screen; expect follow-ups that push past your first answer into mechanics.
Logistics
Fully remote and asynchronous, 5–20 hours per week with room to take more. Work is drawn from a queue rather than scheduled, so consistency across a week matters more than availability at particular hours. Pay is observed at $70/hour for this cohort and is not a guaranteed rate across all task types or tenures.