What the work actually is
You produce and judge the artifacts that come out of a real incident lifecycle, only the first draft is written by a model. A typical task gives you a scenario — a partial outage, a cascading dependency failure, a botched deploy — and asks you to write the blameless postmortem, the runbook, the severity write-up, or the reliability review deck an SRE org would actually accept. Other tasks flip it: the model has produced the artifact and you grade it, mark the specific spans where it goes wrong, and write a correction that explains why a real incident commander would reject it.
The recurring workflows are blameless postmortems and root-cause analyses, on-call runbooks, severity and escalation write-ups, SLO and error-budget reports, remediation action-item trackers, and reliability review decks. Format matters as much as content — spreadsheets need working formulas and sane column design, decks need a slide structure an exec would sit through, documents need timelines with timestamps that reconcile. Reviewers who write beautiful prose in a broken spreadsheet don't last.
What the screen looks for
Ethos runs an AI voice screen. It probes for incidents you personally ran or reviewed, not incidents you read about: what page woke you, what the detection gap was, what you wrote in the timeline, which action items actually shipped. Expect follow-ups that press on specifics — error budget math, the difference between a contributing factor and a root cause, why an action item is unassignable. The other half of the screen is evaluation judgment: whether you can articulate a rubric for a bad postmortem rather than just saying it feels wrong, and whether you can distinguish a model output that is confidently plausible from one that is correct.
Logistics
- Fully remote and asynchronous; tasks are claimed from a queue on your own schedule
- Flexible commitment, typically 5–20 hours per week, more if you want it
- Observed rate for this role is around $80/hour — banding varies by task type and review tier and is not guaranteed
- Work is done in the lab's tooling; expect a calibration period with graded sample tasks before volume opens up