What the work involves
You will be handed documents that a language model produced — or that were seeded to look like internal engineering material — and asked to rule on realism. Not whether the prose is clean, but whether a staff engineer at a real company would have written this, in this structure, with these omissions, at this point in a project's life. Typical artifacts include design docs and RFCs, incident postmortems and timelines, migration plans, API deprecation notices, on-call runbooks, PR descriptions and review threads, and architecture decision records.
The judgments that matter are the unglamorous ones. A postmortem that has no messy contributing factors and a single tidy root cause is a tell. A design doc with no alternatives-considered section, or one where the rejected alternatives are obvious strawmen, is a tell. A code review thread where every comment is resolved politely and nothing is left open is a tell. You are asked to name these specifically, in writing, so the lab can turn your judgment into training signal.
What the platform screens for
- Recency. Ethos cares that you currently write and review these documents, not that you did five years ago. Expect questions about the last design doc you reviewed and what you pushed back on.
- Depth under follow-up. The AI voice screen will take a general answer and ask you to make it concrete — which system, which failure, what the doc actually said.
- Written reasoning. Ratings without a written rationale are close to worthless here. Sample written work is often requested after the call.
- Calibration. Whether you can hold a consistent standard across dozens of documents rather than drifting stricter or looser as you fatigue.
Logistics
Fully remote, asynchronous, no fixed schedule. Volume arrives in batches and can be uneven — some weeks carry steady work, others little. Most reviewers treat it as 5–15 hours per week alongside a full-time role. The $150/hour figure is the rate observed on this listing; rates on evaluation platforms vary by batch, task type, and calibration performance, and nothing here is a guarantee of volume or earnings. The process includes an AI voice screen before any paid work.