What the work actually looks like
You take a queue of samples — a model's answer to a technical question, a document extract, two competing outputs to rank — and apply your professional judgment against a written rubric. Most tasks want more than a label: a short written justification, a citation to a standard or source, and a note on what specifically is wrong when something is wrong. Batches vary. One project may be short binary judgments at volume; another may be a handful of long-form critiques where a single item takes forty minutes. Guidelines get revised mid-project, and the expectation is that you re-read them and adjust rather than annotate from memory.
What the screen is looking for
AfterQuery's intake is domain-agnostic in name but not in practice — the value is in the specificity of your expertise, so the screen pushes on what you actually did in your field, not what you know about. Expect follow-ups that go one layer deeper than your first answer: name the standard, name the failure mode, describe the judgment call you got wrong. Written clarity is weighted heavily, because your justifications are the deliverable that downstream teams read. Consistency also matters: the platform checks whether you apply a rubric the same way on item 200 as on item 3, and whether you flag ambiguity instead of quietly guessing.
Logistics
- Fully remote, asynchronous, project-based. No fixed shifts.
- Work arrives in batches; availability tends to be lumpy, with quiet weeks between projects.
- Typical commitments run 5–20 hours a week when a project is live, self-scheduled.
- Pay observed in the $25–60/hr range, set per project by domain and task difficulty. Rates are not guaranteed and vary between batches.
- Qualification tasks are common before a paid project opens; they are usually short and unpaid or paid at a reduced rate.