What the work actually is
This is evaluation authoring, not client work. A typical task starts with you inventing a problem you have genuinely solved on the job — a mix bus that collapses in mono, a design system token structure that breaks at three breakpoints, a checkout flow with an accessibility failure that only shows up under screen reader focus order. You write the scenario with enough constraints that a competent expert would arrive at a defensible answer, author that answer yourself, and then write a rubric granular enough that a second reviewer scoring the same model output would land within a point of you.
The grading half is where most of the difficulty lives. Frontier models produce answers that read fluently and cite the right vocabulary while being subtly wrong — a plausible-sounding compressor attack time, a WCAG criterion cited for the wrong reason, an information architecture rationale that assumes a user research finding nobody made. Your job is to catch that specifically and write feedback explaining what expert judgment the model failed to apply, not just mark it down.
What the screen looks for
AfterQuery's screening is AI-led and follow-up heavy. It is testing whether your expertise is current and hands-on rather than remembered: expect it to take a claim you make and push two or three layers into the specifics — signal chain, tooling versions, why you chose one approach over the obvious alternative. It also probes evaluation judgment separately from craft. Plenty of excellent practitioners cannot articulate why one answer is better than another in writing, and that gap disqualifies here. Prior peer review, design critique, code review, or teaching experience is the strongest positive signal after domain depth.
Logistics
- Fully remote and asynchronous; no standing meetings, work claimed from a queue
- Contributors commonly report 5–15 hours a week, sized around existing employment
- Pay observed in the $100–170/hr range, banded by track and demonstrated depth — not guaranteed, and calibrated during onboarding tasks
- Written output is the deliverable, so expect writing to occupy more of the hour than you might assume