What the work actually involves

You write benchmark tasks that a frontier model should plausibly fail. That means starting from a real research problem in your subfield — a training setup, an evaluation protocol, a systems bottleneck, a derivation — and shaping it into a self-contained task with an unambiguous statement, a reference solution you could defend to a reviewer, and a rubric that a second expert could apply and land on the same score. Most of the time goes into the rubric and the disambiguation, not the first draft. You will also review and validate other contributors' tasks, which means flagging underspecified prompts, solutions that are right for the wrong reason, and rubrics that reward surface form over correctness.

What the platform screens for

  • Verifiable first authorship. Expect to name the paper, the venue or preprint, and what specifically was yours. Vague gestures at "co-authored work" stall the screen.
  • Hands-on code, not supervision of code. Screeners probe whether you personally wrote training loops, evaluation harnesses, or systems code. Annotation and labelling experience does not count here.
  • Depth over breadth. A candidate who can go three follow-ups deep on one area (RL, diffusion, inference optimization, mechanistic interpretability, whatever it is) outperforms one who names eight.
  • Rubric judgment. Can you say what distinguishes a 4 from a 5 on a task you wrote, in terms someone else could apply?

Logistics and pay

Fully remote and asynchronous, with no standing meetings. Contributors report anywhere from 5 to 40 hours a week, set by them; consistent weekly output matters more than volume. Work is submitted through the platform's task interface and goes through expert review before acceptance, so early tasks often come back with revision requests.

On pay: the listing text quotes $70–$80/hr, while the observed band for this role on Sidequest has been $140–$150/hr. Rates on AfterQuery vary by specialization, review track record, and task difficulty tier, and senior benchmark-authoring work in scarce subfields has been observed at the higher end. Treat neither figure as guaranteed — confirm your rate in writing before you start.