The work

You are the quality gate between a generated product spec and the benchmark that depends on it. Each task hands you an AI-written specification for a web application — a task tracker, a booking flow, an internal dashboard — and asks whether it could actually be built and then verified. You check whether requirements are complete and unambiguous, whether acceptance criteria are testable rather than aspirational, and whether the described UI and data model hold together. Many items arrive as sequences: a base product plus follow-on feature requests. Your job includes confirming that each increment builds coherently on the prior state rather than silently contradicting it or assuming features that were never specified.

A second thread of the work is alignment with browser-based test scenarios. Specifications are paired with automated tests that a model's output will be judged against, and mismatches between the two quietly corrupt the benchmark. You flag cases where a test asserts behavior the spec never mentions, or where the spec promises something no test can observe. Output is structured written feedback against project rubrics — specific, cited to the line or requirement, and actionable enough that someone can repair the spec without a follow-up conversation.

What the screen looks for

  • Verifiable product background: 3+ years shipping or specifying software, with PRDs, user stories, or acceptance criteria you actually authored or reviewed.
  • Fluency in modern web application patterns — auth, state, CRUD, permissions, empty and error states — and the ability to spot what a spec left out.
  • Evaluation judgment under follow-up: can you defend a call that a requirement is ambiguous, and distinguish genuine ambiguity from reasonable implementer discretion?
  • Written precision. Rubric-based review rewards concise, evidence-anchored notes over general impressions.

Logistics

Fully remote and asynchronous, with no fixed shifts. Reviewers typically commit 10–20 hours per week and are paid hourly at observed rates of $30–60/hr, with placement in the band reflecting depth of product experience and calibration performance. Volume on benchmarking initiatives is bursty — expect onboarding calibration rounds, then batches that arrive with short turnaround windows. Rates and availability are as observed on the platform and are not guaranteed.