What the work actually is

You receive a prompt — write an acceptance criteria set for a checkout flow change, draft a release note for a permissions update, summarise a trade-off decision for an executive audience — along with one or more AI-generated responses. You score each response on clarity, tone, instruction adherence, and usefulness to the stated audience, then write a rationale that explains the score in terms a second reviewer could apply to a similar case. The rationales are the deliverable; the numbers alone are close to worthless to the customer.

Most of the difficulty sits in a specific failure mode: fluent, confident product prose that doesn't actually answer the prompt. A spec that reads well but leaves the edge cases undefined. Acceptance criteria that restate the feature description instead of stating testable conditions. A release note that talks to engineers when the prompt named the audience as end users. You are also asked to flag factual inconsistency, scope creep beyond prompt constraints, and cases where a response is complete but unusable for the reader it was written for.

What the platform screens for

micro1's screen is AI-led and leans on follow-up questions. Expect to be pushed on the concrete artefacts you've authored — a real PRD, a real acceptance criteria set, a real decision memo — and on whether you can articulate why one piece of product writing serves its audience better than another. The listing is unusually explicit that delivery, programme, and scrum backgrounds are not a fit: the engagement wants people who have written and reviewed product specifications, not people who have run the process around them. Prior AI or RLHF experience is genuinely optional and the screen treats it that way.

Logistics

  • Fully remote contractor engagement, asynchronous with project managers and reviewers.
  • Task-based volume rather than fixed shifts; contributors typically commit a predictable weekly block.
  • Observed pay band for this listing is $90–140/hr, varying by project and assessed depth — stated as observed, not guaranteed.
  • Written English at native level is treated as a hard requirement because the output is prose critique.