What the work involves

This is a hybrid writing and evaluation seat. Roughly half the day is production: authoring prompts designed to stress a specific capability (tone control, persona consistency, structural constraints, refusal edge cases) and writing gold-standard responses that a model should aspire to. The other half is assessment: reading model output against a rubric and scoring it on fluency, coherence, factuality, creativity, and bias, then writing the justification that makes your score usable by a researcher. Expect batches with defined turnaround, periodic rubric revisions mid-project, and calibration rounds where your scores are compared against other annotators'.

The feedback you write matters more than the score itself. "Response 2 is better" is worthless; "Response 2 maintains the second-person register the prompt established, while Response 1 drifts into third person at paragraph three and introduces a factual claim about the 1974 statute that the source does not support" is the deliverable.

What the platform screens for

  • Verifiable writing history. Three or more years in English content writing, copywriting, or editorial work — named outlets, clients, or in-house roles, and what you actually produced.
  • Range across register. Whether you can move between technical explainer, brand voice, and narrative prose without collapsing everything into one house style.
  • Concrete LLM experience. Not "I use ChatGPT daily" but specific failure modes you have observed and can describe: sycophancy, hedging, invented citations, instruction drift over long context.
  • Rubric discipline. Whether you can hold a stated criterion even when your taste disagrees with it, and flag the conflict rather than silently rescore.
  • Real availability. The listing asks for 40 hours/week across 3–6 months. Screens probe this directly and inconsistency here is a common rejection reason.

Logistics

Fully remote and largely asynchronous, with adaptable hours — but the 40-hour commitment is genuine, not nominal, and this is not designed to sit alongside another full-time role. Engagements run 3–6 months on contract. AfterQuery is YC-backed and places contributors on projects for a large AI lab partner; work is batched through their platform. UK and US candidates are the stated target. Bilingual capability in Chinese, Spanish, Portuguese, Arabic, French, or German is preferred, not required.