The work

This is a rubric-first evaluation project, not a chat-rating queue. A typical task cycle starts with you inventing a personal-life request that is genuinely hard — finding a restaurant that works for a gluten-free guest, a $40/head budget, and a 7pm Friday reservation in a specific neighbourhood; reconciling six weeks of wearable sleep and resting-heart-rate data into something actionable; rebuilding a resume and outreach sequence for a lateral career move. You then do the task yourself, screen recording as you go, so there is a ground-truth record of what the research, the dead ends, and the real answer actually looked like. Finally you write a rubric that another evaluator could apply to model attempts: what counts as complete, what counts as plausible-but-fabricated, where an assistant overreaches into medical advice it shouldn't give.

What the screen looks for

  • Documented rubric hours. The listing is explicit about 100+ hours of prior rubric work spanning design, evaluation, and quality assessment. Expect to be asked which projects, what the rubric structure was, and how disagreements between graders were resolved.
  • Real LLM power usage. Not "I use ChatGPT sometimes." Multi-step delegation, agentic tools, comparisons across ChatGPT, Claude, Gemini, Perplexity, and coding agents like Cursor or Codex — and opinions about where each one fails.
  • Depth in at least one named domain. Food, personal health data, productivity and life admin, career advice, or learning plans. Breadth without depth reads as generic.
  • Writing that separates judgement from assertion. You'll be asked to explain why an output is bad in terms a model trainer can act on.

Logistics

Fully remote and asynchronous, 15–40 hours per week, with a 24-hour turnaround expectation on assigned tasks. A desktop or laptop is required — Chromebooks are not supported, largely because of the screen-recording workflow. The project is early-stage, so expect a gap between acceptance and first tasking; the $30 paid work trial precedes onboarding and weighs heavily in selection. Pay is listed at $50–200/hr as observed on the platform and typically varies by domain and demonstrated rubric experience.