The work itself

You read a conversation — anywhere from one to seven turns, often just one — and score the assistant's response against a detailed rubric. Did it answer the question that was asked, or an adjacent easier one? Could a family member with no medical background act on it? Did the tone leave the user calmer and holding something concrete, or anxious and no further along? Budget roughly 2–13 minutes per item depending on length. Every rating below the top option needs a written justification that quotes or points to the specific failure and says what a better response would have said; a one-line verdict like "too vague" will be rejected.

The centre of gravity is sycophancy. That means catching the response that folds the moment a user pushes back, that retreats into "you know your body best" instead of naming a risk, or that reframes a dangerous decision as empowerment. These answers read as warm, careful and responsible, which is exactly why automated evaluators score them as helpful and why the project needs human readers. You are not being asked to judge whether the medical content is correct — clinicians run that track on the same conversations. Supplying clinical opinion you don't have is a failure mode here, not a bonus.

What the screen looks for

Mercor's screening is AI-led and heavily weighted toward writing samples and worked judgment. Expect to be shown a conversation and asked to rate it and justify the rating in your own words, under time pressure. Screeners are checking whether you can distinguish a genuinely unhelpful answer from one that merely sounds hedged, whether you can hold a rubric steady across many items instead of drifting toward your own instincts, and whether your prose is specific enough to be actionable. Prior annotation, RLHF, content-review, teaching, editing, patient-advocacy or QA experience is preferred but not required.

Logistics and the AI rule

  • Fully remote and asynchronous; you work through the platform's task queue.
  • Volume arrives in batches, some of them time-boxed — the listing asks for the ability to concentrate hours when a batch drops rather than a fixed weekly schedule.
  • Observed pay on this platform spans $20–160/hr and varies by track, item type and assessed skill level; generalist annotation sits at the lower end of that range. Rates are as observed, not guaranteed.
  • Using an AI tool to draft your comments is prohibited and ends your work on the project. It is checked for. The dataset exists to measure what AI systems miss, so AI-written feedback destroys its purpose.