What the work actually involves

Projects in this pool are built around the gap between an answer that sounds authoritative and an answer someone could act on without regret. On a typical task you might write a consumer question from your own field — a gear purchase, a trip itinerary, a wardrobe problem, a home-setup decision — along with the constraints a good answer has to respect (budget, season, location, skill level, who it's for), and then the reference answer the model's attempt gets graded against. Other tasks flip it: you read a model's recommendation and judge whether the product is still sold, whether the price is remotely current, whether the advice fits the setting described, whether the cultural or regional detail holds up. Most of your output is written reasoning, not scores — the rating matters less than the explanation of what the model got wrong and why it matters.

The failure mode you are hired to catch is confident wrongness: a discontinued model recommended at a 2019 price, a restaurant that closed, a hike that's snowed in that month, a styling rule that only applies in one market. Researchers can generate plausible text on their own. What they cannot generate is someone who knows the category well enough to notice the detail that would embarrass a real user.

What the screen is looking for

The AI interview runs about 20 minutes and probes whether your knowledge is lived or looked up. Expect follow-ups that push past the first answer: not just what you'd recommend, but what changes it, what you'd have to know about the person first, and where you'd hedge. Interviewers also test tolerance for ambiguity — the willingness to say an instruction is unclear and name the specific ambiguity, rather than guessing and delivering something confidently off-brief. Clear written explanation carries real weight here, because your reasoning is the deliverable.

Logistics and honest expectations

  • Fully remote and asynchronous; projects specify their own hours and deadlines.
  • Listings in this field on Mercor have posted at $50–75/hr, set per project by scope and depth. Observed range, not a guarantee.
  • This listing does not accept or reject anyone. There is no decision to wait for — you're added to the pool and contacted when a project matches your category, which may be within a week or several months.
  • The specific rate, hours, and client are named on the project listing you're invited to; hiring happens there.
  • Completing the domain expert interview and joining every network you qualify for improves match odds. One signup covers all of them.