The work

You listen to text-to-speech audiobooks the way a paying listener would, then stop and account for every moment the illusion cracks. A typical session means several hours of audio against a reference text: marking the exact span where the narrator says el for él, where a date reads as a raw digit string, where an abbreviation like Dr. expands as the wrong word, where a clause carries declarative intonation through a question, where a breath or splice artifact interrupts a paragraph. Each finding gets a category and a start/end timestamp, plus a holistic judgment at the end: would you have kept listening, and where exactly did you stop wanting to?

The hardest part is not detection, it is consistency. Hour one and hour four have to apply the same threshold for "unnatural phrasing," and your severity calls need to hold across chapters and across days. Reviewers who drift — flagging everything early and nothing late, or vice versa — produce data the modelling team cannot use.

What the screen looks for

Mercor's screening is AI-led and probes for verifiable specifics rather than enthusiasm. Expect it to test whether your Spanish is genuinely North American in intuition (regional lexis, seseo, voseo absence, how you'd treat a Peninsular pronunciation in a US-targeted title), whether you actually listen to audiobooks and can name narrators and titles, and whether you can reason about severity — is a single misstressed proper noun worse than persistently flat sentence-final intonation? Prior annotation, transcription, proofreading, or linguistics work is expected in some form; formal QA credentials are not.

Logistics

  • Remote, location-flexible, though candidates living in a Spanish-speaking region are preferred.
  • Around 20 hours per week, roughly 4 hours per day, largely async against throughput targets.
  • Working English sufficient to write clear issue reports and talk to the team — full bilingualism is not required.
  • Observed band $15–20/hr; rates on Mercor vary by project and are not guaranteed.
  • You will need decent headphones and a quiet place to listen; laptop speakers will hide the artifacts you are being paid to catch.