The work
You listen to AI-narrated audiobooks the way an actual audiobook fan does — for pleasure, for pacing, for whether the narrator carries the scene — but with an analytical ear and a tagging tool open. When something breaks the listen, you stop, mark the exact span, choose a category, and write a short note on what went wrong. The error taxonomy covers text-fidelity failures (skipped words, inserted words, mispronounced proper nouns, numbers read as digits instead of quantities, abbreviations wrongly expanded), prosody failures (question intonation on a declarative, wrong emphasis, sentence-final falls mid-clause, unnatural pauses at line breaks), and signal-level problems (clicks, clipping, sudden timbre shifts, truncated phonemes). You also give a holistic listener-satisfaction read: where the narration held, where it broke the spell, and whether you'd have kept listening.
The volume is real. Four hours a day of close listening is fatiguing in a way that surprises people, and the value of the work depends entirely on your span-marking staying as precise in hour four as it was in hour one.
What the screen looks for
Mercor's screening is AI-led and conversational, and it pushes on specifics. Expect to be asked what you've listened to recently and what you thought of the narration — vague enthusiasm reads as someone who doesn't actually listen. Expect probes on whether you can name a concrete failure and describe it in terms a model team could act on: not "it sounded robotic" but "it read 'Dr.' as 'drive' in a medical thriller" or "it put rising intonation on every list item including the last." Prior annotation, transcription, proofreading, or linguistics work is a genuine plus because it shows you've held a taxonomy in your head under volume. Availability is checked plainly: about 20 hours a week, sustained.
Logistics
- Remote, location-flexible, with a preference for candidates living in a natively English-speaking region.
- Roughly 20 hours per week, about 4 hours per day, part-time and asynchronous within reasonable turnaround expectations.
- Observed pay band $20–25/hr; rates are set by the platform and not guaranteed.
- You'll need decent headphones and a quiet place to listen — artifact detection is unreliable on laptop speakers.