What the work involves

You work across two Amazon Sonic audit collections, both delivered through Mercor. In the Multilingual Transcription Audit, you listen to a recording of human or agent speech in Hindi, read the transcription an annotator submitted, and decide whether it faithfully captures what was said. You tag every applicable error code — not just the first one you notice — assign a pass or fail verdict, and write a short rationale in English that leads with the critical error so a reviewer immediately understands the failure. In the Forced Alignment Audit, you review audio that a machine has already segmented at the word level. For each segment you check that the start and end timestamps land where the word actually begins and ends and that the reference text matches the spoken form, then fix what is wrong: move a boundary, split a segment, merge two, insert a missing one, or delete one that contains no word.

This is not production work. You are not transcribing from scratch and you are not being paid by the line. Tasks run at 5.00 audit hours, which signals the expectation: repeated listening, waveform-level attention, and a written reason for each judgment rather than throughput.

What the platform screens for

  • Language variety, specifically. Native or near-native Hindi as spoken in India — either raised in a region where that variety dominates, or fully fluent with five-plus years of residence there. This is stated as a hard requirement.
  • English as a working language for the rubric. Conventions, error codes, and every rationale you write are in English. Expect to demonstrate that you can read a dense rulebook and write feedback in it, not just speak conversationally.
  • Rubric discipline over instinct. Screens probe whether you can apply a written standard the same way on judgment 400 as on judgment 4 — including when you privately disagree with the standard.
  • Acoustic judgment. Can you separate genuine ambiguity in continuous speech (assimilation, elision, fast speech word boundaries) from a clear annotator mistake? That distinction decides pass versus fail.

Prior transcription, subtitling, localization, linguistic annotation, or QA work is called out as a strong signal. Phonetics or linguistics training and comfort with waveform tools help but are not required.

Logistics

Fully remote and asynchronous — you take tasks from the queue and submit on your own schedule, with no fixed shifts. A short calibration phase comes before production work, and passing it is how you get released to paid volume. Pay is stated as observed at $18/hr against audit hours, not a guarantee; queue availability on Sonic collections fluctuates, so treat this as flexible contract work rather than a fixed weekly income. Headphones and a quiet listening environment are practical necessities.