What the work involves

You are reviewing other people's work, not producing your own. Two collections run in parallel. In Multilingual Transcription Audit, you play a recording of human or agent Mandarin speech, read the transcription an annotator submitted, and decide whether it faithfully represents what was said — applying a fixed error-code set, assigning pass or fail, and writing a short rationale. In Forced Alignment Audit, you inspect audio already segmented word-by-word by an automated system and verify that each start and end timestamp lands where the word actually begins and ends, and that the reference text matches the spoken form. When the machine is wrong you fix it: move a boundary, split a segment, merge two, insert a missing one, or delete an empty one.

Tasks are budgeted at 5.00 audit hours each, which tells you the expected pace. This is careful listening work — replaying a disputed 200 ms boundary, deciding whether a neutral-tone syllable was elided or the annotator simply dropped it, distinguishing genuine acoustic ambiguity from a plain mistake. Where a task fails, you tag every applicable error code rather than stopping at the first, and you surface critical errors first so the reason for failure is immediately legible to whoever reads your note.

What the platform screens for

  • Native or near-native Mandarin as spoken in mainland China — born and raised in a region where that variety dominates, or fully fluent with five-plus years of residence there. This is a hard gate.
  • English reading and writing at working standard. The rulebook, the error codes, and your rationales are all in English.
  • Evidence you can hold to a written rubric across hundreds of judgments instead of drifting toward instinct. Transcription, subtitling, localization, linguistic annotation, or QA history is a strong signal.
  • Ear discipline — segmentation in connected speech, sandhi, regional colouring, background noise.

Phonetics or linguistics training, prior AI annotation or model evaluation work, and familiarity with waveform editors (Praat, Audacity, ELAN) help but are not required.

Logistics

Remote and asynchronous, paid hourly at an observed rate of $22/hr — rates on Mercor vary by collection and are not guaranteed. Expect a short calibration phase before production work unlocks; calibration is where most candidates are filtered, not the interview. You will need quiet working conditions and decent closed-back headphones — laptop speakers are not sufficient for boundary work. Hours are self-scheduled, but volume is only awarded consistently to auditors whose agreement scores hold up.