What the work involves
You work inside two audit collections. In Multilingual Transcription Audit, you play a recording of human or agent speech, read the transcription an annotator submitted, and decide whether it faithfully captures what was said. You apply a fixed error-code taxonomy, issue a pass or fail, and write a short rationale in English. Critical errors are named first so a reader can see immediately why something failed. In Forced Alignment Audit, you review audio already segmented at the word level by an automated system, and for each segment confirm that the start and end timestamps land on the actual word boundaries and that the reference text matches the spoken form. Where the machine is wrong you adjust the boundary, split a segment, merge two, insert a missing one, or delete one containing no speech.
Both collections are budgeted at 5.00 audit hours per task. That figure sets the tone: this is not volume piecework. You are expected to tag every applicable error code rather than stopping at the first one, and to distinguish genuine acoustic ambiguity — overlapping speech, coarticulation, a swallowed final syllable — from an annotator who simply misheard or normalised the text incorrectly.
What the screen looks for
- Brazilian Portuguese as a native or near-native variety. Either raised where it is dominant, or fully fluent with five or more years of residence there. This is stated as a hard requirement.
- English reading and writing at working standard. The rulebook, the error codes, and your rationales are all in English.
- Consistency under a written standard. Screeners probe whether you follow the rubric when your instinct disagrees with it, and whether you can articulate the difference between an error and a judgement call.
- Ear for boundaries. Expect questions about where one word ends and the next begins in connected Brazilian speech, and about what counts as a real timestamp error versus an acceptable tolerance.
Prior transcription, subtitling, localization, linguistic annotation, or QA work is a strong signal. Phonetics or linguistics training and familiarity with waveform tools help but are not required.
Logistics
Fully remote and asynchronous, with work drawn from a queue rather than scheduled shifts. A short calibration phase precedes production work — expect your early judgments to be compared against gold standards and your rationale quality reviewed. Pay has been observed at $25/hr for this collection; rates on Mercor vary by collection, locale scarcity, and calibration outcome, and are not guaranteed. You will need reliable internet and good headphones; laptop speakers are not adequate for boundary work.