The work
You sit on the review side of a speech pipeline. In the Multilingual Transcription Audit collection you play a recording of human or agent speech, read the transcript an annotator submitted, and decide whether it faithfully represents what was said in Japanese. You tag every applicable error code — not just the first one you hit — assign pass or fail, and write a short rationale in English that makes the decision legible to someone who does not speak Japanese. Critical errors go first in the writeup.
In the Forced Alignment Audit collection you work at the word level. An automated aligner has already cut the audio into segments with start and end timestamps; you confirm each boundary lands where the word actually begins and ends, and that the reference transcription matches the spoken form rather than the written one. When the machine is wrong you fix it: nudge a boundary, split an over-merged segment, merge two fragments, insert a missing word, or delete a segment containing no speech. Japanese makes this non-trivial — mora timing, devoiced vowels, sokuon, particle elision in fast speech, and the absence of orthographic word spacing all sit right on top of the boundary decisions.
What the screen looks for
The platform screen is AI-led and checks three things. First, verifiable language provenance: native or near-native Japanese as spoken in Japan, either by upbringing or by five-plus years of residence. Second, English strong enough to read a detailed rulebook and write rationales in it — expect to produce written English during screening, not just claim it. Third, rubric discipline: whether you can apply a fixed error taxonomy consistently across hundreds of judgments instead of falling back on "it sounded wrong to me." Prior transcription, subtitling, localization, linguistic annotation, or QA work is a strong signal; phonetics training and waveform-tool experience help but are not required.
Logistics
- Fully remote and asynchronous; you pick up tasks from a queue.
- Tasks run at 5.00 audit hours each — this is deliberate. It is careful listening work, not volume piecework.
- A short calibration phase precedes production work, where your verdicts are compared against gold standards.
- Quiet environment and decent headphones are effectively required; you will be resolving boundaries at the tens-of-milliseconds level.
- $38/hr is the observed rate on this listing, not a guarantee; pay and task availability can change with collection volume.