The work
You will receive Mandarin audio clips — call recordings, interviews, unscripted conversation, read speech, sometimes noisy field recordings — and produce verbatim or lightly cleaned transcripts in simplified Chinese according to a written style guide. The guide is the job: it will specify how to handle filler words, false starts, laughter, code-switching into English, numbers and dates, proper nouns, and inaudible segments. A second stream of work is review — auditing another transcriber's output, marking errors by category, and deciding whether a file passes, needs correction, or should be rejected as unusable audio.
Expect regional accents (Northeastern, Sichuanese-inflected Mandarin, Taiwanese Mandarin), speakers talking over each other, and audio that is genuinely hard rather than merely fast. Speaker diarisation and timestamping are common add-ons. Some projects extend into evaluating machine-generated transcripts: comparing an ASR output against the audio and rating it, or writing the gold-standard reference that a model will be scored against.
What the screen looks for
- Native or near-native Mandarin listening comprehension, including dialect-influenced speech, plus accurate written Chinese
- Evidence you can follow a detailed style guide precisely rather than transcribing to your own preference
- Judgment about ambiguity: when to tag `[inaudible]`, when to guess, when to flag a file as out of scope
- English good enough to read guidelines, file structured issue reports, and communicate with project leads
Prior transcription, subtitling, court or medical reporting, translation, or linguistics work is the usual background, but demonstrated accuracy on the platform's own audio test carries more weight than a CV line.
Logistics
Fully remote and largely asynchronous, with per-batch deadlines rather than fixed shifts. Contributors commonly report 10–25 hours a week when a project is live, with gaps between projects; work is contract, not salaried. You will need a quiet environment, reliable internet, and headphones good enough to distinguish sibilants and low-volume speech — laptop speakers will not survive an audit. A Chinese IME you can type quickly in is assumed.