What the work actually is
You receive audio files — some clean studio reads, many closer to real life: phone calls, overlapping speakers, dialect-inflected Mandarin, code-switched Mandarin-English sentences. Your job is to render what was said, verbatim, into text that matches a written style guide: how to handle 嗯/啊 fillers, false starts, laughter, inaudible spans, speaker labels, numbers and measure words, traditional vs. simplified output, and where English tokens sit inside a Chinese sentence. A large share of the task is not typing but decision-making under an ambiguous guideline, then applying that decision consistently across every file in the batch.
Batches usually include a proofing pass — either self-review or reviewing another annotator's transcript — and periodic gold-standard audits where your output is compared against a reference. Sustained accuracy on those audits is what keeps work flowing and drives movement within the pay band.
What the screen looks for
micro1's process is AI-led: a conversational interview, often followed by a short transcription sample. It is checking that you are genuinely native or near-native in Mandarin, that you can hear and reproduce accented or noisy speech rather than guess at it, and that you understand the difference between verbatim transcription and tidying someone's grammar. Expect follow-ups that push on specifics — name the tool, name the dialect, describe the tag set you used. Vague claims of "years of transcription experience" collapse quickly under that pressure.
Logistics
- Fully remote, contractor engagement, asynchronous — you pull batches and hit deadlines rather than keep shift hours.
- Realistic throughput on difficult audio is roughly 4–6× real time, so a one-hour file is most of a working day; plan commitments accordingly.
- You will need a quiet workspace, reliable headphones, stable internet, and a Chinese IME you are fast in.
- No prior AI or machine-learning background is required or expected.