What the work actually involves
You receive video sessions — a robotic arm gripping, lifting, placing, colliding, slipping — and your job is the audio channel. That means two distinct tasks running together. First, restoration and editing: pulling down room tone, mains hum, handling thumps on the mic, clipping and distortion, so the usable signal is clear and consistent from session to session. Second, annotation: marking what is happening in the sound. A servo whine, the click of a gripper closing on nothing, the scrape of a mis-seated part, the difference between an object being set down and an object being dropped. Your written notes on those events are the labels the model learns from, so precision in language matters as much as precision in the edit.
Expect ambiguity. Real-world capture is messy, mic placement is inconsistent, and two contributors will hear the same 200ms transient differently. Much of the value you add is flagging those cases to project leads rather than guessing, and helping tighten the annotation protocol so the next hundred sessions are labelled the same way. You will also be asked to keep traceable records of what you changed — which edits were applied, in what order, to which file — because a denoised track with no edit history is unusable as training data.
What the platform screens for
micro1's screening is AI-led and conversational. It probes concrete craft: which DAW you work in and why, how you approach a noisy non-music source, how you'd describe a mechanical sound in writing so another annotator reaches the same conclusion. No AI or machine learning background is expected or asked about — domain depth is the whole test. Screeners tend to follow up on any claim you make, so name tools, plugin chains, and specific past projects rather than describing your process in the abstract. Experience with field recording, foley, dialogue cleanup, or forensic/technical audio review reads more strongly here than a music-mixing-only history.
Logistics
- Fully remote, contractor engagement, project-based rather than salaried.
- Largely asynchronous — you work through assigned session batches on your own schedule, with periodic sync on annotation standards and edge cases.
- Hours are typically flexible and variable by batch volume; observed pay for this listing sits in the $50–90/hr range, set by the platform and not guaranteed for any individual contract.
- You supply your own DAW, plugins, and monitoring setup. Accurate headphones or monitors matter more here than on most remote audio work, since the target detail is often very quiet.