What the work actually involves
Most days you are reading Spanish that a model produced and deciding whether it is correct, natural, and appropriate for its context — three separate judgments that often disagree. A sentence can be grammatically flawless and still read like it was translated from English; a colloquial rewrite can be idiomatic in Buenos Aires and jarring in Monterrey. You will rate outputs against a rubric, rewrite the weak ones, and write the short justification that tells a researcher why the output failed: calque, register mismatch, subjunctive misuse, false cognate, invented regionalism, dropped clitic. Annotation passes are also common — tagging text for syntactic structure, semantic roles, or error type according to a codebook that changes between projects.
What the screening looks for
- Verifiable Spanish work history. Translation, editing, proofreading, terminology, corpus annotation, or NLP data work — with specifics: language pair, subject matter, variety of Spanish, volume.
- Metalinguistic vocabulary. You should be able to name what is wrong, not just feel it. "Uses aplicar para instead of solicitar, an anglicism" beats "sounds off."
- Rubric discipline. Screeners test whether you can hold a scoring scale steady across dozens of items rather than drifting toward your personal stylistic preferences.
- Regional honesty. Nobody is a native speaker of every variety. Candidates who state which varieties they can adjudicate and which they cannot score better than candidates who claim all of them.
Logistics
Fully remote and largely asynchronous, run through Mercor's own platform: an AI-led interview first, then task calibration if you pass. Hours are typically flexible and self-scheduled, but projects come with weekly throughput expectations and quality thresholds — sustained availability of roughly 15–20+ hours a week is what most of these engagements ask for. Work is contract, project-scoped, and can pause or ramp without much notice. The $50/hr figure is what this listing has been observed at; effective rates vary by project, task type, and calibration performance.