What the work actually is

You will read German output from a frontier model and decide whether it is right — not just grammatical, but idiomatic, register-appropriate, and faithful to source where a source exists. A typical batch mixes short tasks: rating two candidate translations against each other, rewriting a summary that is technically accurate but reads like translationese, flagging a dialogue turn that slips from Sie to du mid-conversation, marking a compound noun the model invented that no German speaker would form. Beyond ratings, you write the justification: a short, specific note that a researcher who does not read German can act on. Annotation passes ask you to tag features — case and agreement errors, verb-final violations in subordinate clauses, semantic drift, Anglicisms, tonal mismatch — against a rubric that changes between projects.

The literary side is real. Projects here care about genre convention, period register, and whether the model can hold a voice across a long passage, so experience in literary editing, criticism, or serious copywriting matters more than a generic translation background.

What the screen looks for

  • Verifiable German depth. Native or near-native command, plus the metalanguage to name what is wrong: not "sounds off" but "dative where the verb governs accusative," or "Bavarian lexical item in a neutral High German brief."
  • Documented professional work. Titles you edited, outlets you wrote criticism for, annotation schemes you have applied. Vague "years of experience" answers rarely survive follow-up.
  • Rubric discipline. Whether you can apply someone else's scoring guide instead of your own taste, and say clearly when the guide is underspecified.
  • Calibration. Consistent scores across similar items, and the willingness to give a middling rating rather than defaulting to the extremes.

Note that the posted description contains a visible copy-paste artifact — it mentions French-language tasks in the opening paragraph while every requirement and responsibility is German. Treat the role as German throughout; it is worth asking the screener to confirm.

Logistics

Remote and asynchronous, contractor engagement, work drawn from a queue with per-batch deadlines rather than fixed shifts. Contributors commonly report 10–20 hours a week, with volume that rises and falls as projects open and close. The $50/hr figure is the band observed on this listing, not a guarantee; rates on Mercor vary by project, assessment result, and task type. Expect a written or recorded screening, then a paid or unpaid calibration task before live work.