What the work involves
Most tasks start with a realistic business artifact: a contract clause rendered into Japanese, a 決算 summary spreadsheet, a client-facing slide deck that has to land with a Japanese executive audience. You either produce the gold-standard version yourself or evaluate what the model produced against it. The judgment that matters is rarely binary — a translation can be lexically correct and still wrong because the 敬語 level insults the recipient, because 御社/弊社 usage is inverted, or because a direct English refusal was carried over into Japanese where indirection was required. Your written rationale explaining why a version fails is often worth more to the lab than the corrected text itself.
Work spans several recurring streams:
- Business and legal translation review, including contracts, 稟議 documents, and compliance material
- Localization and tone-adaptation QA — register, honorifics, formality ladders, audience calibration
- Cultural-context and etiquette assessments, including scenarios where the model gives advice that is fluent but socially wrong
- Native-fluency writing evaluation of long-form Japanese output
- Terminology and glossary management, including consistency enforcement across a document set
- Bilingual client-facing materials where both language versions must carry the same weight
What the platform screens for
Ethos runs an AI voice screen. Expect it to push on specifics: what you translated, in what domain, for whom, and what you did when the source text was ambiguous. Generic claims of bilingualism get probed until they produce a concrete example — a term you had to render without a clean equivalent, a client correction you accepted or pushed back on. The screen also tests evaluation instinct: given two plausible Japanese outputs, can you articulate a ranking that another reviewer could reproduce? Fluency alone is not the bar; the bar is fluency plus the ability to explain your judgment in writing that survives disagreement.
Logistics
Fully remote and asynchronous, with 5–20 hours per week and room for more if you want it. Work is drawn from a queue rather than scheduled, so you set your own hours. The $100/hour rate is what contributors on this project have reported; rates on AI evaluation work vary by task type and stream and are not guaranteed. Expect competence with Word, Excel, Google Sheets, PowerPoint, and Slides to be assumed rather than taught — formatting quality is part of what's being graded.