What the work actually is

You are producing training and evaluation data for frontier language models, in both Japanese and English. A typical shift is a queue of analytical prompts: read a block of content and break it into logical components, interpret a sales-by-month table and argue which location grew most (and defend the time window you chose), or construct a constrained scheduling puzzle of the postman-and-neighborhoods variety. For each item you write the model-facing question, the gold answer, and an explanation that shows the reasoning step by step — the explanation is often worth more than the answer, because it is what the model learns from. Other days you review model output and annotate where the chain of reasoning breaks, with enough specificity that someone else could reproduce your judgement.

Turing states no prior specialized domain experience is required. What it is genuinely screening for is whether you can hold a claim to account: validate assertions through online research, notice when a question is ambiguous rather than answering it anyway, and write cleanly in both languages. Professional writing or analyst background — business analyst, research analyst, journalist, technical writer, editor, translator — is listed as preferred and does tend to show in the sample work.

What the screen measures

  • Real bilingual fluency, not passive comprehension. Expect to be asked to produce analytical prose in Japanese, not just read it.
  • Arithmetic and data interpretation under scrutiny — percentage growth, base effects, small-sample volatility. Getting the number right matters less than saying why you framed the comparison that way.
  • Logic construction, meaning you can build a puzzle with exactly one valid solution and prove it, not just solve one someone else wrote.
  • Annotation discipline: feedback that cites the specific failing step rather than grading a response as vaguely weak.

Logistics

Fully remote contractor engagement — no medical or paid leave, no employment relationship. Up to 40 hours a week is preferred for the contract duration, with 2–5 hours daily overlapping America/Los_Angeles (UTC−8). For most of Japan and East Asia that means early-morning or late-evening hours; be honest in the screen about which end you can sustain. You need your own desktop or laptop and reliable internet. Pay is not disclosed in this listing; Turing's rates vary by language, task tier and geography, and extension depends on measured quality.