What the work actually is

You spend your day inventing and answering the kind of question a frontier model gets subtly wrong. Two categories dominate: quantitative interpretation (given monthly sales across locations, which grew the most — and what time window makes that claim defensible?) and constraint logic (delivery-route puzzles, ordering rules, elimination problems). For each item you supply the task, the verified correct answer, and a step-by-step rationale that makes the reasoning auditable. Some tasks are authored in Korean or require you to check that an English-language reasoning chain survives translation intact. Other days lean toward review: reading two model responses, deciding which reasons better, and annotating precisely where the weaker one broke down.

Research is part of the job. Claims embedded in source content need verification online, and a task is only usable if its answer is genuinely unambiguous — a surprising amount of the craft is noticing that your own question has two defensible answers and repairing it before submission.

What the screen looks for

  • Genuine Korean and English working fluency. Not conversational Korean — the ability to write technical explanation and argue about nuance in both.
  • Arithmetic and data honesty. Percentage change, base effects, small-sample volatility. Screens frequently hand you a tiny table and watch whether you state your assumptions or just assert a winner.
  • Explanation quality over answer quality. Right answer, hand-waved reasoning scores worse than a careful walkthrough.
  • Written feedback that localises the error. "Response B is wrong" is worthless; "Response B treats the Oak-before-Pine rule as applying within a day rather than across the schedule" is the deliverable.

No prior AI or specialist domain background is required. Professional writing experience — analyst, journalist, editor, translator, technical writer — is the common thread among people who do well.

Logistics

Fully remote contractor engagement: no medical or paid leave, no equity, hourly. Availability of up to 40 hours per week is preferred, with 2–5 hours per day overlapping America/Los_Angeles (UTC-8) for handoffs and calibration discussions. You need your own desktop or laptop and reliable internet. Contracts are project-scoped with extension possible on performance. Pay is undisclosed on this listing; Turing typically states a rate at offer stage, and you should ask before committing to a 40-hour week.