What the work involves
You write and answer analytical problems that current models tend to get wrong, then document the correct reasoning in enough detail that a model can learn from it. In practice that means things like reading a monthly sales table across several locations and arguing which one actually grew the most — including which window to measure and whether to discount early-period volatility — or laying out constraint puzzles (visit order, scheduling rules, deduction chains) and showing the step-by-step path to the single defensible answer. Other days you'll be on the critique side: reading a model's output, deciding whether the conclusion is right, wrong, or right-for-the-wrong-reason, and writing annotations that pinpoint the exact step where it went off. Some tasks run in Norwegian, so the same standard of reasoning and clear prose has to hold in both languages.
What the platform screens for
Turing's process is assessment-heavy and takes roughly 80–90 minutes end to end: an automated analytical challenge (~30 min), a business writing assessment in English (~30 min), then a Norwegian language proficiency assessment (~20 min). No specialised domain background is required — the screen is testing whether you can decompose a messy prompt, do the arithmetic and the online fact-checking, and explain your reasoning in writing that a reviewer can follow without asking you a follow-up question. Professional writing history (analyst, journalist, editor, translator, technical writer) and comfort with Excel or Google Sheets both help, and a degree is preferred rather than required.
Logistics
- Contractor/freelance engagement — no paid or medical leave, extension depends on performance and project demand
- Fully remote; availability of up to 40 hours/week is preferred
- 2–5 hours per day must overlap US Pacific (UTC-8), which for a Norway-based contributor means afternoon-into-evening work
- You supply your own desktop or laptop and a reliable connection
- Pay band is not disclosed in this listing; ask for the per-hour or per-task rate and the expected task volume before you commit
Who tends to do well
People who enjoy being pedantic in a useful way — who will notice that a growth comparison is meaningless without a stated baseline, or that a puzzle as written admits two answers, and say so instead of guessing. Self-direction matters: task queues arrive with written guidelines and little live supervision, and the quality bar is enforced through review rather than coaching.