What the work actually is
This is adversarial research design, not subject-matter expertise and not writing. Each task begins with a fact you can point to in a primary source — a line in a municipal register, a figure in a table buried in an annual report PDF, a date on a gravestone index, a quantity in a patent family. You then work backwards to build a natural-language question whose answer is that fact, constructed so that the path to it defeats a state-of-the-art browsing agent with full web access and multiple attempts. Clues must span several fact types (dates, people, places, organisations, works, events, records, quantities) and each must be independently checkable on its own.
The deliverable is mostly documentation. A validation record shows the obvious searches a capable agent would run, the exact queries, and what those queries returned — proving the shortcut doesn't exist. Sourcing is expected at page, table, and section granularity; a homepage link is a rejected task. Familiarity with JSON and structured delivery formats helps, since submissions land in a defined schema.
What the screen looks for
Turing's assessment is the real gate, and it tests research tradecraft over credentials: can you locate primary records, navigate government and institutional databases, archives, and registries, and cite precisely from inside a PDF? Expect probing on how you distinguish a question that is genuinely hard from one that is merely ambiguous or unanswerable — benchmark items need short, stable, objectively verifiable answers, and a question with two defensible answers is a defect. Backgrounds that map well: reference librarianship and special collections, archival research, investigative journalism and professional fact-checking, OSINT, due diligence and KYC, patent and prior-art or legal-discovery search, genealogy, and competitive quizzing or puzzle-hunt construction. Prior LLM evaluation, red-teaming, or benchmark construction experience is listed as a qualification, not a nice-to-have.
Logistics
- Fully remote, paid in USD, must be located in the United States
- Eight-week project, ideally up to 40 hours/week, asynchronous
- Observed pay is approximately $60 per approved task under the current task mix — approved being the operative word, since incomplete evidence trails don't pay
- Master's degree or more than three years of relevant experience; native or near-native written English
- Start follows successfully passing the assessment