What the work actually involves
You spend your day on two linked tasks. The first is prompt authoring: constructing questions hard enough that a frontier model plausibly fails them — comparative electoral system trade-offs, treaty interpretation, coalition formation mechanics, the difference between what a statute says and how it has been implemented. The second is evaluation: reading model outputs closely and marking where they invent a citation, collapse a contested question into one partisan framing, quote a constitutional provision that was amended, or reason from an election result that has since been superseded. Every judgement you record needs an evidence trail — a source, a specific passage, a reason the model's claim fails — because your rationale is training signal, not just a score.
Politics is unusually hard to annotate well, and Turing knows it. A large share of the work is handling normative questions without smuggling in your own priors: describing how proponents and critics of a policy each reason, distinguishing empirical disputes from value disputes, and flagging when a model's confident tone outruns genuine scholarly consensus. Expect adversarial test-case design, benchmark dataset construction, and written documentation of your labelling decisions so other annotators can apply them consistently.
What the screen looks for
- A master's degree or higher — the field is open, but politics-adjacent disciplines (political science, public policy, IR, government, law, economics, journalism) are preferred.
- Three or more years of professional work in policy, government, political research, consulting, academia, journalism, think tanks or international affairs.
- Depth that survives follow-up questions. Screeners push past the first answer to see whether you can cite specific institutions, dates, cases or data rather than gesture at themes.
- Written English good enough that your rationales read as usable feedback for researchers who are not political scientists.
- Prior LLM, prompt engineering or AI evaluation exposure is a plus, not a gate.
Logistics
Remote contractor assignment, no medical or paid leave, eight-week duration. The commitment is stated as 40 hours per week with a minimum four-hour overlap with Pacific time, so this is not comfortably stackable with another full-time role. Pay is not disclosed in the listing; treat any figure you hear elsewhere as observed rather than promised. Entry runs through a take-home assessment that Turing shares after application, followed by internal review.