The work

You will be assigned one of two tracks. Authoring means writing original questions in your subdomain — National Security, Public Policy, Business History, Environmental History, or Latin American History — that test conceptual command rather than date recall. Each item needs a self-contained prompt, one defensible correct answer, nine distractors that a well-read specialist could plausibly select, a difficulty rating (Medium for intro undergraduate, Hard for advanced undergraduate, Expert for postgraduate and above), a step-by-step chain-of-thought solution in markdown, and one to five references from peer-reviewed journals or university repositories.

Verification means reading pre-written items adversarially: is the question answerable from the stem alone, is exactly one option correct, is any distractor secretly defensible, does the stated difficulty match reality? You edit where needed and document what you changed and why. Reviewers in this format spend much of their time on interpretive ambiguity — the historiographical claim stated as settled fact, the policy question where a competing school of thought makes a second option correct.

What the screen looks for

  • Verifiable credentials and subdomain specificity. A PhD or doctoral candidacy in History, Political Science, or International Relations; master's degrees are considered where subdomain depth is exceptional. Expect to name your dissertation area, advisors, and publications.
  • Depth under follow-up. AI-led interviews probe a claim two or three layers down. Vague answers about "comparative methods" fare worse than a precise account of one debate you can argue both sides of.
  • Distractor craft. The hardest skill here is writing wrong answers that are wrong for a reason a specialist would have to reason through — not obviously implausible, not arguably correct.
  • Sourcing discipline. Citations must be real, locatable, and load-bearing for the answer.

Logistics

Fully asynchronous, no fixed hours, no meetings. Expect around 10+ hours per week, with volume varying by domain demand and how much rework your items require in review. Pay in the $44–56/hr band has been observed for this listing; rates are set by the platform and vary with track, credential level, and assessed quality — treat the band as reported, not promised. Work is task-based, so continuity depends on maintaining acceptance rates.