What the work actually is

You write questions that a strong model gets wrong, then you prove what the right answer is. In practice that means picking a text, period, or debate you know well; constructing a prompt that requires close reading, contextual inference, or the reconstruction of an argument; running it against the model; and iterating when the model answers correctly. Each submission travels with an answer, a reasoning trace, and citations to primary texts or reputable scholarship — a claim you cannot source is a claim you cannot submit. Reviewers send items back, and applying that feedback cleanly is part of the job, not an exception to it.

The hard part is difficulty calibration. Obscurity is cheap: asking for the fourth stanza of a minor Elizabethan lyric produces a wrong answer without testing anything. What the project wants is items where the model has all the material in front of it and still misreads — ambiguous pronoun reference in a translated passage, a philosopher's argument that superficially resembles a more famous one, an attribution question where the conventional account is contested in the literature. Ambiguity is also disqualifying in the other direction: if two competent scholars would defend different answers, the item fails.

What the screen looks for

  • Verifiable depth in a named discipline. Expect follow-ups that go a layer past your first answer — sources, dating, historiographical disputes, the state of the debate. Vagueness is read as absence of expertise.
  • Citation discipline. Editions, translations, page or line references, and an honest distinction between primary evidence and secondary interpretation.
  • Evaluation judgment. Can you tell a genuinely hard item from a trivia question, and can you explain why a model's plausible-sounding answer is wrong?
  • Writing precision. Questions that admit exactly one defensible answer, stated in unambiguous prose.

Logistics

Fully remote, asynchronous, contractor engagement. Pay is per accepted task rather than per hour; the $40–90/hr band reflects rates observed by contributors on comparable micro1 humanities projects and depends heavily on how fast you can produce items that clear review on the first pass — it is not a guaranteed rate. A weekly minimum submission volume applies. Roles are typically filled within 48 hours and first tasks are expected within 24–48 hours of onboarding, so this suits people with immediate capacity rather than those planning around a future term break. No prior AI experience is required.