What the work actually involves

You build the material an AI lab uses to find out whether a model understands teaching. That means authoring scenarios that span the real arc of instructional work — a unit plan for a mixed-ability Year 8 class, a reading intervention for a student two grade levels behind, a formative assessment that has to survive a 40-minute period with no aide in the room. For each scenario you write a reference answer and, critically, the pedagogical reasoning behind every decision: why this scaffold and not that one, why the exit ticket is three questions instead of ten, what a teacher would actually do when the timing collapses.

The second half of the job is grading. You read model-generated lesson plans, feedback comments, differentiation strategies and parent communications, and score them against rubrics for instructional quality, developmental appropriateness and factual accuracy. The most valuable thing you contribute is catching output that is fluent and plausible and would still fail on Monday morning — activities that assume equipment nobody has, group work with no accountability structure, "differentiation" that is just the same worksheet with a larger font, feedback that would demoralise a struggling fourteen-year-old.

What the screen looks for

AfterQuery's screening is AI-led and follows up on whatever you say. It is looking for a verifiable credential and real classroom time, then for whether you can articulate pedagogical reasoning rather than recite frameworks. Expect to be pushed on specifics: which grade, which subject, what you did when a lesson failed, how you would rewrite a weak model answer and why. Naming Bloom's or UDL is not evidence of depth; explaining what you changed for one particular student is.

Logistics

  • Fully remote and fully async — no live classes, no scheduled calls, no grading nights
  • Minimum commitment around 10 hours per week, with hours you set yourself
  • Paid hourly via Stripe on a weekly cycle; the $35–70 range is as observed on the platform and varies by task type, subject scarcity and calibration performance
  • Ongoing project work rather than a fixed-term contract, with task volume varying by lab demand