What the work actually is
You author teaching scenarios — a B1 learner overgeneralising the past simple, a young-learner class where the coursebook explanation of 'used to' is technically wrong, an IELTS Writing Task 2 script that needs graded feedback — and then write the reference answer a strong teacher would give. The reference answer is not the interesting part; the pedagogical reasoning behind it is. Labs want the trace: why this correction now, why this metalanguage at this level, why you'd let a minor error pass to protect fluency. You also grade model output against those references, and flag the specific failure the category exists to catch: an explanation that is linguistically accurate but would leave an A2 learner more confused than before.
What the platform screens for
- Verifiable credentials. TEFL, TESOL, CELTA, or a state ESL licence, plus at least two years of classroom or online teaching. Expect to name the awarding body and the contexts you taught in.
- Level calibration under follow-up. Screeners will hand you an explanation and ask which CEFR band it suits, then push on why. Vague answers about 'lower levels' get probed.
- Error-correction judgment. Which errors you treat, when, and what you deliberately ignore — with a reason that isn't just 'it depends on the learner'.
- Written clarity. Everything you produce is read by an annotator or a researcher who is not an ESL teacher. Rambling reasoning is unusable regardless of how correct it is.
Logistics
Fully remote and fully async — no live classes, no fixed shifts, no synchronous meetings. Work is claimed from a queue and hours are set week to week, with a stated floor of about 10 hours. Pay is observed in the $35–70/hr range, weekly via Stripe; the band typically reflects task complexity and calibration track record rather than seniority alone. Project work is ongoing, but volume fluctuates with lab demand — treat it as reliable supplementary work rather than a guaranteed full load.