What the work involves
You build the test material and then judge what the model does with it. On a typical task you author a realistic instruction scenario — a ninth grader's argumentative paragraph with a buried thesis, a student who has confused summary for analysis in a response to Their Eyes Were Watching God, an English learner whose sentence-level errors are obscuring a genuinely sophisticated claim — and specify the level, the assignment, and what the student needs next. Then you write the feedback you would give and, crucially, the reasoning behind it: why you address the organizational problem before the comma splices, why you name what the student did well in specific terms rather than generic praise.
The second half is evaluation. You read model-generated feedback and literary analysis and grade it: is the reading of the text defensible, is the feedback prioritized correctly, would a student who followed it produce better writing? A recurring category is feedback that is warm, encouraging, and useless — praise that names nothing, suggestions the student cannot act on, a revision note that rewrites the sentence for them instead of teaching the move. Flagging that gap, and articulating precisely why it fails, is most of the value you add.
What the platform screens for
- Verifiable credential and classroom time. Certification with an English or language arts endorsement, plus two or more years actually teaching, not tutoring or curriculum design alone.
- Depth under follow-up. Screeners push past your first answer — expect to be asked what you would say to a specific student, then why that and not something else.
- Evaluation judgment. Whether you can separate feedback that sounds good from feedback that changes the next draft, and whether you can defend a literary reading as sound, thin, or wrong.
- Written register. Your prose is the product; sloppy writing in the screen is disqualifying in a way it would not be for other roles.
Logistics
Fully remote, fully asynchronous, no live sessions and no evening grading blocks. You set your own hours week to week with a floor of roughly 10 hours; work is ongoing rather than a fixed contract. Pay is hourly via Stripe on a weekly cycle, with an observed band of $35–65/hr — placement within it typically reflects level taught, subject depth, and calibration performance rather than being negotiated up front.