What the work involves
You write the science-teaching situations that frontier labs use to probe whether a model actually understands science or just sounds like it does. A task might be: a student asks why a helium balloon rises but a hydrogen one rises faster, a class conflates heat with temperature after a calorimetry lab, or a model proposes a demonstration mixing household chemicals. You author the scenario, write the reference answer with the reasoning chain spelled out — not just the correct endpoint — and then grade model output against it.
The grading half is where judgment matters most. A lot of model science reads beautifully and is quietly wrong: a plausible-sounding mechanism, a conservation argument applied where it doesn't hold, a lab protocol missing eye protection or ventilation, an analogy that installs a misconception more durable than the confusion it replaced. Your job is to catch those and write a rationale a reviewer who isn't a chemist can follow.
What the platform screens for
- Verifiable credentials. Teaching licence or certification with a science endorsement, and at least two years in a real classroom. Expect to name the issuing state or authority and the courses you taught.
- Disciplinary depth. You'll be pushed on one of biology, chemistry, physics, or earth science until you either show command or run out. Claiming all four is usually read as claiming none.
- Misconception fluency. Whether you can name the specific wrong models students hold and why the usual explanation fails to dislodge them.
- Safety instinct. Whether you flag hazards in a protocol without being prompted to look for them.
- Written clarity. Reference answers and grading rationales are the deliverable; prose that wanders doesn't survive review.
Logistics
Fully remote, fully asynchronous — no live sessions, no calls with students, no grading deadlines at 11pm. You set your hours each week against a floor of about 10. Pay is hourly in the observed $45–75 range, disbursed weekly through Stripe; where an individual lands in that band typically tracks discipline scarcity and task complexity rather than tenure. Work is ongoing and project-based, so volume rises and falls with what labs are testing that month.