What the work actually is
You write physics problems that a strong model cannot shortcut, then you grade what the model does with them. A typical batch alternates between authoring — a clean problem statement, stated assumptions, a full worked solution, and a defensible final answer with units — and adjudication, where you read two or more model attempts at someone else's problem and mark exactly where the reasoning breaks. The interesting failures are rarely arithmetic: a model will conserve energy in a collision that isn't elastic, drop a boundary term, apply a non-relativistic dispersion relation past its validity, or produce a numerically correct answer through two cancelling sign errors. Your feedback has to name the step and the physics, not just flag the line.
Problem design is the harder half. The brief asks for items that reward conceptual reasoning over grinding, which in practice means problems where the setup forces a choice of frame, limit, or approximation, and where a wrong choice produces a plausible-looking wrong number. You will also be asked to sanity-check difficulty: a problem that every model solves and a problem no model can parse are both low-value.
What the screen looks for
- Verifiable depth in named subfields. Expect follow-ups that go one layer past your first answer — if you claim statistical mechanics, be ready to discuss ensemble equivalence or where a partition function argument fails.
- Willingness to state limits. Physics is broad; a candidate who scopes their strong areas honestly reviews better than one who claims all of it.
- Solution hygiene. Dimensional analysis, limiting cases, and stated assumptions are the currency of this work.
- Ambiguity handling. Whether you can tell an underspecified problem from a wrong answer, and what you do about each.
Logistics
Fully remote, contractor, asynchronous — no standing meetings, work claimed from a queue. Contributors commonly report 10–20 hours a week, with volume varying by project. Pay in the $50–80/hr band has been observed for this role; final rates depend on subfield, review tier, and calibration performance, and are set by the platform, not guaranteed here. Expect a paid or unpaid sample task involving one authored problem and one model-response critique before you are onboarded.