What the work involves
Most days split between authoring and reviewing. Authoring means writing original problems hard enough to break a strong model — multi-step, abstract, or proof-based items where the answer cannot be pattern-matched from a textbook — and then writing your own complete solution with every justification spelled out. Reviewing means reading model-generated proofs closely enough to catch the failure modes that matter: an unjustified step, a case left unhandled, a limit interchange nobody earned, a correct final number reached by invalid reasoning. Your annotations have to name the error and locate it, not just mark the response wrong.
Two other streams appear on most project rosters. Computational tasks ask you to design closed-ended prompts with verifiable numeric answers, implement the solution in Python using approved scientific libraries, and document the rationale. Theorem-prover tasks ask you to translate informal statements and proofs into Lean and confirm the formal proof compiles — this is where candidates most often overstate their experience, and it is checked.
What the screen looks for
Turing's screening is AI-led and leans on follow-up. Expect to be asked which areas of mathematics you can defend at research depth and which you would only grade at undergraduate level, then pushed on specifics inside whatever you claim. Expect at least one scenario where a model output is subtly wrong and you must say precisely where. Concrete answers — a named theorem, a named library, the actual shape of a Lean tactic block, a real error you once caught — score. Broad claims of comfort across all of mathematics do not.
Logistics
- Fully remote, contractor engagement — no medical or paid leave.
- Commitment tiers of 20, 30, or 40 hours per week, at least 4 hours per day.
- Four hours of daily overlap with PST is required regardless of tier.
- Your own desktop or laptop and a stable connection.
- Contract extension depends on performance and project demand; pay is not disclosed in the listing.