What the work involves

You will spend most of your time on two linked tasks. The first is authoring chemistry problems at the level a qualifying exam or a graduate seminar would set — mechanism elucidation from spectral data, thermodynamic or kinetic derivations, coordination chemistry and ligand field reasoning, separation and detection method design — each paired with a complete, defensible solution and the reasoning chain that gets there. The second is evaluating what a model produces against that solution: deciding whether an answer is correct, correct-by-accident, or plausibly worded nonsense, and writing feedback that names the exact step where the chemistry breaks.

The hardest part is usually the second. Frontier models write fluent chemistry. A proposed mechanism can have correct arrow-pushing for four steps and then invoke a pentavalent carbon; a rate law can be dimensionally consistent and thermodynamically impossible. Your value to the project is catching that and articulating it precisely, not simply marking the final answer wrong.

What the screen looks for

  • A Master's or PhD in chemistry, chemical engineering, biochemistry or an adjacent field, and at least two years in academic research, industrial R&D, or applied chemistry.
  • Genuine subfield depth — the screen will follow up on whatever you claim, so specificity about your actual bench or computational work beats breadth.
  • Item-writing judgment: whether you can build a question that is hard for a reason, not hard because it is ambiguous or under-specified.
  • Written precision. Feedback that a model trainer can act on is the deliverable; prose that gestures at "this seems off" is not.
  • Publications or technical reports help, as does prior exposure to AI/ML work, but neither is a gate.

Logistics

Fully remote and asynchronous, organised as project-based batches rather than shifts. Contributors typically pick up work in blocks of several hours and turn tasks around within a stated window; volume fluctuates with which subfield a given project needs. Observed rates run $50–100/hr, with the upper end tied to scarcer specialisations and review-level tasks — treat that as reported range, not a guarantee. Expect a paid or unpaid calibration task before steady work begins.