What the work looks like
You take on discrete, project-scoped tasks that sit between machine learning and computational materials science. In practice that means writing and evaluating ML pipelines against simulation-derived data — DFT relaxations and band structures, MD trajectories, Monte Carlo sampling output — and producing property-prediction, materials-discovery, or structure-property modelling artefacts with the reasoning made explicit. Tasks frequently ask you to build a defensible baseline, state your assumptions, and document why a given featurisation, target, or validation split is or isn't appropriate for the underlying physics. Some tasks are generative (write the reference solution); some are evaluative (judge whether a model's output or a chain of reasoning holds up against what the simulation data can actually support).
What the screen is actually checking
The first-author publication requirement is a real gate and is verified — the platform wants people who have defended methodology in review, not people who have read about it. Beyond that, screening probes whether your materials knowledge and your ML knowledge are genuinely joined: whether you know why random k-fold on a materials-project-style dataset leaks, why an interatomic potential fit on near-equilibrium configurations fails in the repulsive regime, what a DFT functional choice does to the label noise on your target property. Expect follow-ups that go one layer deeper than your first answer. Vague fluency in both fields separately reads as a fail; specific, falsifiable statements about a system you actually computed read as a pass.
Logistics
- Fully remote, fully asynchronous — no standing meetings, no timezone requirement.
- 10–20 hours per week on accepted projects, over a 2–3 week engagement window.
- Project-based; you accept or decline individual assignments.
- Pay observed in the $150–200/hour range, not guaranteed and generally tied to task complexity and calibration performance.
- Deliverables are code plus written methodology, so reproducibility and clear documentation matter as much as model performance.