AI jobs for software engineers and STEM experts
Frontier AI labs pay engineers, data scientists and security researchers to grade model-written code, build the repositories and test harnesses models are measured against, and audit the benchmark tasks that train them.
- Open roles
- 110
- Median rate
- $80/hr
- Publish an hourly rate
- 89%
listings, updated continuously
from 98 published bands
the sample every median here uses
What each STEM and coding specialty pays
Every listing we track, grouped by specialty. Rates are read from the platform's own posting — the median is of published hourly bands, and the range runs from the lowest published floor to the highest published ceiling. Per-task listings are counted but not priced.
| Specialty | Roles | Median | Published range | Distribution |
|---|---|---|---|---|
| Machine Learning, AI & Data Science | 14 | $110n=13 | $50–280 | |
| Security & Red Teaming | 9 | $95n=8 | $50–250 | |
| General Software, IT & AI Evaluation | 6 | $83n=6 | $30–160 | |
| Software Engineering | 37 | $80n=33 | $25–280 | |
| Cloud, DevOps, Networking & IT | 12 | $80n=12 | $45–165 | |
| Data Engineering & Analytics | 9 | $80n=9 | $30–120 | |
| GPU & Kernel Engineering | 5 | $80n=5 | $60–170 | |
| STEM PhD Coding | 5 | $70n=5 | $70–70 | |
| Application Users — Developer Tools | 5 | $60n=5 | $55–70 | |
| Nuclear, Civil & Physical Engineering | 8 | Too few to price | — |
Open STEM and coding roles
The listings behind the figures above, richest published ceiling first.
What engineers do in AI
Grade model-written code
Review AI-generated patches, pull requests and coding-session traces for correctness and sound workflow, and write rubric-anchored feedback a lab can act on.
Build the environment a model is tested in
Turn a real repository into a task: a pinned commit, a Dockerised environment, a failing-to-passing test signal and a golden solution you sharpen.
Audit benchmark tasks before they ship
Check reference patches, test harnesses and grading scripts for correctness, reproducibility and leakage, then deliver a written verdict on whether the task enters the set.
Write the problems models still fail
Author executable scientific-computing, kernel or security challenges that frontier models get wrong more often than right, with the criteria that define a correct answer.
Work directly with an AI lab
Some listings are full-time roles embedded with a lab's research team writing instruction specs and benchmarks — W-2, and sometimes hybrid on-site in a specific city.
Who this work is for
Backgrounds in demand
- Software engineers, full-stack & backend37roles
- ML engineers, data scientists & AI researchers14roles
- Cloud, DevOps, SRE & network engineers12roles
- Security researchers & red teamers9roles
- Data & analytics engineers, BI analysts9roles
- GPU, CUDA & accelerator-kernel engineers5roles
- Nuclear, civil & process-safety engineers8roles
Commonly asked for
- Code review & patch grading
- Reproducible test harnesses
- Docker-based task environments
- Benchmark auditing & leakage checks
- Reference solutions & rubrics
- Python, Java, Go, C, C#, Rust
- CUDA, Triton & accelerator kernels
- Infrastructure as code
- Vulnerability research
- Statistics & experiment design
- Rubric-based evaluation
- Written technical verdicts
Drawn from the listings themselves and not counted — the skills field on a listing is free text, and no ranking of it would mean anything.
Platforms hiring engineers and STEM experts
- micro1Roles34Median$85Published range$30–280
- EthosRoles6Median$85Published range$80–150
- MercorRoles38Median$80Published range$25–250
- AfterQueryRoles20Median$80Published range$30–200
Figures are calculated over the STEM and coding listings each platform currently has open, on the same basis as the rate report.
Also join these expert networks
- PROFILE ENTRYEthos
One profile with Ethos, matched to paid opportunities across fields as they open — not an application to a single listing.
Create profile on Ethos → - PROFILE ENTRYHandshake AI
One profile with Handshake AI, matched to paid opportunities across fields as they open — not an application to a single listing.
Create profile on Handshake AI →
How AI expert work actually works
- 4 min read
How Coding and STEM AI Evaluation Works
Check specifications, reproduce results and distinguish a convincing explanation from code or reasoning that actually works.
- 4 min read
How to Evaluate AI Agents and Multi-Step Tasks
Inspect the result, the tool-use record and the constraints an agent had to respect. A confident final message is only one piece of evidence.
- 5 min read
How to Pass an AI Evaluation Interview or Assessment
Prepare to demonstrate your reasoning, follow the assessment rules and explain your experience without scripts or invented credentials.
- 6 min read
Mercor vs micro1: How to Compare AI Evaluation Platforms
Compare the application route, project requirements and engagement terms. Choose a specific opportunity rather than a platform-wide promise.
Questions, answered plainly
Mostly four things: grade model-written code and coding-session traces against a rubric; build the repositories, containers and test suites a model is measured in; audit benchmark tasks for correctness and leakage before they enter a training or evaluation set; and write the hard problems models still fail, with the criteria that define a correct answer. Engineering judgement is what is being bought.
The median across the listings we track is at the top of this page, and each specialty's median and published range is in the table above it — every one with the number of listings behind it. Not every listing here is hourly: some are paid per accepted task, and those are counted in the total but enter no median. Rates are read from the platform's own posting; where a platform publishes nothing, we show nothing.
Not for most of the board. Much of it asks for working software engineers — backend, full-stack, QA — and grades on engineering judgement, not on ML. The ML and data-science rows are the exception and say so in the title. Each role page separates what a listing requires from what it prefers.
Most listings here are project-based contract work paid by the hour and described as remote. A run of Mercor listings is paid per accepted task instead — codebase tasks, dual-use red-team prompts — and some are full-time W-2 roles, sometimes hybrid on-site in a specific city. Some are restricted to the United States. Each role page repeats what its listing says; we publish no remote percentage, because a listing that says nothing about location is not the same as one that says “anywhere”.
Other fields
New STEM and coding roles, by email
One message on the days STEM and coding listings open. Nothing on the days they do not.


