What the work actually involves
You write HR problems hard enough that a capable model gets them wrong, then you document why. A task might be a performance-management scenario where the documented write-ups don't support the termination, a compensation-banding question where internal equity and a competing offer pull in opposite directions, or an FMLA/ADA overlap where the naive answer is confidently illegal. Each item needs a defensible reference answer with reasoning — not just a key — because the value to the model is the explanation of the correct path, including which facts change the answer.
The second half of the job is evaluation. You read model responses that sound polished and decide whether they are actually right: whether the classification of exempt vs. non-exempt holds, whether the advice would survive an EEOC charge, whether the recommendation is generic HR blogspeak dressed up as analysis. Feedback is expected to be specific and citable — the regulation, the practice standard, or the fact pattern the model ignored.
What the screen looks for
- Concrete HR history: which functions you owned, at what headcount, in what jurisdictions. Vague "HR generalist" claims get probed hard.
- Whether you can state a defensible answer and then name what would change it. Rigid rule-recitation and mushy "it depends" both score poorly.
- Awareness that employment law is jurisdictional. Candidates who answer US-federal questions as if they were universal, or state-law questions as if they were federal, lose ground.
- Writing that is tight and unambiguous. Ambiguous prompts are unusable as training data.
Logistics
Fully remote and asynchronous, contract, project-based. Volume fluctuates with active projects; contributors commonly work in blocks of a few hours rather than fixed shifts. Pay for this listing has been observed around $50/hr — rates vary by project and are set by the platform, not guaranteed. Expect written calibration rounds and reviewer feedback on your first batch before volume opens up.