What the work actually involves
You build tasks that a competent paralegal could answer and a language model probably cannot. That means writing scenarios with enough factual grain to be gradeable: a discovery deadline computed under a specific rule set, a Bluebook or ALWD citation with a planted defect, a privilege log entry that omits the basis for withholding, a subpoena that names the wrong custodian. Each task ships with a reference answer and a rubric explaining why alternative answers are wrong — not merely that they are.
The second half of the job is evaluation. You read model output against your rubric and write feedback that a non-lawyer reviewer can apply consistently. Most model failures in this domain are plausible-sounding: an invented pin cite, a state procedure applied in federal court, a filing deadline that ignores the weekend rule, confident advice that crosses into unauthorized practice of law. Naming the failure precisely is worth more than a long critique.
What the platform screens for
- Verifiable paralegal background: where you worked, what practice areas, what you personally handled versus supervised.
- Procedural specificity under follow-up — screeners push on rules, deadlines, and citation conventions to see whether recall is real or paraphrased.
- Whether you can tell a wrong answer from a differently-worded right one, and can articulate the line.
- Awareness of the UPL boundary and of jurisdictional variation, since much of the training value lives there.
Logistics
Fully remote and asynchronous — no standing meetings, no fixed hours. Contributors typically work in blocks of a few hours and are paid hourly at an observed $50–60/hr; rates and available volume vary by project and are not guaranteed. Expect a calibration round of a few sample tasks before steady assignments, and expect written feedback on your early submissions. Work is contractor-based, sourced by AfterQuery for its model-training clients.