What the work involves

Once you're in the pool, the tasks vary by project but cluster around a few shapes. You write legal prompts hard enough to expose a model's weak spots — fact patterns with a buried limitations issue, a jurisdictional split, a statute amended in the last two years. You grade model responses against a rubric, or write the rubric yourself. You produce reference answers with reasoning shown, so the label carries the why and not just a verdict. And you audit for the failure modes specific to legal output: hallucinated case citations, real cases cited for holdings they don't contain, a rule stated correctly for the wrong jurisdiction, a confident answer where the honest answer is that authority is unsettled.

What the assessment screens for

The qualifying assessment is unpaid and that is stated plainly on the platform — treat it as the cost of entry to the pool. It is testing whether you can do two separable things: reason about a legal question correctly, and evaluate someone else's reasoning about it. Those are different muscles. Reviewers look for whether you check citations rather than recognising them, whether you can articulate why an answer is wrong in terms a non-lawyer annotator could apply consistently, and whether you distinguish an error of law from a defensible position you happen to disagree with. Vague grading — 'good analysis, 4/5' — reads as unqualified regardless of your credentials.

Logistics

  • Fully remote and asynchronous; work is claimed from a queue rather than scheduled.
  • Volume is project-driven and uneven — busy stretches followed by quiet ones. Most contributors treat it as supplementary rather than primary income.
  • Rates within the band track specialisation and task complexity; niche practice areas and rubric-design work sit higher than bulk response grading.
  • A JD, LLB, or equivalent qualification plus practice or academic experience is the baseline expectation. Bar admission is commonly asked about and matters more for some projects than others.