What the work involves

You'll build reference artifacts and critique model output across the deliverables BI analysts actually ship: executive KPI dashboards, monthly business performance decks, ad hoc SQL pulls, data quality and validation checklists, metric definition glossaries, and self-serve reporting templates for non-technical stakeholders. A typical task might give you a messy prompt — "summarize Q3 marketing performance for the exec team" — and ask you to produce the gold-standard deck, or to evaluate a model's attempt and write down precisely where it went wrong: a churn metric defined inconsistently between slides, a SQL pull that silently drops rows via an inner join, a chart that implies causation from a correlation.

The written justification matters as much as the score. Ratings that say "the query is incorrect" without naming the join semantics or the null-handling assumption are of little use to the lab. Expect to write a paragraph or two of specific, reproducible reasoning per item.

What the platform screens for

  • Depth under follow-up. The AI voice screen asks about a dashboard you built, then keeps going: what grain, which metrics, how you handled late-arriving data, what stakeholders got wrong about it.
  • SQL fluency spoken aloud. You'll be asked to reason through joins, window functions, and aggregation traps verbally, without an editor.
  • Metric judgment. Whether you can define a business metric precisely enough that two analysts would compute it identically.
  • Artifact craft. Evidence that you actually build the spreadsheets and decks yourself, not just review them.

Logistics

Fully remote and asynchronous — tasks are queued, you pick them up when you want. Flexible 5–20 hours per week, with more volume available to consistent contributors. The $80/hour figure is what this listing observed; rates on evaluation platforms vary by task type, calibration performance, and project phase, and are not guaranteed. Some projects add a paid calibration round before live work begins.