What the work actually is

You write BI problems the way they arrive in real jobs: a messy source schema, a stakeholder request that is half-specified, a metric definition nobody agrees on. Each scenario runs from data model through the dashboard build and the rollout decision — who sees it, what filters default to, what the executive summary tile is allowed to claim. Then you write the reference answer, and the reference answer is the real deliverable: not just the chart you'd build but the reasoning for the mark type, the axis treatment, the aggregation grain, the metric definition you rejected and why.

The second half is grading. You read model-produced dashboards and SQL and judge them on two axes that come apart constantly: analytical accuracy (is the number right, does the join fan out, is the denominator what the question asked for) and chart integrity (does a truncated axis manufacture a trend, does a dual-axis pairing imply a correlation the data doesn't support, does a stacked area hide a declining segment). The brief explicitly asks you to flag dashboards that look polished and still mislead — that failure mode is the point of the project, not an edge case.

What the screen looks for

  • Depth in at least one of Tableau, Power BI, or Looker that survives follow-up: LODs and context filters, DAX filter context and CALCULATE, or LookML explores and symmetric aggregates.
  • SQL you can defend — window functions, join grain, and how you validated a dashboard number against the warehouse.
  • Written reasoning. You are documenting why, and the screen reads your prose as a work sample.
  • Evaluation judgment: can you separate a chart you'd have built differently from a chart that is actually wrong.

Logistics

Fully remote and fully async — no standups, no live calls with labs. You pick hours each week against a 10 hour floor; contributors report clustering work into two or three sessions rather than daily slices. Payment is weekly via Stripe. The $50–90/hr band is as observed on the platform and typically tracks tool depth, SQL strength, and whether your reference answers need editing before they ship. Work is ongoing rather than a fixed contract, and volume varies with lab demand.