Data Lab
Nobody hands you a clean table and a question
A question bank gives you five tidy rows and tells you what to compute. Data Lab gives you one Malaysian ride-hailing warehouse — 12 connected tables, 300,000 rides, dirty on purpose, no indexes to start — and you work it chapter by chapter, the way you'd work a real one.
Beta — included free with Basic for now.
Clean it, join it, aggregate it, make it fast — then read a brief with no column list and no stated metric, and defend the number you come back with.
Clean it
Find the nulls, reconcile four spellings of one city, spot the duplicate riders, and catch the orphaned foreign key — on a warehouse big enough that these actually cost something.
4 stepsChapter 2Join it
Join rides across the estate, catch the fan-out trap that quietly doubles a total, and mix INNER and LEFT joins on purpose instead of by accident.
4 stepsChapter 3Aggregate it
Pick the right grain, rank and share-of-total with window functions, and build a running total — the shape every real dashboard number is actually built from.
4 stepsChapter 4Why your query is slow
Read an EXPLAIN ANALYZE plan on a 300k-row table and find the seq scan that's costing you.
5 stepsChapter 5Add the index
Create a composite index on the same 300k-row table and watch the plan — and the timing — change.
3 stepsChapter 6Read the brief
One stakeholder brief, no column list, no stated metric — find the right tables across the whole warehouse and defend your number.
1 step + stakeholder brief