Menu▾

Data Lab

Nobody hands you a clean table and a question

A question bank gives you five tidy rows and tells you what to compute. Data Lab gives you one Malaysian ride-hailing warehouse — 12 connected tables, 300,000 rides, dirty on purpose, no indexes to start — and you work it chapter by chapter, the way you'd work a real one.

Beta — included free with Basic for now.

Clean it, join it, aggregate it, make it fast — then read a brief with no column list and no stated metric, and defend the number you come back with.

Chapter 1

Clean it

Find the nulls, reconcile four spellings of one city, spot the duplicate riders, and catch the orphaned foreign key — on a warehouse big enough that these actually cost something.

4 steps
Chapter 2

Join it

Join rides across the estate, catch the fan-out trap that quietly doubles a total, and mix INNER and LEFT joins on purpose instead of by accident.

4 steps
Chapter 3

Aggregate it

Pick the right grain, rank and share-of-total with window functions, and build a running total — the shape every real dashboard number is actually built from.

4 steps
Chapter 4

Why your query is slow

Read an EXPLAIN ANALYZE plan on a 300k-row table and find the seq scan that's costing you.

5 steps
Chapter 5

Add the index

Create a composite index on the same 300k-row table and watch the plan — and the timing — change.

3 steps
Chapter 6

Read the brief

One stakeholder brief, no column list, no stated metric — find the right tables across the whole warehouse and defend your number.

1 step + stakeholder brief