Ch.3: Your BigQuery Query Is Waiting for Compute
Outline
- 0:00 Bytes are not the whole story
- 0:45 What a slot is not
- 1:19 What a slot actually is
- 1:51 One query, many workers
- 2:42 Slot time vs wall time
- 3:32 The shared pool
- 4:15 On-demand still uses slots
- 4:59 Reservations and capacity
- 5:37 Fair scheduling
- 6:07 Slot-bound queries
- 6:53 I/O-bound queries
- 7:42 Shuffle-bound queries
- 8:20 Reading execution details
- 9:06 Choosing a pricing posture
- 9:38 Slots explain compute
Transcript
0:00 Imagine two BigQuery queries that scan almost the same number of bytes. One comes back in about 4 seconds. The other crawls for 2 minutes while the whole team starts asking who launched the monster job. Oh, man, that one. Same bytes, wildly different wait, and the billing console still swears they cost the same, because it only counts bytes. So bytes are clearly not the whole story. The missing half is the slot: the compute currency BigQuery actually spends to run your query. And that currency is the lens for the whole episode.
0:35 Right. Bytes explain what you read off disk. Slots explain how a query actually spends compute, and whether the capacity was free the moment it needed it. So before we say what a slot is, it helps to kill the wrong pictures first. A slot is not a virtual machine you log into. It is not a CPU core you pin. It is not a fixed worker that belongs to your query forever. Oh right, the hidden-server one. So drop that mental image of a box with your name on it. And drop the warehouse you resize by hand, too. Picture either of those and, you know, BigQuery's behavior feels random, when it really is not. True.
1:13 A more abstract picture actually survives every hardware change underneath. So if it is none of those things, what is it? A slot is BigQuery's virtual compute unit. It bundles CPU, memory, and network together, then hands that bundle to a piece of query work for as long as the work needs it. Right, so it is rented compute, and the docs linked in the description spell out exactly what a slot bundles. You never touch the actual hardware. Yes. Google can swap the whole fleet overnight and nobody relearns any hardware math. And that is the point: a slot is currency, not a machine.
1:51 Now take that definition and watch a real query run, because that is where slots start moving. A single query almost never runs as one serial task. Picture slots as a crew of temp workers you rent by the second. The scan splits into pieces they read in parallel, aggregations happen in stages, and joins and sorts move data between those stages. Sure. So you are hiring more of that crew whenever the plan has parallel work to hand out, and paying them all to stand around when it doesn't. True, and the wider that fan-out goes, the more it grabs. A stage splitting into thousands of parallel inputs pulls a lot of slots right there. But a narrow stage, or one stuck waiting on the one before it, gets nothing from extra capacity.
2:37 Yeah. So extra slots are dead weight unless the plan can hand out parallel work. Then that fan-out is also why we measure slot time, not just a slot count. Slot time is compute multiplied by how long you hold it. Let's say a query grabs a thousand slots for 10 seconds. That is a big burst of parallel compute packed into a short window. Can two queries really burn the same slot time and still feel completely different to the user? They can. The slot time comes out roughly the same, but the person staring at the screen waits minutes instead of seconds in the second one. Woah. So slot time tells you how much compute happened, and wall time tells you how long the human actually waited.
3:22 Honestly, it took me years to really feel how separate those two axes are. I always thought tuning one would move the other, and it just does not. Now, so far the query has been alone in the room. In production it almost never is. For example, a customer-facing dashboard query times out at 9 in the morning because the nightly job is still, Holding most of the slots, while three analysts and somebody's forgotten experiment named final underscore two all hit refresh. Sure. Same SQL, busier pool, slower result.
3:54 That is the unsettling part: nothing in the query changed. Yeah, and you know the feeling: the query that was fast all week suddenly isn't, and nobody touched the SQL. That one breaks people who trust local benchmarks, where the machine is theirs and the numbers repeat. A slow run might just be a crowded pool, not a bad query. So all that contention shows up differently depending on how you pay. Start with on-demand. On-demand bills you on bytes processed, which is the number most people notice first.
4:26 Right, but underneath, the query is still spending slots. The meter just hides them behind the bytes number. So you're renting by the query, yeah. On-demand does not mean slot-free. It means you did not buy a fixed slice of capacity for this workload. So the tradeoff is simplicity for control. For bursty or exploratory work the simple model is great, and you tune bytes: prune partitions, select fewer columns, stop wildcarding tables. And that is the moment performance isolation stops being a detail you can ignore.
4:59 That leads to the alternative: reservations make capacity explicit. Okay, so what changes day to day? You buy slot capacity through reservations and, I mean, hand it to specific projects. Take a company where production dashboards and somebody's ad hoc experiment share one pool. A reservation splits them so the experiment cannot starve the dashboard. Oh, so it is really a scheduling decision dressed up as a billing one. The nightly job, the analyst notebooks, and the customer dashboard each get their own lane.
5:32 Yes. One team needs steady throughput; another needs bursts. Reservations let both work under one capacity plan. Now, once you have reservations, something still has to decide who gets the slots in any given second. And BigQuery does that live. Capacity gets reallocated as jobs arrive, finish, or change shape mid-flight. Hold on. It can pull slots away from a job that is already running? It can shift free capacity the instant another job releases it, so a query can suddenly speed up halfway through.
6:02 Every slot stays busy, the sharing stays fair, and no single query owns the machine room. So with that scheduling picture, the performance symptoms finally pull apart. The first one is slot-bound. Slot-bound sounds like this. Imagine the plan is ready to fan out to a thousand workers, but only two hundred slots are free, so most of the work just sits in line. Uh-huh, and that is not the same as reading too many bytes. Not at all. I mean, adding a filter helps only if it cuts the work. If the real bottleneck is a crowded pool or a stage starved of parallelism, the filter misses entirely. Sure.
6:41 I watched a teammate lose an afternoon convinced his SQL was broken, when the real problem was that nothing else would let go of its slots. The result is you reach for timing or isolation, not "scan less." Now, slot-bound is about capacity. The second symptom sits at the opposite end of the pipe, which is I/O. This is where the old advice still earns its keep. Select fewer columns. Use partition filters. For instance, do not wildcard a year of tables to answer a question about yesterday. Okay, but we have all sat in the meeting where a staff eng just says buy more slots.
7:17 Why is that the wrong move here? Hmm, fair, and I think that instinct is exactly the trap. The slot model sits next to bytes-scanned thinking, it never replaces it. If the query spends all its time feeding workers input it never needed, extra capacity just helps you waste money faster. Which is a genuinely expensive way to be wrong, a reminder that more capacity never answers an I/O problem. Now, bytes read is one axis. The 3rd symptom hides on a completely different one, which is data movement.
7:50 Movement, meaning the data physically relocates between stages. Consider a join that explodes 10 million rows before any filter trims them. Matching keys, grouped values, sorted ranges, all of it has to move so the right rows land together. Oh, wow. So the bytes look fine, but the whole apartment is getting moved just to find one receipt. Exactly. That is the train wreck under the hood. Joins, groups, windows, and sorts create the movement. When movement dominates, reshape the query instead of adding capacity.
8:21 So all three of those symptoms leave fingerprints in the same place, the execution details. So which fields actually tell you which problem you have? More than one, actually. Look at the stages that, Dominate slot time. Look at pending, active, and completed units, and whether work is stuck waiting. Then check shuffle output, spill, and input size. So, I mean, if one stage is eating most of the compute while the rest sit idle, that is the tell? Yes, and the query-plan doc we linked names every one of those fields.
8:54 The job is to classify the pain. Is this mostly I/O, mostly shuffle, mostly a crowded pool, or just a query that will not parallelize the way you hoped? Name which one it is, and now you know which lever to pull. Now, reading those details is the skill. Choosing a pricing posture is the decision that sits on top of it. And it is not about which model is morally better. It is which one fits the workload. On-demand fits when usage is exploratory, spiky, or honestly still unknown. A finance partner would put it bluntly: if you cannot forecast it, do not commit to it.
9:31 Capacity fits predictable workloads, isolation needs, dashboard reliability, or finance planning, but bad SQL still matters because pricing only changes who controls the compute. So pull it together. Slots are not some optional footnote to BigQuery performance. They are the missing compute lens. Bytes scanned tell one part of the story. Slot usage, slot time, waiting, and shuffle tell the execution half. And a slot was never a server you babysit. It is the currency BigQuery spends on your query work, shared across stages, jobs, and whole teams.
10:04 So once you see it that way, slow stops being mysterious. It means too much input, too much movement, too little free capacity, or a plan that cannot parallelize. Coming next, we look at Capacitor, the columnar format that makes BigQuery cheap when you ask for the right columns and brutal when you ask for everything. Thanks for listening to Learning Podcasts.