System Design Ch.3: Caching Patterns
Outline
- 0:00 Fast Per Request, Expensive in Aggregate
- 0:35 One Expired Cache Key
- 1:40 Two Versions of Reality
- 2:04 The Staleness Budget
- 3:15 Cache-Aside
- 3:55 The Race Condition
- 4:35 Write-Through
- 5:36 Write-Behind
- 7:14 TTL Is a Business Rule
- 7:50 The Thundering Herd
- 8:52 Jitter
- 9:31 Let One Request Pay the Miss
- 10:17 CDN and Personalization
- 10:56 Memcached vs Redis
- 11:20 The Miss Path Is the Real Design
- 11:50 Next: Message Queues
Transcript
0:00 Your database query takes exactly eight milliseconds. You look at your metrics, and I mean, that sounds incredibly fast, right? Oh yeah, it feels totally cheap, almost free, honestly. Right, but then a million users asked that exact same question today, and suddenly you are spending hours of expensive database compute time just, you know, repeating work your system already did. And here is the real paradox we're exploring today. Sometimes the massive catastrophic incident that brings your whole system down Yeah.
0:30 It isn't caused by a bad query plan or like a missing database index. No, it's caused by a single expired cache key. Just one key that suddenly sends 10,000 identical reads back to the database at the exact same millisecond. Which is just terrifying. Yeah. It's the defining moment when caching stops looking like this clever little micro -optimization, you know, and it reveals itself as a load-bearing pillar of your architecture. Because the absolute fastest request your system will ever process is not the one with the most highly optimized database index.
1:00 It's the one that never reaches the database at all. Exactly. Completely bypassed. Because, at its core, caching is just the ultimate exercise in intentionally managing stale reads. Right. The structural change. Yeah. Think of it like keeping your most frequently used tools sitting right there on your workbench, instead of walking all the way back to the main tool room every single time you need a wrench. I love that analogy. Because the speed gain doesn't come from having a technologically superior wrench.
1:27 No. The speed comes entirely from avoiding the repeated trip. But the thing is, the moment you set up that workbench, the exact second you store a computed result somewhere cheaper and closer to the read path, you fundamentally fracture your system's reality. Oh, that's a good way to put it. You fracture reality. You do. Because you now have two distinct versions of the truth. You have the authoritative source of truth sitting in your durable database, and you have the easily accessible cached copy sitting on the workbench.
1:56 And from that exact moment forward, caching is no longer a latency problem. It is a staleness problem. Exactly. But you're saying the real design question we should be asking is actually, how wrong is this value allowed to be, and for how long? That's feisty. Because data inherently changes over time. I mean, if the data in your database never mutated, caching would be trivial. You'd just load it once and never think about it again. Exactly. But in the real world, the authoritative source of truth gets updated.
2:27 And the instant that database commit happens, your cached copy is officially a lie. A highly efficient, rapidly served lie. Yeah. So as a system designer, you have to look at every single piece of data and ask a product question. Is it acceptable for my system to lie to the user about this specific value? And if it is, for how long? 10 seconds. 10 minutes. Right. Or is this data so financially or operationally critical that it can never be wrong, not even for a single millisecond? That tolerance dictates literally every engineering choice that follows.
3:02 Okay. So if we've intentionally fractured our reality into two copies, the immediate structural problem becomes synchronization. Like how do we keep the workbench updated when the tool room changes? That's the core mechanical challenge. And the most standard approach I see in all these system design manuals is the cache -aside pattern. Oh, absolutely. The classic. Yeah. The application code checks the cache first. If it's a miss, it goes to the database, reads the value, puts it into the cache, and then returns it to the user.
3:29 But structurally, looking at how distributed systems operate under load, doesn't that leave a massive window for race conditions? Oh, it leaves a gaping window if you aren't careful. Which is why understanding the mechanics of cache-aside is so critical here. But the trade-off is that the application code itself now completely owns the responsibility for data freshness. Which sounds risky. Let's look at the race condition I mentioned. Sure. Imagine Thread A experiences a cache miss. It reads a value of, say, 10 from the database.
4:01 But then it experiences a momentary garbage collection pause. Okay. Happens all the time. Right. Meanwhile, Thread B updates that database value to 20 and obediently invalidates the cache key, just like it's supposed to. Wait, so the cache is empty now? Yep. Then Thread A wakes up from its pause and writes its stale value of 10 into the cache. Oh, wow. So your cache is now permanently out of sync until the entry's natural expiration. Exactly. It's a permanent lie. That is terrifying, because it turns every single code path that touches data into a high-stakes data freshness decision.
4:34 It really does. So if application-owned invalidation is that prone to race conditions and human error, I mean, I'm looking at the write-through pattern as a potential fix. Right. Just wire them together. Yeah. The application writes new data, and it updates the cache immediately in the same synchronous operation as the database update. It eliminates the race condition entirely, but I, I'm guessing the hidden tax is write latency. The write latency combined with, like, potentially massive cache churn.
5:02 Because every write hits the cache. Exactly. With write-through, you move the cache directly into the critical write path. Yeah. That's the upside. But the penalty is steep. Every single database right now pays an extra network hop and operational cost to update the cache. So writes get sluggish. Very sluggish. And crucially, if you have a high write, low read system, you end up constantly storing and updating data in memory that nobody will ever actually request. Right. You're paying a premium latency cost for a read benefit you aren't actually utilizing.
5:34 Exactly. It's just wasted effort. Which naturally leads me to wonder about write-behind caching. Because on paper, it looks like a silver bullet for that exact latency problem. Oh, it looks incredible on paper. Right. The application writes the update to the cache in memory, immediately acknowledges success back to the user, and then an asynchronous process flushes those changes to the actual database in the background. Making the write path lightning fast. Exactly. But what am I missing? Where is the hidden tax there?
6:02 The hidden tax is the catastrophic risk of silent data loss. Ah. Yeah. Write-behind makes writes incredibly fast, and it gives you the ability to beautifully batch your database updates, which really reduces load. But it is playing with fire for anything resembling critical state. Because it's volatile memory. Let's trace the mechanics. The system accepts the write, stores it in the volatile cache, and tells the user, success, your data is saved. But it's not really saved. No. If that cache node loses power or crashes before the asynchronous flush to the durable database actually executes, that data is gone forever.
6:39 Wow. So the system literally lied to the user about durability. Completely lied. Write-behind is a powerful tool, but strictly for data, where losing a small rolling window of recent writes is an acceptable business risk. Like video view counters or maybe buffered analytics streams? Exactly. But it is an absolute disaster for processing payments or updating user passwords. Okay. So we have these intricate paths for getting data into the cache, but eventually we have to follow the failure path. Because eventually every cache item has to die.
7:12 Right. It stops being trustworthy. Which brings us to the time to live, the TTL. And based on what we just discussed about staleness, a TTL isn't really a technical setting at all. It's a business rule wrapped in an integer. That is exactly what it is. The TTL is the blunt instrument we use to enforce the staleness budget we defined earlier. Right. So a user's uploaded avatar image can probably be cached with a TTL of 12 hours. Sure. Because a slight delay in a profile picture updating is a trivial user experience issue.
7:41 Nobody cares. But their financial account balance or like flight seat availability, the TTL on those might need to be zero. Zero. Absolutely. Yeah. Exactly. That hot key reaches its TTL and expires. In the exact millisecond that key drops out of memory, you might have 10,000 concurrent in-flight requests all checking the cache for that video. And every single one of them gets a cache miss at the exact same moment. Yep. The thundering herd. The thundering herd. All 10,000 requests simultaneously pivot and stampede the database.
8:16 Oh, man. They're all desperately running the exact same expensive query to compute and rebuild the exact same value. So the cache was supposed to act as your shield. But the expiration event effectively synchronized all the misses. Exactly. It manufactured a massive, highly coordinated traffic spike that your database connection pool was never sized to handle. The database locks up, the queries time out, and the whole system collapses. So mechanically, how do we mitigate the thundering herd? Because just making the database bigger isn't sustainable.
8:47 No, you can't outscale that. There are a few primary mechanical patterns. The first is adding randomness or jitter to your TTLs. Like varying the expiration times. Exactly. If you load a massive batch of keys into the cache during a startup sequence, you do not want them all expiring at exactly midnight. Right. That would be a huge spike. So you add a random variance of maybe plus or minus five minutes to each key's TTL. That way they expire organically over a smeared window of time rather than a single violent drop.
9:17 Okay, but jitter doesn't solve the problem for a single massively popular hotkey dropping. No, it doesn't. Because if the viral video key expires, jitter doesn't stop the 10,000 requests that hit in the next millisecond. Correct. For a single hotkey, you need a locking mechanism, often a distributed mutex lock in the cache itself. How does that work? When the key expires, the very first request that gets a miss acquires a lock. It is granted exclusive permission to go to the database, run the heavy query, and rebuild the cache.
9:49 And what about the other 9,999 requests? They hit the lock and are forced to wait for a few milliseconds, just pulling the cache until the new value is populated. Ah, so only one does the work. Exactly. So the overarching rule of system design here is basically, never let 10,000 requests all pay the heavy database compute cost for the exact same mess. That is the golden rule. Never do it. Okay, so far we've kept this strictly inside the application architecture. But pushing the staleness problem closer to the user inherently breaks application level controls, right?
10:22 Which forces us to shift to edge caching, CDNs. Content delivery networks, yes. But for a senior engineer, the real complexity of a CDN isn't the geography. It's the invalidation mechanics and key fragmentation. That is where the design actually gets difficult. The moment you push dynamic or semi-dynamic data to the edge, you are stripping away the application's ability to easily reach in and clear a specific user's stale state. I see teams spend weeks arguing over which technology to deploy. It means the debate is almost entirely secondary.
10:58 I mean, Memcached is brilliant in its simplicity. A pure, multi-threaded, in-memory, key-value cache. And Redis is much broader. Right. It's single-threaded but offers rich data structures, sorted sets, pub-sub, and optional disk durability. Which is why it solves so many adjacent architectural problems. So, let's synthesize this. Caching is not a magic trick to make poorly written SQL queries run fast. It is a strict, structural trade-off framework where you deliberately introduce a second, inherently stale copy of your data simply to eliminate repeated work.
11:34 Yeah. A cache hit is a wonderful optimization. But the miss path is where the actual system design happens. Because when the cache inevitably disappears or a key expires, your underlying database still has to survive the impact. It all comes down to the miss path. So, in Chapter 4, we are going to pivot our focus. We are moving away from speeding up repeated synchronous reads. And we are going to dive into how to decouple work entirely.