System Design Ch.1: Framework Over Tools
Outline
- 0:00 Introduction
- 0:20 The Tech-First Trap
- 2:18 The Payload Analogy
- 2:55 The 4-Stage Framework
- 3:57 Step 1: Functional Requirements
- 4:38 Step 2: Non-Functional Requirements
- 5:48 Step 3: Capacity Estimation
- 7:18 Step 4: The Write Path (Base62)
- 8:02 The Read Path (Why Caching Is Mandatory)
- 8:26 Async Analytics (Never Block the Redirect)
- 9:14 The Revelation: Every Decision Has a Reason
- 9:58 What's Coming in the Series
Transcript
0:00 Picture this. You're staring at a grid of faces on a Zoom call, or maybe you're sitting in a conference room with a whiteboard that hasn't been erased since, I don't know, last Tuesday. Oh yeah, the one with the permanent marker stains. Exactly. And the product manager leans forward and says, our team, we need to build a URL shortener. Here we go. Right. Instantly, before a single person in that room has asked how many requests per second this thing actually needs to handle, the entire meeting devolves. I mean, someone is aggressively defending Redis for the caching layer.
0:32 Another engineer is arguing that, you know, Memcache is more lightweight. And the junior developer is quietly asking if the whole thing can just be rewritten in Rust. Yes. We have all lived through that exact meeting. And honestly, starting with the technology stack instead of the system requirements is the single most common and almost always the most fatal mistake in backend engineering. It really is. And the irony is the more experienced the team is with specific technologies, the faster they tend to fall into that trap.
1:01 Oh, totally. Because an engineer hears a problem and their brain immediately like pattern matches it to a tool they already know and love. They just bypass the actual design phase completely. Well, welcome to this deep dive. Today, we are embarking on something massive. This is chapter one of a definitive 18-part series focusing entirely on system design for backend engineers. It's going to be quite the journey. It really is. And because we know who is listening, we are making a major assumption about you right out of the gate.
1:30 You already know your stuff. I mean, you know Python, you write Java, you can navigate GCP and AWS blindfolded. You probably dream in Terraform configurations at this point. Right, exactly. If we asked you to spin up a Spark cluster or trigger an Airflow DAG, you wouldn't break a sweat. So we are not here to teach you how to set up a database. The mission today is to unpack the structure underlying reasoning process for designing distributed systems entirely from scratch. Because, you know, knowing how to configure Kafka is essentially useless if you don't know the underlying mathematical reason why you're implementing it in the first place.
2:06 Or when you shouldn't touch it at all. Exactly. Technology selection is the absolute final step of system design. It should never be the first. Let's frame this with an analogy. Because picking your tech stack first is like, well, it's like trying to put Formula One racing tires on a dump truck. That's a great way to put it. Right. You're arguing over the brand of the rubber before you even know if you're hauling 10 tons of gravel across a dirt quarry. Or, you know, trying to break a land speed record on asphalt.
2:34 The tires you need are entirely dictated by the chassis. And the chassis is dictated by the payload. And the payload is the requirement. Right. Engineers who build systems that actually survive the brutal reality of production, they don't start with a shopping list of tools. They rely on a very rigid, sequential framework. Okay. So break that framework down for us. What does it look like? So first they start by defining the requirements. Then they use those requirements to derive mathematical constraints.
3:03 After that, they sketch a design that satisfies those specific constraints. And then, only at the very end of the process, they choose the technology that fits the sketch. Exactly. So let's put that to the test. Let's prove why the tech first approach fails by looking at what happens when you do it right. Right. We'll stick with that URL shortener example, but let's apply it to an incredible scale. Let's look at Bitly. Oh, Bitly. Yeah, their scale is insane. They're processing over 10 billion link redirects every single month.
3:32 10 billion. That is just, that's a staggering baseline. It really is. When you break that down into daily operations, I mean, you're looking at roughly 4,000 requests hitting the routing servers every single second, constantly. 4,000 a second. And every single one of those redirects needs to resolve in milliseconds. Okay. So instead of jumping to Redis or Rust, let's walk through the framework. First up, defining the functional requirements. What does the system actually do? And the secret here is keeping the scope incredibly tight, right?
4:03 Absolutely. For a URL shortener, the functional scope is just three actions. One, the system must generate a shortcode from a long URL. Okay. Simple enough. Two, it must redirect a user from that shortcode back to the original long URL. And three, it optionally tracks clicks for user analytics. And that's it. That's it. If an action doesn't fit into one of those three buckets, it is completely out of scope. Right. So once we establish those basic rules, the next step is defining the non-functional requirements.
4:33 This is where the actual engineering starts to take shape, I think. This is where we define the physics of the environment our system has to survive in. Yeah. We're defining the constraints. So for a URL shortener, the system is going to be incredibly read heavy. The primary non-functional requirement is roughly a 100 to 1 read to write ratio. Because people are clicking links vast magnitudes more often than they're actually logging in to create new ones. Exactly. Another constraint. The redirects must be fast.
5:02 Let's say under 100 milliseconds. Makes sense. Nobody wants to wait for a link to load. Right. And then there is high availability. Think about it. If a URL shortener goes down, every single link it has ever generated across the entire internet is suddenly a broken page. Oh man, yeah. The blast radius of downtime isn't localized to a single app. It literally breaks the internet for your users. Exactly. So it has to be highly available. And finally, it must scale horizontally to handle unpredictable viral spikes in traffic.
5:31 Okay. So we know what it does and we know the constraints. Now we transition to capacity estimation. And I know engineers who absolutely dread this phase. Oh, they hate it. It feels like making blind guesses in a vacuum. Right. But let's actually run the numbers. Let's assume we have 100 million new URLs being generated per month. Okay. So to break that down without getting lost in the weeds, you just use the standard engineering mental cheat code. Which is? There are roughly 2.5 million seconds in a typical 30-day month.
6:01 So you divide your 100 million new URLs by 2.5 million seconds and boom, you are looking at 40 writes per second. Wait, 40 writes? Suddenly this isn't a terrifying, unimaginable data problem. I mean, 40 writes a second is a completely manageable trickle. It is. But then we apply that non-functional constraint we defined earlier, the 100 to 1 read ratio. Ah. So you take those 40 writes, multiply by 100, and you get 4,000 reads or redirects per second. Exactly. Now let's calculate storage. If each database record, which includes the long URL, the short code, and some basic metadata, is about, let's say, 500 bytes.
6:41 Then we're generating 50 gigabytes of new data every month. Right. And the underlying goal of this estimation phase is not to calculate your storage needs down to the exact byte. It's about finding your order of magnitude. Right. Like, are you building an architecture for hundreds of requests per second or hundreds of thousands? Because the infrastructure required to store 50 gigabytes of data a month is completely different from the infrastructure needed for 50 terabytes a month. Okay. So with the math firmly in hand, we can finally draw the system.
7:11 We move to the high-level design. Let's look at the right path first. We need to generate unique short codes. And I think the standard approach is base 62 encoding, right? Yeah. Base 62. The mechanism behind it is elegant because it just uses the standard alphanumeric alphabet. You have A through Z uppercase, A through Z lowercase, and digits 0 through 9. Which totals 62 distinct characters. Exactly. And if you mandate that every short code is exactly 7 characters long, the math is 62 to the power of 7.
7:43 That yields over 3.5 trillion possible combinations. 3.5 trillion, which means we won't run out of unique URLs for decades, even at 100 million new links a month. That's awesome. So that's the right path. Now, the read path. We know we have 4,000 reads per second, and we have a strict sub-100 millisecond latency limit. Right. This is our hot path. And sitting a database naked in front of 4,000 reads a second and hoping it responds in under 100 milliseconds is, well, it's a gamble you will lose. Definitely.
8:15 That read-to-write ratio mathematically dictates that a caching layer must sit in front of the database. And lastly, the analytics pass, the click tracking. Right. So if you process click analytics synchronously, meaning the user clicks the link, the server writes the tracking data to a database, and then finally redirects the user, you will instantly violate your latency constraint. Because the database write will block the redirect. Exactly. Therefore, the analytics pipeline must be asynchronous.
8:42 You drop a message in a queue, redirect the user immediately, and let a background worker process the click data whenever it has capacity. See, this is the massive aha moment for me. When you walk through the framework, you realize the cache wasn't a technology preference. We didn't add a cache just because someone in the meeting really liked Redis. Oh, not at all. The cache was forced into existence by the mathematical reality of a 100 to 1 read ratio. And the asynchronous analytics pipeline wasn't like a trendy architectural choice.
9:13 It was physically demanded by the sub-100 millisecond latency constraint. Exactly. Every single architectural decision has a mathematical or logical reason you can point to on a whiteboard. The constraints dictate the design. The framework just completely strips away personal bias and forces you to justify every line you draw on the architecture diagram. Well, this first chapter has laid down the foundational skeleton. We've gone from the chaos of starting with a tech stack to the structured discipline of defining functional requirements, establishing non-functional constraints, running the capacity math, and letting those numbers draw the high-level design.
9:49 And this framework is just the beginning. I mean, the systems we interact with every day are infinitely more complex than a URL shortener. But the reasoning process remains exactly the same. Over the next 17 chapters, we are going to escalate the complexity drastically. We aren't going to read you a laundry list. But just to give you a sense of where we are heading, we're going to start by tackling what happens when your data gets too big for one machine. Oh, yeah. We'll dive into replication, the CIP theorem, and what happens when your databases start disagreeing with each other.
10:23 From there, we scale up to real-time problems. We'll explore the architecture behind massive chat systems using WebSockets. And we'll dissect the dreaded fan-out trade-off that makes or breaks social media news feeds. Not to mention caching strategies, cache invalidation, the thundering herd problem. Message queues, Kafka versus traditional queues, exactly once delivery. Right. And eventually, we are going to use this exact same four-stage framework to design colossal infrastructure. We're talking distributed job schedulers, multi-region databases, and we cap the entire series off by architecting a full -scale search engine from scratch.
11:00 These are the exact whiteboard challenges that separate mid-level developers from truly senior back-end engineers. Every single one of those massive systems, no matter how intimidating it appears from the outside, can be completely demystified by applying the constraints-driven framework we discussed today. Absolutely. Which brings us back to the core thesis of this deep dive. Requirements dictate constraints. Constraints dictate the design, and the design dictates the technology. Never, ever start a meeting by picking the tech stack.
11:29 Let the math and the physics of your system make the decisions for you. If you internalize that process, you won't just build systems that look elegant on a whiteboard. You will build resilient systems that survive the chaos of reality. Well said. Thank you for joining us on this deep dive. We will see you in Chapter 2.