System Design Ch.8: Notification Systems at Scale

Outline

Transcript

0:00 Picture this. It is, uh, it's two in the morning. Oh man, the absolute worst time for things to break. Right. You are completely locked out of your account. You're exhausted. And you just desperately need a password reset. So you click the button and you wait. And every single minute the ticks by feels like your account is just permanently broken. Exactly. But behind the scenes, your urgent reset email isn't lost. It's actually trapped in a massive queue sitting directly behind a batch of like 2 million weekly marketing digests.

0:30 Which is an absolute nightmare scenario for any platform. I mean, when you mix those streams, you're basically signaling to the user that their account security is exactly as important as a promotional coupon. It's a total failure of system design. And that is why in today's deep dive, we are tackling chapter eight of the system design for backend engineers series. We're talking about designing a notification system. Yeah. And the core thesis here, the thing you really need to keep in mind, is that a notification system is not just a dumb pipe for sending bytes to a provider.

1:01 Right. If you just need to send an email, you call an API. That's easy. Exactly. The actual system design challenge is building a decision system. It's deciding what deserves to be sent, how urgently it needs to go out, through which channel, and crucially, when silence is the correct result. Because user attention is your most scarce resource. Every single time you interrupt the user, you're spending their trust. Spot on. And that means we need to explicitly draw a boundary between notifications and chat.

1:30 We covered chat systems back in chapter seven, right? Chat is live, two-way communication between active users. Yeah. Whereas notifications are interruptions. Precisely. They tap you on the shoulder when you're not looking. So to set the stage for you listening, I want to quickly preview the full architectural map you're going to build today. Just visualize these major regions in your head. First, you have the intent intake. Then the policy gate. Next is the template renderer, followed by priority lanes, channel workers, external providers, and finally, the send ledger.

1:59 We're going to unpack this map piece by piece. And we promise we will return to this exact diagram at the end once all these pieces actually make sense together. Let's start at the very beginning of the pipeline, which is intent ingestion. We need a strict boundary between product services and the delivery platform, right? Yeah, absolutely. Yeah. Your login service shouldn't know how to talk to APNS for Apple or FCM for Android or your email provider. It just shouldn't care. It's not its job. Right. Instead, it emits what we call a notification intent.

2:32 Like a structured fact. Something like, um, password reset requested for user 123. Precisely. This intent includes a stable event ID, a type, a priority level, the raw template data, and this is critical, a deduplication key. And then the notification service takes that intent and owns the entire rest of the life cycle. Yep. But the deduplication key is where things get really interesting. Well, let me push back on that a bit because using a time window for deduplication sounds a bit dangerous. How so?

3:00 If you're holding a notification for a minute to group social likes, you're introducing state into an otherwise stateless ingestion pipeline. Doesn't that create a massive memory overhead during a viral event? Oh, for sure. Like if a celebrity posts a photo and you get a million likes in 60 seconds, buffering all those payloads in memory just to deduplicate them sounds like a fast track to an out of memory crash. That is a very real danger, which is exactly why you never buffer the full payload in memory for aggregation.

3:29 Ah. Okay. So how do you handle it at scale? The actual implementation relies on a fast, distributed key value store, typically Redis. When the intent arrives, you store the heavy payload in blob storage or maybe a temporary database table. Okay. That makes sense. Then you generate a strict deduplication key. This combines the user ID, the notification type, the business object ID, and the time window. And then you use a Lua script in Redis. Exactly. You use a Lua script to atomically increment a counter or append the new event ID to a set. And you bind that by a time to live.

4:02 So the platform is really just maintaining pointers during the time window, not the actual heavy data. You got it. When the window closes, a separate worker gathers the pointers, fetches the payloads from storage, and collapses them into a single aggregate intent. Like, user 1, 2, 3, 10,000 people liked your post. That's a much better user experience anyway. Yeah. And this idempotency model is non-negotiable. In distributed systems, producers are going to retry. Message queues will redeliver. Workers will crash.

4:33 Right. Without a strict idempotency boundary, normal system failures turn into duplicate spam. You click reset once, the worker crashes twice, you get three emails, it completely ruins trust. Okay. So the platform safely ingests the intent. It's deduplicated. The immediate next step in our map is the policy gate. Right. This is where we decide if the message should actually go anywhere at all. But wait, if every single intent has to join against a user preference table to check quiet hours or opt-outs, you've just created the biggest single point of failure in your architecture.

5:04 Oh yeah. A huge bottleneck. How do you scale that without the database melting down? You decouple the read path from the write path. User preferences are a heavily read-biased data set. You do not do a relational join for every single notification. So you flatten it. Exactly. You flatten the user's routing preferences into a document and cache it globally. You might use Redis read replicas distributed across your availability zones. Or maybe localized caching within the worker nodes. Yeah. Using a PubSub model to invalidate stale preferences.

5:36 Which means the policy gate can evaluate rules in memory. And that is critical because user preferences are not just some UI settings page. No. In a proper architecture, they are a core correctness constraint of the entire system. Right. If a user has quiet hours enabled, the policy gate makes the decision. It holds the push notification until morning, downgrades it to an email, or suppresses it entirely. And, you know, the complexity usually arises when product managers want to bypass those preferences.

6:07 Oh, I've had that exact argument. Right. They argue their notification is a critical transactional alert, like an account overdraft, so it needs to bypass quiet hours. And while some transactional messages are genuine exceptions, if you allow every product team to flag their specific feature as critical, the designation just becomes meaningless. It's the boy who cried wolf at scale. Exactly. The marketing team decides their weekend promo is quote unquote critical, the user's phone buzzes at 3am, they get furious, and they revoke all notification permissions at the OS level.

6:38 Yep. Game over. So centralizing this policy gate protects the user from your own internal organizational chart. You're enforcing a unified attention budget. Well said. So once an intent survives the policy gate, meaning it's valid, deduplicated, and allowed to send it, we have to ensure our 2am password reset doesn't get stuck behind that massive backlog of marketing digests. Which brings us to priority lanes. Are we talking about physical isolation here? Yes. Like distinct Kafka topics for each lane backed by isolated Kubernetes pods for the workers?

7:10 Because theoretically you could just use a single priority queue data structure. You could. But my instinct, and I think yours too, is that shared infrastructure means shared failure modes. Yeah. A memory leak in the bulk worker could take down the emergency worker. Exactly. Your instinct is spot on. A single undifferentiated queue is easy to draw on a whiteboard, but an absolute nightmare to operate. We physically separate the work into different lanes with completely separate worker budgets and autoscaling rules.

7:38 I always think of this like a highway system. You build an emergency lane for account recovery and security alerts. You have a normal traffic lane for social activity and comment replies. And you have a bulk freight lane for weekly digests and massive marketing blasts. Right. By decoupling these at the infrastructure level, operational isolation is practically guaranteed. If a marketing campaign drops 2 million pending emails into the bulk freight lane, it only saturates the Kafka partitions and worker pods assigned to bulk freight.

8:07 Your emergency lane remains completely unaffected. You can aggressively throttle or even pause the bulk lane without ever touching the security alerts. But what about starvation? Most of the time, that emergency lane is going to be nearly empty. Are those dedicated workers just sitting around idle? Yes. And you gladly pay for that idle compute. In an emergency lane, you are optimizing for predictable latency, not maximum throughput. You want them ready to go. You want those workers sitting there doing absolutely nothing so they can process a password reset the exact millisecond it arrives.

8:38 Okay. So our urgent reset is flying down the emergency lane. Now it needs to be transformed from raw facts into a specific format for the user's device. Which brings us to the template renderer. We centralize this for consistency and localization, right? Instead of the login service sending a massive string of English text, it sends a translation key and the variables. And then the platform handles translating that into 40 different languages. Makes sense. The renderer also shapes the message for the distinct personalities of each channel.

9:08 Mobile push, which goes through platform services like APNS or FCM, is cheap and fast. But it's a non-guaranteed signal. Right. Apple explicitly frames remote notifications this way. Push is not a database. Email, conversely, is slower, durable from the user's perspective, and highly searchable. It's where you put receipts and long, detailed explanations. Which brings up a massive privacy boundary. Oh, yeah. Because push, email, and SMS move through external third-party networks, your payload shouldn't casually contain sensitive data.

9:44 Never. If it's a security alert, the message shouldn't say, someone logged in from this IP address using this specific device model. It should be a deep link into the secure app that just says, review recent login activity. Exactly. The notification is not the durable container of sensitive state. It is just the tap on the shoulder. You're spending a fraction of the user's attention to pull them back into an environment you control. And that calculation changes drastically when you look at SMS. Oh, SMS is a whole different beast.

10:12 It is immediate, highly intrusive, expensive, and heavily regulated. Which perfectly transitions us to the next stage of the map. The channel workers. Right. The message is rendered, and now we have to actually interact with the outside world. For external providers, my initial instinct might be to just wrap the API call in a standard retry loop and call it a day. But at this scale, treating all failures the same feels like a recipe for a thundering herd. Yeah, that will get us rate-limited or outright banned by the provider.

10:42 Repping everything in a blind retry loop is super dangerous. A clean design uses distinct channel adapters that translate our internal platform model into the specific API of the provider. And within those adapters, we have to distinguish between transient provider timeouts and permanent errors. Exactly. Assuming you understand the mechanics of exponential backoff and jitter for the transient stuff, how do we handle the permanent failures? A permanent error, like an invalid device token from APNS or a hard bounce from an email server, should be dead-lettered immediately.

11:15 You stop trying. Just cut it off. Right. More importantly, you emit an event back to the user profile service to prune that contact method. If you keep hammering a provider with invalid tokens, they will throttle your entire account. And this naturally brings up fallbacks. I've seen systems where SMS is treated as this sort of universal panic button, like the email bounced, shoot them a text. That is massive anti-pattern. Fallbacks are product decisions, not automatic infrastructure mechanisms. So SMS is not a universal fallback.

11:45 No, absolutely not. It spends vastly more user attention and company money. Escaping to a higher friction channel should only be reserved for urgent, high-value flows. Like account recovery or two-factor authentication. Right. Sending a weekly product digest via SMS just because an email bounced is a great way to lose a user forever. So how do we track all of this? How do we definitively know if we actually sent the password reset or if it failed or if it was suppressed by quiet hours? This is where the send ledger comes in.

12:16 The best mental model for the send ledger is like a package tracking record for your asynchronous pipeline. That's a perfect analogy. It records the state transitions. Pending, sending, accepted by carrier, failed, suppressed, or dead-lettered. But building this ledger introduces a massive synchronization challenge, doesn't it? Because we're relying on asynchronous webhooks from providers like Twilio or SendGrid to update the ledger. And network routing means those webhooks can arrive completely out of order.

12:45 Yep. What happens if the delivered webhook from a provider arrives at our ledger before the send webhook? Well, if you just naively update the database with whatever webhook arrives, your ledger ends up in an impossible state. Right. The send ledger has to be designed as a strict state machine. You define valid state transitions. A message can go from sent to delivered. So if a delivered event arrives for a message that is still marked pending... The ledger either buffers the event or uses vector clocks or timestamp reconciliation to realize the sent event was delayed and updates the final state accordingly.

13:20 It provides a durable truth of what the platform attempted and what the provider reported. Yes. Which leads to the biggest trap of observability. Just because the ledger says a payload was accepted by a carrier or delivered to a device, that does not mean the human actually opened the notification or read the text. Exactly. It just proves what our infrastructure accomplished. Yeah. Accepted by APNS is not displayed to the user. Your support and operations teams need to clearly understand that boundary.

13:47 The ledger stops arguments during an incident by providing factual, queryable history, but it cannot measure human perception. We have covered a massive amount of architectural ground today. As promised in the first 60 seconds, I want to bring you back to that full architecture diagram. We can finally see how it all connects. Let's walk the map one last time. Okay. Take us through it. It starts with product services emitting structured notification intents. Those intents hit the policy gate, which enforces your scarce attention limits and user preferences from a distributed cache.

14:19 Right. Then the template renderer safely creates the channel-specific content, maintaining privacy boundaries. From there, the work flows into isolated priority lanes, ensuring your emergency traffic is completely protected from your bulk freight. Like the highway lanes. Yep. Channel workers then pull from those lanes, handle the specific adapters for your external providers, manage backoff and process fallbacks. And through it all, the send ledger acts as a strict state machine, permanently recording every attempt, transition, and outcome.

14:51 It really is a chain of deliberate, highly isolated gates. It earns trust by making decisions explicit and observable. Exactly. Which brings us to the end of this topic and sets up the next step in the series beautifully. In chapter 9, we are going to tackle the fan-out problem from a completely different angle. Oh, that's going to be a good one. We'll be designing a news feed, exploring the massive trade-offs of fan-out on write versus fan-out on read at social scale, and handling real-time feed updates.

15:18 Until then, keep those priority lanes isolated and respect your users' attention.