Ch.18: System Design: Alerting Without Alert Storms

Outline

Transcript

0:00 Checkout is failing. Your phone gets 40 pages, mostly about individual servers. The checkout owner gets none. We detected plenty. We still missed the incident. Let's design the alerting system that reaches the right engineer, without making them sort the outage from their lock screen. So follow the missing page: collectors measure checkout, rules judge the harm, alert managers organize notifications, and an incident service reaches the owner. We'll return to this map with a failure to catch at each boundary.

0:30 Our checkout is fictional. Google's reliability books and the Prometheus project ground the design. Source links are in the description. First, functional requirements: collect measurements, detect customer harm, combine related alerts, and notify an accountable owner. If the page reaches the team but nobody takes it, have we actually handed the incident over? No. Someone must acknowledge it; otherwise escalate. Keep enough evidence to investigate and explain why the page fired. Otherwise we wake the right person and give them nothing useful.

1:01 Next, non-functional requirements: deliver the first notification within 2 minutes of a complete regional checkout outage, assuming steady traffic and working delivery. That's our exercise target, not a promise for every slowdown. Bound collection lag and resource use. And survive a regional failure. I'd rather dismiss a duplicate than discover nobody called while checkout was down. With those requirements, the first temptation is to page on every unhealthy server. But suppose a replica disappears and the remaining replicas handle checkout successfully. What should the on-call engineer actually do?

1:36 Maybe nothing immediately. Keep the failed-replica signal for diagnosis and recovery. Page on checkout failures or unacceptable latency. Prometheus recommends alerting on customer symptoms instead of listing every possible cause. Unless losing that replica removes our last safety margin. Then imminent harm and a necessary action can justify a page. The rule isn't that infrastructure never matters. It's that the interruption needs a reason. That urgency leads to different ways of responding. Imagine disk usage growing steadily.

2:06 There's enough room for a week, and expanding it is routine. That's a ticket. What if the same disk has minutes left and will stop checkout? Now someone must act. Page for action now. Make a ticket for scheduled work. Keep the dashboard for context. Each asks for a different response, not just a different severity color. For the disk with a week left, keep the expansion ticket and the graph. I'd rather not page until we can name the urgent action. Yeah, the graph can be interesting tomorrow, too.

2:37 Checkout can't wait. How do those 40 messages become one useful page? First fix identity. Suppose a checkout symptom persists and sends repeated updates. Give the ongoing condition a stable identity: the rule, service, environment and the region we intend to distinguish. Not the exact error message? That seems useful. Useful in the details. Dangerous in the identity. If every changing message creates a new label set, an ongoing incident can keep introducing itself as a stranger. Put changing explanations in annotations.

3:10 Those descriptive fields don't define identity. Keep the identifying labels stable, and retain instance details without automatically paging separately for every instance. Next, deduplication recognizes another copy of the same alert. Grouping combines different alerts because they share selected labels. That's removing duplicate letters versus putting related letters in one envelope. And the envelope still contains the individual letters. We want one useful notification, not deletion of the affected-instance evidence.

3:37 For example, group checkout symptoms by the service and region the owner can act on. Don't group the whole company just because everything says critical. Nor claim the group proves a common root cause. Matching labels are an organizational decision, not a diagnosis. Even if we choose to group a database alarm with checkout errors, they might still be separate incidents. Grouping changes the delivery path. Suppression changes whether we notify at all. Imagine an on-call lead saying, can't we silence checkout while its database is down?

4:08 I think that's too broad. A bad checkout deployment could happen during that database incident. Suppressing every checkout symptom would hide the second failure. Then inhibit only known dependency noise: while the matched database alert fires, suppress those connection notifications. Keep customer harm visible. That's conditional suppression. A silence is time-bounded instead. For example, cover one cluster during 20 minutes of maintenance, with an owner, a reason and an expiry. Not a wildcard someone forgets after lunch.

4:37 But keeping harm visible still leaves the urgency question. Now define a service-level objective, or SLO: our reliability target. Say checkout should succeed for 99.9% of eligible requests over 30 days. The allowed failure fraction is one 10th of 1%. That's our error budget. Eligible matters. In our exercise, exclude invalid-card attempts from both counts. An eligible purchase that fails because checkout is broken still counts. Define that policy before looking at the graph. So this is a request budget, not permission to stop checkout for a fixed amount of time. Right. Losing a busy afternoon can harm far more requests than losing a quiet night.

5:18 With that budget, for example, 1% of eligible checkout requests fail. Divide that by the allowed point 1%. We're burning the budget 10 times faster than the target permits. 10 Times doesn't sound like an outage siren until you attach a clock. At constant request volume, with that failure rate continuing, a full 30-day budget lasts 3 days. Those assumptions matter. Traffic changes, and we may already have spent some budget. Burn rate tells us how aggressively we're consuming reliability, not an exact deadline printed on the incident.

5:51 Okay, but when does that number justify waking someone? Then use two views of the same symptom. A longer window asks whether the harm is substantial. A shorter one asks whether it's still happening. Require both for that paging rule. Like a leak detector that checks both how much water accumulated and whether water is still arriving. The puddle alone could be yesterday's leak. Right, the long window measures the puddle; the short one checks the continuing leak. Google's workbook starts with an hour and 5 minutes, both above 14.4 times normal burn.

6:24 That's a suspiciously specific number. Where did it come from? Spending 2% of a 30-day budget in an hour, when request volume stays steady. The budget policy gives us the number. A separate slower rule catches lower sustained burn. But does an hour-long lookback already break our 2-minute objective? No. Our threshold works out to 1.44% errors. Picture an hour of healthy requests. Failed ones start replacing them; complete failure fills that share in about 52 seconds at steady traffic. Oh, we already have the history. We don't need a fresh hour.

6:59 But an explicit continuous-failure timer adds waiting. Repeated outages shorter than that timer could keep resetting it while customers keep suffering. Yes. Brief waiting can tame a rule that keeps firing and clearing, but it costs detection time. And 10% failures cross our hour threshold much later than total failure. The 2-minute target needs its stated outage conditions. Those percentages raise a different problem when almost nobody uses checkout. Say two overnight requests produce one failure.

7:30 50% errors. Does that deserve the same response as thousands of failed purchases? Not automatically. It depends on the transaction's importance and whether the evidence supports urgent action. A rare, high-value operation might demand an immediate response to one failure. A throwaway check might not. Sure, importance still matters. When real requests leave too little evidence, synthetic probes can exercise a critical path with test transactions. Separately from the customer denominator. Otherwise successful fake traffic can make real customer failures look smaller.

8:01 And one probe doesn't cover every checkout path. Those requests and probes still need a collection path to the evaluator. Imagine their measurements stuck behind a central queue. Regional collectors avoid that dependency by keeping recent measurements locally. Our example workload is roughly 67,000 samples a second. My first thought is to centralize ingestion at that scale. I like having one place to buffer everything. But would that make every region wait for it before paging? Not in this design. The regional evaluator reads local storage.

8:33 Historical export can use buffering, but its backlog doesn't delay detection. So we accept separate regional collection to keep that shared queue off the emergency path. Now the local collectors have to survive the load themselves. And those collectors can still choke on their own labels. For example, add a customer identifier to every metric and the number of separate measured streams grows with the customer base. That count is cardinality. We'd buy monitoring capacity every time we acquired a user.

9:02 Potentially, yes. Labels create separate series, not free searchable notes. Keep user-level investigation in an appropriate log or trace system, with its own access controls. Then enforce a series budget at ingestion. Otherwise one new label can overload the very system meant to tell us we're overloaded. And labels decide which measurements we combine. For example, one replica handles 100 requests and fails 10. Another handles 9,000 900 and fails none. Averaging their percentages says 5% failed.

9:32 But 10 failures out of 10,000 requests is point 1%. That first replica handled only a small share of traffic. Yes. Sum failures and requests over the same scope, then divide. One more trap: a restarted counter drops back to zero. Compute each replica's rate before combining them, so that drop is handled as a reset. That gives us a test: restart a busy replica without changing customer outcomes. The aggregate shouldn't suddenly invent improvement or failure. But a correct calculation still needs evidence.

10:04 For example, the error series disappears while checkout is still broken. Prometheus can deactivate the symptom alert when the expression no longer returns it. We mustn't turn that into a recovery message. Oh, no. We'd be congratulating ourselves for losing the thermometer. What keeps the incident open? Our incident workflow requires fresh healthy evidence or recorded human confirmation. If the integration can't enforce that, disable automatic resolution. And what tells the owner that we've lost the measurements?

10:32 A separate coverage alert for missing expected targets or stale samples. Holding the old symptom open for a fixed time alone would not prove recovery. Now follow the page on our clock. The threshold takes roughly 52 seconds. Budget collection and evaluation too, and we reach 1 minute 17 before grouping. Only 43 seconds left for gathering alerts and getting through to the contact. Give the new group 10 seconds, then 20 for delivery. That's 1 minute 47 overall. 13 seconds spare. One retry could use that up.

11:05 Yes. These are exercise allowances, with collection working, not measured promises. Existing groups have a separate update timer; test those rather than assuming this first-page budget applies. And start the clock at the customer failure. A fast notification after a slow rule still arrived late. That address points back to a service directory: checkout maps to an owning team and its on-call schedule. The page needs affected service and region, user impact, when it began, a useful dashboard and the first investigation step.

11:36 Say the team has renamed itself. Does the old route still reach the scheduled person? Then a correct alert can still reach nobody useful. Review ownership changes like routing changes. A default route should reach an accountable fallback owner, not quietly discard unmatched alerts. And the runbook should help distinguish likely failures, not say check the logs. That's not a runbook; that's the job description. And the person needs to receive it before that runbook helps. Provider acceptance isn't yet delivery, and delivery isn't acknowledgment.

12:07 Say the first engineer takes responsibility, then goes offline during a shift handoff. Has the incident ended? No. Acknowledgment names responsibility; resolution ends the incident. Configure the handoff and acknowledgment timeout so work can't silently lose its owner. So repeatedly texting the same unreachable phone isn't escalation. Who takes over if nobody acknowledges in the first place? The incident service tries the next responsible contact according to policy. Test that schedule. Acknowledgment can also time out and re-trigger, depending on the service's configuration.

12:39 Now consider a delivery timeout. For example, the incident provider accepts an event, but its response is lost. We can't know from our timeout whether it created the incident. Yes, retry with the same stable deduplication key, rather than a fresh identifier. PagerDuty documents matching keys folding subsequent alerts into an unresolved alert. Retain the attempt history so delivery failures are visible. That makes retries safer, not magical. A stable key doesn't guarantee one phone call, and it doesn't prove a human responded.

13:11 Bound the retry delay, watch the oldest undelivered critical event, and have an explicit alternate contact path for a provider failure. That delivery path needs redundant alert managers too. For example, the evaluator sends the checkout alert to both North and South, our two Alertmanager instances. Not a load balancer that picks just one. Both desks get the alarm itself; sharing the call log wouldn't be enough. What happens when North and South can't talk? They may notify independently rather than wait for agreement.

13:41 Normally, a separate link shares notification and silence state. Prometheus explicitly favors duplicates over missing a critical notification. I'd accept that failure mode, but only if both sides really receive the input. And none of this makes an unreachable delivery provider reachable. But even when delivery survives, combining regions can hide what it should deliver. For example, our smaller region is down while the larger one is busy and healthy. The global ratio makes the outage look minor.

14:11 So global health doesn't clear the smaller region. I'd keep a symptom rule there whenever the business requires regional protection. And put its observer outside the failure we're trying to survive. What if every evaluator lives beside the failing checkout service? We'd lose the evidence along with the service. Name that failure domain, then keep an independent observer and a notification route outside it. But a surviving route can still be muted by configuration. Imagine a maintenance silence intended for one database.

14:42 Someone removes the service filter and leaves only the production environment. The preview suddenly includes checkout and unrelated production services. That's a much bigger change than muting database noise. Review the actual matches, restrict who can approve it, and keep the prior configuration available to restore. And record who made the change, why, and when that silence expires. Yes. A configuration can parse perfectly while suppressing the wrong service. Test what stops arriving, not just whether the file loads.

15:13 Beyond configuration, watch the live paging path for failure. Run a known test signal through it and have an independent observer expect its arrival. If that heartbeat disappears, the observer uses a different notification route. That test arrives. Can we now trust every checkout rule? No. It shows this test reached its destination on this attempt, not that every rule is correct. Now fail its collector or provider; the independent observer must report that gap through a different route. Then the watchdog test has its own expected failure, not just a green health endpoint.

15:48 Exercise the real integration with a clearly labeled test and an agreed destination. Before declaring recovery, return to the deployment failure we almost hid. Say the database recovers, but that broken checkout release remains. Which incident can we close? The database incident, once its recovery is established. Not the checkout incident: fresh requests still fail. That's the payoff from preserving separate evidence inside the group. So the group isn't a single truth value. One explanation can end while the customer problem continues.

16:19 Exactly. Correlation suggests a possible cause. Contradictory evidence must stay visible. One recovery message must not close unrelated customer harm. Finally, test what arrives when the primary owner is missing from the schedule. We expect the backup owner to receive the page, not a successful send to nobody. And for missing samples, we expect a coverage warning where urgent, never a false checkout all-clear. Match the actual messages to those predictions. Also let a maintenance silence expire during a live incident.

16:51 Do we release one useful notification or a wall of them? Test suppression ending as carefully as suppression starting. Keep the same outcome-based tests for recovery, sparse traffic, collector loss, partitions and provider failure. Those tests lead to the final design review. Track missed significant incidents, unnecessary interruptions, time to the correct owner, acknowledgment delay and delivery failures. A lower page count is useful only if the important incidents still arrive. The easiest way to achieve zero noisy pages is to turn the pager off.

17:22 Excellent dashboard. Terrible service. Compare reviewed incidents against the pages they should have produced, and inspect what woke people without a useful action. That feedback changes rules, grouping and ownership. It doesn't just turn every threshold up until the team stops complaining. So hand over the full map. Platform owns regional evidence, evaluators and alert managers; checkout owns its harm and recovery rules. Incident delivery owns routing, retries and escalation. Who owns the watchdog when its alternate contact goes stale?

17:52 Platform, in this exercise, with a named on-call route. Otherwise the independent observer has nobody responsible for hearing it. Then test that stale contact before handoff. Independence alone didn't assign anyone the job. And keep historical export outside this emergency route. Each boundary has a failure test and an accountable owner, not just a box on a diagram. Before the next incident, pick one unwanted page. What immediate action did it actually require? Name who could perform that action, and trace whether the notification reached them soon enough.

18:24 Follow that failure through the system before adding another rule. For the measurements behind that investigation, our OpenTelemetry series follows how engineers diagnose systems they didn't write. Thanks for listening to Learning Podcasts.