Ch.20: System Design: Stop a Stolen Cookie After Logout
Outline
- 0:00 A stolen session still works
- 0:16 A copied cookie escapes logout
- 0:41 Follow the stolen secret
- 1:17 What the system must do
- 1:46 What must survive failure
- 2:29 Login is an event; a session persists
- 3:03 One device, one opaque secret
- 3:47 The browser boundary
- 4:27 Create, rotate, remember
- 5:00 One protected request
- 5:41 The tempting cache
- 6:19 Revoke one device
- 6:58 The late gateway
- 7:31 Pay for the guarantee
- 7:50 Two doors, one authoritative desk
- 8:27 Log out everywhere
- 9:09 Two commit orders
- 9:33 Expiry is another boundary
- 10:00 The connection still open
- 10:39 Four failure drills
- 11:21 Prove it without logging the secret
- 12:00 Return to the attack
Transcript
0:00 Alex sees an unfamiliar laptop in the account's session list. Someone has its browser session secret. Alex remotely signs out that device from the phone. The button says done. The attacker sends another protected request, and another gateway lets it through. That's the failure we are designing out. Deleting a cookie on Alex's phone cannot reach a copy already stolen from the laptop. If we acknowledge revocation, what exact request must the system reject next? That promise will decide the architecture.
0:29 The next new request carrying that stolen secret, whichever region receives it. We'll see how a current session check makes that stale regional cache lose. That's what I expect “done” to mean. Now put that new-request promise on one map. The stolen laptop secret travels through a gateway in another region toward Alex's data. Alex's phone sends a separate revoke request to a session authority backed by durable storage. Regional caches receive updates, perhaps late. Keep this picture. We will test its boundaries one by one, then return to this exact diagram when the stolen request makes its last try.
1:05 The attacker and Alex take different routes. If the cache in that other region is one event behind, which path wins? That question is already more interesting than the button label. Now give Alex functional requirements that match those routes. Each device gets its own login; Alex sees both sessions and can revoke the laptop alone or every device if the compromise spreads. Unrevoked sessions stay signed in until expiry. What does the phone need after a lost response? The committed answer, not a guess.
1:35 For example, if the phone loses the acknowledgment, it can ask for the laptop's current state without reactivating it. Alex can keep using the trusted phone after a targeted revoke. That lost-response case is why the non-functional requirements matter. The secret must be unguessable and protected in transit. We need account isolation, a responsive check, and no stale-cache approval after revoke. An operator must explain a denial without finding the stolen cookie in the logs. The denial reason, yes; never the secret.
2:06 Here is the user-facing boundary: once the revoke is committed and acknowledged, every new authorization check for that session must deny. A check ordered just before the revoke may finish afterward; the button cannot travel backward in time. A sensitive write can check again near commit, though that is a mitigation rather than a promise to undo an earlier admission. That boundary makes yesterday's password check insufficient. We cannot rerun it on every page view, so what does the gateway see instead?
2:37 A session secret issued after authentication. According to NIST's session guidance, that continuity proves possession of this session, not that Alex still holds the device. And a stolen copy passes the possession test until the server stops honoring it. Exactly. Logout is not just clearing the local cookie; the attacker still holds a copy. What record can we revoke without throwing Alex's phone out too? So the server keeps a record for each device. Give Alex's laptop and phone different random secrets; each record holds its account, issue time, expiry and active-or-revoked status.
3:14 The browser value is opaque, with no email or role inside it. A security reviewer might ask, “Why not keep the raw token in that row? Lookup would be trivial.” I admit the simple lookup tempts me too. But if the table leaks, wouldn't every row become a ready-made login? Exactly. Store a keyed, one-way fingerprint for lookup instead of the bearer secret. A leaked row then cannot authenticate by itself; protect the fingerprinting key too. Separate records let us kill just that browser session and keep the phone.
3:46 Now send that secret as a cookie limited to our host. Secure means HTTPS only. HttpOnly keeps our JavaScript from reading it. Does that settle the injected-script problem? Only direct reading. A script on our site can still send a request with the cookie attached. Oh, so the browser can carry the secret even when the script cannot inspect it. And another site? SameSite limits when a different site can make the browser attach it. OWASP recommends these controls. They reduce theft and misuse, but the laptop's leaked value still needs the server's revocation check. That's the unpleasant distinction: safer delivery does not cancel a secret that escaped.
4:26 Once those browser limits are set, Alex signs in on the phone and gets a fresh random session. If the phone carried an identifier planted before login, we discard it; the attacker might already know that old value. So we never turn an attacker-known identifier into Alex's new login. The session list shows both devices separately, but those labels are clues, not proof of who holds a copied cookie. Right. We target the laptop's server record, then test the cookie itself. A neat device name cannot make that decision for us.
4:59 With that new phone session, Alex requests the account page. Its cookie reaches the gateway, which derives a lookup key and asks the session authority whether the record is active, unexpired and bound to Alex. Only then can the gateway forward identity inside the trusted service boundary. What if an outside caller supplies that identity header directly? The gateway strips the outside header before adding its own. The application still checks permissions after authentication. For example, an attacker replay must fail its session check before any document permission is even considered.
5:32 Right. That caller cannot become Alex by spelling Alex in a header. But could a cached yes still make the stolen session look active? Now I think a nearby cache is tempting: remember that session as active and skip the home read. It would reduce latency, but could it make the confirmation lie? For example, the laptop's session is cached as active in two regions. Alex signs it out in region one. The event to region two is delayed. A stolen request reaches region two before the cache notices. If it trusts the old “active” entry, our opening failure returns.
6:07 Oh, so even a short cache lifetime leaves a window after Alex sees “done.” That timeout cannot prove the logout worked; the next step has to make the stale answer lose. Now that the cache has failed, what can the revoke button change that a lagging gateway cannot overrule? The authoritative session record. Alex taps “log out that laptop.” We verify the phone's current session and that the target laptop record belongs to Alex. Then we commit revoked there, and only after that durable commit do we tell the phone it is done. Suppose the response vanishes on the way back; the retry reads the committed revoked state and never re-enables it.
6:46 So invalidation is merely a hint. Broadcast it and close known live connections for speed, sure, but the committed record proves the next gateway must deny. What if a gateway misses the event entirely? Next, test the missed event. Send the copied laptop secret to the lagging region. Its cache still says active. Can the gateway return Alex's data? No. To honor the acknowledgment guarantee, this request checks the latest authoritative session status. The stale cached record can help locate the session's home, but it cannot authorize on its own.
7:18 The authority reports revoked, so the gateway rejects it. Hold on. The cache can be wrong and the replay still loses. But if that home region is down, will we block Alex's healthy phone too? Now the outage cost: yes, the gateway refuses even Alex's good-phone request when its home cannot answer. The revoke write and every later authorization read go to that session's authoritative home, so a later check sees the committed result. A cached yes cannot settle it. I think that availability cost will hurt in production.
7:49 Picture two doors with copies of a guest list. Crossing off a name at one desk does not update both copies immediately; for our after-acknowledgment promise, both doors call that desk before admitting the guest. A weaker product could disclose a delay during which a stale copy still admits requests. Ours adds latency and blocks when the desk is unreachable; an earlier admitted guest can still finish walking through. Right. When that desk is down, this design keeps the rightful owner out alongside the thief.
8:20 Rough success state. Does the same current-check rule work for every device at once? But denying one laptop leaves copies from other devices unknown. Alex now chooses “log out everywhere.” Scanning every session row and racing new logins is messy. What single account-level fact can invalidate the old group? Imagine Alex's phone, laptop and forgotten tablet were issued under account generation four. Revoke-all advances the authoritative value to five before acknowledging. Every new gateway check compares a session's generation with that current account value, so all three old records fail even if a cache event is late.
8:58 That catches the tablet too, and signs out Alex's current phone as “everywhere” promises. But what if a fresh login started just before the generation moved to five? Take the two orders. If login commits under four first, revoke-all moves the account to five next. That login response might arrive late with an already dead session. So the client may have to sign in again, even though the login succeeded before the revoke. Yes. If revoke-all commits first, the new login creates under five. In either order, authorization compares the current generation before admitting work. Now that every device can be revoked together, could a short expiry spare the authoritative read?
9:38 There are two clocks. The idle one resets with activity; the absolute one reaches a fixed end even if the attacker keeps clicking. Both run on the server, as OWASP advises. Before either ends, that stolen cookie still works unless we revoke it. Expiry limits exposure; it cannot make a sign-out confirmation immediate. Now suppose the attacker opened a live connection just before Alex's revoke. The gateway checked once at setup. If we merely push a close message, could that old connection still ask for protected data while the message crawls across regions? Yes. Closing it quickly helps, but every new protected action over the connection must ask the current session authority. The same goes for server pushes that reveal Alex's data.
10:25 A connection that was valid at noon does not get a lifetime pass after Alex's revoke. That's the same contract on a different transport. A stolen token cannot ride an already-open socket around the check. Let's make the test runner act as the thief. 1: lose the phone response. 2: delay the cache event. 3: cut off the authority's region. 4: start an allow check just before Alex revokes. Which case can finish after confirmation? The 4th. It was admitted before revoke, so it can still complete later.
10:58 For a high-risk write, we check again before commit, which narrows that window but cannot erase completed work. The other tests expect a safe retry, a denied stolen request, and a blocked healthy phone during the outage. That 4th test makes me wince. The UI can be truthful about new checks and still not undo a request already inside. With those schedules in hand, the gateway needs observable proof. Consider a test that replays the attacker's cookie against every region after acknowledgment and expects denial.
11:30 Reorder invalidation events and crash the cache. A success banner in the phone UI is not evidence of global revocation. On call, I need to tell a revoked-session denial from an unreachable-home denial without preserving the bearer value. Record a non-reusable session reference, gateway decision and reason. For this replay, the reason reads “revoked”; a partition reads “authority unavailable.” And now the test can tell a blocked thief from a region that could not verify anyone. The same red result needs two different fixes.
11:59 Finally, Alex's phone has its confirmation. The laptop's copied secret reaches the other region before its invalidation event. Follow that request: the gateway asks the authority, which reads revoked and denies it. The event can arrive later without changing the answer. The cache can speed routing, but it cannot grant access from a stale yes. We pay for an authoritative read on every fresh protected operation, and losing the home blocks requests. That is the cost of a truthful confirmation. And paying that cost makes each new stolen-session check deny in every region.
12:35 Thanks for listening to Learning Podcasts.