Ch.21: System Design: Ticket Sales Without Overselling

Outline

Transcript

0:00 Mina and Theo click the same numbered concert seat. Mina's card is authorized, then checkout spins. Theo's map stays green. Did capture land? Then we risk confirming both buyers, or giving Theo the seat while Mina's payment is unresolved. Who decides? The seat row decides ownership; the provider knows payment. Authorization is not a ticket. So Mina waits and Theo sees green, but neither has a safe answer. Exactly. Let's follow what happened. That order crosses this whole map. The waiting room meters people into checkout.

0:32 A browse cache serves the seat map. A reservation service owns the inventory row. Payment is outside its database; a recovery worker checks uncertain results before ticket delivery. It's tempting to infer that orderly admission prevents the seat race. But both fans can pass the gate and still click the same chair. Right. The queue decides who reaches the door. It does not decide who gets that chair. True. The provider cannot join our seat transaction either. I want that gap visible on this map. The inventory record names the holder, while the provider still has the payment answer.

1:07 We'll return to this same map after Mina's checkout stops answering. First, follow her in. That route leaves us with functional requirements: Mina browses seats, enters checkout when admitted, and requests a brief hold. She pays against that claim. If Theo loses the race, he gets a clear rejection, not another checkout for the same seat. What if Mina abandons the cart? The reservation expires or is released. If she buys, the system creates one ticket entitlement and can retry delivering its confirmation safely. For example, if the email worker times out, Mina may receive the notice again, but she cannot acquire a second ticket.

1:44 If checkout times out, she reloads the same order to learn its state. Back to Theo's green map: it could still say available while Mina is pending. Does that break the requirement? Only if the map grants ownership. Theo's next hold goes to the reservation authority. Available, held, payment pending, and purchased describe distinct states; green merely invites him to ask. So those buyer states turn into a support test. Theo calls with a green screenshot while Mina has a ticket. What could support actually prove?

2:15 The non-functional requirements start with proof: a durable trail of which order acquired the hold, when it expired or became pending, what the provider reported, and when the entitlement was issued. No card details in that trace. And the hard invariant: never confirm two owners for a numbered seat. Keep each seat transaction short, survive the flash arrival, and make repeated browser and worker requests safe. There's a painful preference hidden there. We would rather withhold a seat temporarily and delay Mina's confirmation than sell it twice. We should tell product and support that up front.

2:51 Fairness also needs a product rule. Arrival order or a shuffled pre-sale queue may decide admission. The first green click still cannot claim ownership; that decision is next. The sale opens and a crowd hits the purchase path. Put a waiting room at the edge to meter entry. SeatGeek has published that separation as a way to protect its purchase zone during spikes; the source is in the description. I think product has to decide what “first in line” means. Cloudflare documents a scheduled pre-queue that can shuffle waiting visitors when the sale starts. That's a choice people should hear, not a hidden guarantee of pure arrival order.

3:29 Also, an edge threshold is not a globally exact admission count. Regional decisions can be approximate. The checkout still has to survive an extra burst. All right, admit both buyers together on purpose. If our correctness story needs the queue to separate them, it isn't a correctness story. Yes. Let them race. That gives us a real test of the inventory write, not a lucky queue ordering. Once both shoppers are admitted, a cached map lets browsers pan without locking the reservation table. For example, Theo may see green while Mina's hold already blocks checkout.

4:03 Many of us have wanted to push live updates to every tab and call the map current. A tab can miss one. Theo still clicks green. And that's the moment the map loses the argument. It answers what recently looked free, not who owns it now. Theo's hold request reaches the authoritative record, fails, and the interface says no longer available. We can make that failure fast and polite. We can't make a stale badge into a reservation. Can the database settle two hold calls arriving together? Once those calls reach the seat row, a short database transaction locks it, checks state and expiry, then records her order, the hold deadline and a version before committing.

4:45 Does Theo's request bounce off Mina's row lock, or wait there? PostgreSQL makes that conflicting transaction wait. Once Mina commits, Theo reads held and loses. We lock this seat, not the entire venue. Suppose they both read free earlier. Could Theo still write free-to-held after Mina? No. The check belongs inside the locked transaction. The old browse read is irrelevant. A conditional update could enforce the same decision. I think that's easy to lose: Mina has an order-bound hold, not a ticket, which means the deadline can still take it away.

5:19 Mina's browser shows a countdown. It is useful feedback, but her laptop cannot release inventory. The server stores the expiry. That puts the cleanup worker in an odd spot. Suppose it sleeps through the deadline. Can Theo take the seat before it wakes up? Yes, if a locked transaction sees an expired hold and atomically replaces Mina's claim with Theo's. The late cleanup job checks the version or order identity; it cannot erase Theo. Hm. So expiry is a condition at the next authoritative write, not a timer firing exactly on schedule.

5:52 But that only works while the row remains held. A committed order awaiting payment no longer obeys the old countdown; that's where the next race begins. With Mina's hold still live, her browser times out and the Buy button still works. Does another click create another payment? It must reuse her application order, associated with 1 provider payment intent. A retry of a particular provider request carries the same idempotency key; a lost response cannot prompt us to casually issue a new operation.

6:24 Stripe recommends one PaymentIntent per order or customer session. A repeated idempotency key can even replay an initial error. Wait, really? Does that replayed error prove the operation did nothing? No. Read the existing order and the provider's current intent state. A missing answer is still an unknown outcome. We've all stared at a timed-out POST and wondered whether retry means duplicate work. Here, a fresh payment intent after that refresh would be the incident we'd deserve. We also keep provider calls outside the seat-row lock.

6:58 That leaves the database lock free while the external operation runs. That order leaves a payment tradeoff. Authorization asks the provider to reserve funds without finalizing the charge. Then a short seat transaction checks that Mina owns the live hold and converts it to her pending order. Capture, the request that finalizes the charge, comes after that commit. My first thought was capture first. It might let us show success sooner. What's the failure? If her hold expired and Theo won, we'd have finalized payment for a seat we cannot give her. Oh, right.

7:32 But now our seat row says Mina while the provider's capture response is missing. Two systems know different parts of the story. Yes. There is no transaction spanning our database and the provider. Theo cannot take that pending seat, and Mina cannot receive a ticket until we confirm payment. That leaves an operational problem, not just a code path. That payment order leaves two races. First, imagine Mina's authorization arrives after the deadline. Theo has claimed the seat, so Mina's final seat check fails.

8:04 Then do not capture for Mina. Check the provider's actual intent state and cancel or release the authorization as appropriate. We don't evict Theo to make payment look tidy. What if that authorization reply vanished too? A timeout is not proof of failure. We retrieve the same intent and reconcile it, without sending Mina through a new purchase. She gets a lost-seat result when we know that outcome. Until then, we don't promise an instant reversal. So Mina loses the seat, and we still owe her a clean payment outcome.

8:35 A committed claim creates a different uncertainty. But a committed seat claim creates the second race. We request capture; its reply vanishes. Theo's map is stale and his hold request comes in. He gets rejected. The sales lead pushes back: “Mina might not have paid. Show me why this seat must stay hidden.” What evidence do we need? First, the authoritative row is pending, not held. The old deadline cannot undo that claim. We retrieve the existing payment intent, retry only the same operation when its state permits, and record what the provider actually says.

9:10 And if the provider is unreachable all afternoon? Keep the seat unavailable and Mina pending. Alert an operator with both identifiers; if the provider stays unreachable, escalate the case and update Mina. Time passing triggers investigation, never resale. Release only after the provider proves failure or a reversal completes, followed by a durable state change here. That leaves one unresolved buyer, not two apparent winners. Once payment is confirmed, the reservation store can create Mina's sole ticket entitlement. A delivery queue carries the message; it does not award seats.

9:43 Suppose the delivery worker runs twice because the first acknowledgement disappeared. Do we send two tickets? No. The entitlement is unique to the order and seat. Repeated delivery can resend a notice or link to the same entitlement. It does not rerun checkout. And if the worker arrives before the provider result? It reads the durable order state and waits. A queued “please confirm” job is not evidence of payment. That prevents an invented success screen. But what if capture succeeded and this worker crashes before writing the ticket?

10:14 That's the next failure to test. And that's the crash I would put in a design review. The provider says captured, then our process dies before Mina has a durable ticket. Does the delivery queue know how to fix it? Not by itself. A recovery scan finds an order with no entitlement. If the capture result was not saved locally, it rechecks the existing provider intent. On confirmed capture, it retries the same unique entitlement write. Only after that commits does a delivery job notify Mina. And a duplicate recovery job races it?

10:48 The uniqueness constraint leaves one entitlement. For example, if two workers both see the missing ticket during an outage recovery, one insert wins and the other reads the same ticket. Mina is paid but pending for a while; she is never asked to buy again. Oof. Mina has paid and still has no ticket. Someone should see that before she has to call. Now we have two different on-call queues: capture still unknown, or captured with no ticket. What do you need to separate them? A trace joining her order, the seat version, and the provider intent, without card details.

11:23 Did capture happen? Is the provider reachable? Has a ticket write been attempted? I want a dashboard that separates unresolved captures from paid orders missing tickets. Alert on their age, show the oldest case, and let support open the exact order rather than tell Mina to click Buy again. And watch the inventory cost. Count seats tied up in each state. During a provider outage, we can say why A12 is unavailable without pretending reconciliation made progress. Now return to Mina before the provider answers.

11:54 The occupied seat has a buyer-facing cost: each fan needs a status we can defend. With inventory held during payment uncertainty, Mina's screen has to be honest too. She reloads after the missing capture reply. What does it actually say? Mina sees “payment pending” on her existing order, with no new Buy button. Theo sees “seat unavailable,” even under a green map. And when the payment result finally says success? Purchased, with 1 entitlement. If her seat claim had failed before charging, she gets a failed checkout and the authorization is handled according to its provider state.

12:29 Held meant a short-lived claim; available only described a recent browse read. Not the instant answer either buyer wanted. But every word matches an actual state we can defend when support calls. Finally, we're back at the map. Theo's green browse hint ends at a rejected reservation request. Mina's amber recovery loop runs the other way: payment, reconciliation, reservation. Right: the provider may know capture succeeded while our order still says pending. The worker carries the verified outcome back to durable local state; it doesn't let a timeout choose a winner.

13:05 Only then can ticket delivery or a proven release proceed. The edge absorbs the rush; the seat transaction names one claimant; that return loop keeps the payment promise honest. Thanks for listening to Learning Podcasts.