Ch.15: System Design: Payments Must Never Double-Charge
Outline
- 0:00 Timeout ambiguity
- 0:52 Architecture and repair loops
- 1:46 Functional requirements
- 2:49 Safety, durability, convergence
- 3:42 Authorization, capture, settlement
- 4:48 Payment state machine
- 5:44 Idempotency across retries
- 7:03 The honest unknown state
- 8:06 Duplicate-safe webhooks
- 9:00 Ledger and transactional outbox
- 10:12 Why balanced can still be wrong
- 11:18 Crash walkthrough
- 12:27 Reconciliation
- 13:05 Refunds and chargebacks
- 13:56 Tokenization and PCI scope
- 14:51 Financial observability
- 16:00 Final system blueprint
Transcript
0:00 Picture this: a customer taps Pay on a 79 dollar order. The spinner spins, and then the connection just dies. No success screen, no error. Somewhere out there, a payment provider may have already moved their money. And now you're holding the worst problem in backend engineering. Did it go through? Because both guesses hurt somebody. Yeah. Retry blindly, and you risk billing a person twice for one order. Refuse to retry, and maybe they paid and nothing ships. Double charge or stuck order, and the network will not tell you which nightmare you're in.
0:35 So that's the whole chapter: how you make retries safe when neither you nor the customer knows what the money actually did. By the end, you'll hold the full mechanism: one name per operation, an unknown state it can admit, and three repair loops that always converge. Now, here's the finished architecture up front. Read it by its repair arrows. Left: checkout, with provider-hosted card collection. Middle: the payment orchestrator and its state machine, talking to the payment service provider, the outside company that actually moves money.
1:07 Arrows before boxes? Okay. Then the right side is everything that catches the drops. A webhook inbox for the provider's own reports. A ledger. An outbox feeding the order system. And a reconciliation loop with an exception queue at the end. And three arrows matter more than all the boxes. The provider knows and we don't: a webhook or a lookup repairs that. We committed and downstream never heard: the outbox relay repairs that. Anything still off: reconciliation catches it. Keep the picture in your head.
1:38 We'll build every piece, then come back to this exact map at the end, and those three arrows should feel inevitable. So, the functional requirements, in the order the money moves. Create a durable payment intent, one record saying this purchase exists, with a stable payment ID. Authorize and capture through the provider. And let every operation be retried safely. Then the unglamorous verbs. For example, show payment status without trusting the browser's redirect as financial truth. Record every value movement in a ledger.
2:11 Tell fulfillment when money is actually in. Yep. Accept webhook reports from the provider, including duplicates and events that show up in the wrong order. Recover a payment whose result went missing. And support refunds, partial refunds, and disputes as real operations, not afterthoughts. Plus one requirement teams forget: give operators a genuine exception queue, because some payment, someday, will need a human with a paper trail. The list looks long, but, you know, every line is one promise: the system always knows where the money stands.
2:43 And the provider only sells you the money-moving lines. The remembering is all yours. And the non-functional requirements are where it gets strict. First by a mile: duplicate-money safety. One logical operation must never produce two provider side effects, no matter how many retries, workers, or webhooks pile on. Then durability and auditability. Every accepted intent, every transition, every ledger row survives a crash, and every state change records what triggered it, from where, for how much.
3:13 There's a quieter one I think matters just as much: convergence. Any payment stuck in doubt has to resolve eventually, through a lookup, a webhook, or reconciliation. Doubt is allowed. Permanent doubt is not. Huh, I like that one. And availability with a spine: checkout can accept an order while downstream is limping, but it must never pretend an unknown payment succeeded just to keep a graph green. The whole middle of the map exists to keep that one promise. Now, the first trap hides inside one innocent word: paid. That word is four different facts.
3:46 Authorization, the card network reserving the funds. Capture, actually collecting them. Settlement, the money physically landing in your account days later. And fulfillment, your warehouse shipping the thing. Every integration engineer asks the same thing sooner or later: my provider exposes one call that authorizes and captures together, so why do I need four facts? Because the bundle is a convenience, not a physics change. The authorization can expire if you never capture it in time. A capture request can be accepted and still, rarely, fail afterward, reported later by the provider.
4:20 And settlement always arrives on its own clock. Hold on. The provider says yes to the capture, and the money still might not land? Rare, but documented, the provider docs are linked below. Which is exactly why one boolean called paid is a lie waiting for its moment. So the model is four dated facts, and the proud little paid boolean retires. Say the order ships on captured but the books close on settled: different questions, different days. So let's give the purchase a spine. One logical purchase gets one intent row: a payment ID that never changes, the amount, the currency, and a canonical state.
4:58 Every attempt, and every later refund, hangs off that one identity. And the states walk a strict path from created to authorization pending to authorized to capture pending to succeeded, with honest side doors for needs customer action, failed with a reason, and canceled while canceling is still legal. Plus the door this chapter exists for: unknown. Any pending call whose connection died without an answer lands there, and it gets its own section in a minute. And one discipline turns the diagram into a machine, hmm: every transition is compare-and-set on the row's version.
5:34 Let's say a stale worker tries to drag succeeded back to pending. It just bounces off. Which means concurrency can't corrupt the story. Then comes the load-bearing word: idempotency. Plain version: safe to do twice, because the second time changes nothing. The catch is what the key names. Imagine the classic retry loop minting a fresh UUID on each attempt. People assume any unique value works, and now every retry looks like a brand-new charge. Ouch, yeah. The key has to name the logical operation, not the attempt. Something like capture, order O-4815, version one, and that exact string rides on every retry of that capture, forever.
6:15 With a request fingerprint stored next to it, so a reused key with a different amount gets rejected instead of guessed at. And providers meet you halfway: the big ones store your first result under the key and replay it on retry. One even documents this exact scenario, a capture timing out after it succeeded, where the retry returns the original outcome, not a second charge. With fine print, though. Keys get pruned after a retention window, and they can be scoped to an account or a region. So your own operation record stays the durable source of truth, and the provider's dedupe is a safety net, not the system of record.
6:53 So end to end: the same identity from the checkout edge, through the orchestrator, down to the provider call. Three layers, one name for one operation. So now that door in the machine, the unknown one. The connection died mid-capture, and the provider may or may not have the money. What is the state of that payment right now? Not failed, and not succeeded. You genuinely don't know, and most systems can't say so. I'll confess I shipped one that couldn't. Let me guess: the timeout landed in a catch block that wrote failed?
7:27 And checkout offered a cheerful try again, and support found the double charges before we did. So unknown has to be a first-class state: it parks the payment, blocks new charges, and starts a clock. And it resolves through exactly three doors: ask the provider's API directly, wait for the verified webhook, or let reconciliation sweep it up. So prefer an honest unknown over a false success or a false failure, always. False failure double-charges. False success ships free furniture. Honest unknown just waits, and one that ages pages a human.
8:01 That gives every ambiguous payment a deadline. Now, those webhooks deserve their own machinery, because the provider's reports arrive like weather: duplicated, out of order, retried across days. Which the providers say out loud, by the way, links in the description. Their own docs tell you to expect the same event twice and to fetch anything missing from the API. So the shape is an inbox. Verify the signature first, Junk anything unsigned. Persist the raw event with its provider event ID. Acknowledge fast. Process later, asynchronously, and let the event ID turn duplicates into no-ops.
8:38 And when an event arrives out of order, for example yesterday's authorized event landing after today's succeeded, and wants to move the state backwards? The machine just declines. Webhooks are evidence, not commands, that's the shift: your machine decides what each verified fact means from where it stands. So the inbox can replay all day and the story never bends. So here's where I show you the fix that feels right and isn't. The payment succeeded, and the order side needs to know. Everyone's first instinct, honestly: open a database transaction, update the payment, call the order service right there, commit when it all works. One transaction, perfectly atomic. What breaks?
9:18 Go for it, count them. You're holding row locks across a network call, so if fulfillment answers slowly, your payment table queues up behind somebody else's latency. And the deeper one? The transaction is a local promise. It makes your database atomic, nothing else. The order call either already happened when you roll back, or didn't when you commit. A local commit cannot reach across the network and make two systems agree. So instead, the transactional outbox. In one local transaction, commit three things together: the payment state, the ledger entries, and an outgoing payment-succeeded event in an outbox table.
9:56 A relay publishes it afterwards. And because the relay can crash mid-publish, it may send duplicates, the pattern's own documentation warns you. So each event carries a stable identifier and the consumer dedupes. Delivery is at-least-once. The effect is exactly once. Now money needs its own memory, separate from workflow. That's the double-entry ledger: every movement recorded as a transaction with at least two entries, and the debits and credits balance to zero. In plain terms, money never appears or vanishes.
10:26 It only moves between named accounts, one gaining exactly what another gives up. Right, and let's say we capture our 79: two ledger entries land in one stroke, one on what the processor owes us, one on payment clearing, opposite signs. Posted entries are immutable, permanently; corrections are new reversing transactions, so the mistake and the fix both stay on the record. That's what makes an audit a read, not an archaeology dig. Wait, though, let me poke at the worship. A balanced ledger proves balance.
10:58 It does not prove truth. Book the same capture twice, balanced both times, and your books are beautifully, symmetrically wrong. So keep the split: workflow truth lives in the state machine; value movement lives in the ledger. Neither alone is proof. Idempotency guards the writes; reconciliation checks them against the outside world. Now the payoff run, and wait till you hear how much goes wrong: take a single order with the worst luck in the world, every net firing in sequence. Order O-4815, 79 dollars. The capture request goes out under its stable key.
11:32 The provider captures. The response dies on the way back. The operation goes to unknown. No second charge, order unfulfilled, clock ticking. The provider's webhook arrives, twice. The inbox stores the event once and shrugs at the rerun. Then one local commit moves the payment to succeeded, writes the balanced ledger rows, and drops payment-succeeded into the outbox. Oh boy. And because this order is cursed, the process now crashes before the relay publishes. Doesn't matter, I mean, the event is inside the commit.
12:06 The relay wakes up, publishes it, maybe even twice, and downstream, the order consumer already knows that ID, so it fulfills once. One charge, one shipment, zero heroics. Every mechanism we've built just earned its keep on a single unlucky Tuesday, and really, that's the pitch: the system converges even when everything flakes at once. Then reconciliation is the loop that assumes we're wrong anyway. Take the provider's transaction reports, the fees, the settlement batches, the actual bank payouts, and diff all of it against our states and our ledger.
12:40 Reconciliation is basically couples therapy for your books and the provider's records. Where the provider's statement always thinks it's right. And the annoying part it usually is. So every mismatch, an amount off by a fee, a capture that never settled, a payout that doesn't add up, goes to a durable exception queue with an owner and a documented repair. Never a silent balance edit. So what about money going backwards, when the customer wants that 79 back? Yeah, though not by pressing undo. A refund is a new money-moving operation with its own identity, its own key, its own states, its own ledger movement.
13:19 Partial refunds too, several if needed, up to the captured amount. Which means the original capture never gets edited into something else. History reads charged, then refunded. Two facts, both true. Chargebacks are kind of the same discipline under worse weather. The customer's bank disputes the charge, the amount gets pulled, you submit evidence, and sometimes the money comes back. Its own workflow, its own trail. Imagine wiring that as a status overwrite on the payment row instead. You'd erase the exact history the dispute is arguing about, and the fights are precisely when the records have to be intact.
13:55 One boundary we haven't drawn: the card number itself. The goal is that raw card data never touches our servers at all. The provider hosts the collection fields, the customer's browser talks to those directly, and what we store is a token, a stand-in reference that's useless outside our provider relationship. That's tokenization, and it's what keeps the compliance surface small. PCI DSS is the card industry's security standard, and every system that touches real card data sits inside its scope. Scope is expensive.
14:26 Huh, wait. Systems that only ever touch the token, are those still in scope? Properly isolated, they can sit outside the cardholder-data environment. But the council's own guidance keeps the tokenization environment itself in scope, and the merchant keeps real obligations either way. So: tokens shrink the blast radius, and the final scope ruling comes from a qualified assessor, however confident the blog post sounded. Then, once the workflows exist, the last mechanism: proving they work before production proves it for you.
14:57 Drills. Fire two capture retries concurrently and check the provider saw one operation. Kill the worker right after provider success and watch unknown resolve. Replay yesterday's webhooks out of order. Let's say nothing breaks: good, now break the refund path on purpose. And there's a voice worth staging here, your on-call a month after launch: the dashboard is green, requests are all two hundreds, and finance says we've been 11 dollars off since Tuesday. Oh, man. That's the gap: the service being up and the money being right are different claims, and standard dashboards only watch the first. So financial observability watches the money facts: unknowns aging past their window, webhook backlog growing, idempotency conflicts spiking, any ledger transaction that fails to balance.
15:42 With thresholds that are yours, not universal. A marketplace's normal and a subscription service's normal are different planets, so you alert on your own baseline drifting. And every alert exists because some assumption on the map can rot silently. Which is our cue to go back to the map. So, same map as minute one, and every box earns its label now. Hosted collection keeps card data out. The orchestrator owns one identity per operation, a state machine that refuses nonsense, and the same key on every retry.
16:15 The inbox turns webhook weather into deduplicated evidence. The ledger remembers every movement, immutably. The outbox makes the local commit and the downstream news one atomic act. And reconciliation diffs our whole story against the bank's. And the three repair arrows from the start: provider-knows-we-don't heals by lookup or webhook. We-committed-they-never-heard heals by relay. Everything else heals by reconciliation. Honestly, that's beautiful design in the least flashy way. No single clever trick, just ambiguity boxed in and forced to resolve.
16:47 Exactly-once isn't a network feature you buy. It's a business outcome you assemble. And the 79 dollars that started all this? Charged once, shipped once, reconciled to the penny, even on its cursed Tuesday. Next time, we go multi-region: one database stretched across continents, and why letting two regions accept writes at once is the hardest promise in distributed systems. Thanks for listening to Learning Podcasts.