Ch.9: Change Protobuf Schemas Without Breaking Production

Outline

Transcript

0:00 A team ships a Protobuf schema with a field called customer note, field number 7. Months later somebody deletes it, and a new engineer reuses seven for shipping address. Every service stays up. Every parser succeeds. And the archived messages still have field 7 inside them. So an old customer note, imagine something like leave it on the back porch, can come back out of storage labeled as a shipping address. That's a quiet little disaster: nothing to page you, nothing to roll back. Right, and the data is simply wrong.

0:32 Protobuf's official docs are blunt about it: a field number has permanent history. We've all treated a proto file like ordinary source code, though. Rename a field, delete dead ones, reuse what looks free. Every one of those instincts is safe in Python or Java and dangerous here. Which is exactly what today fixes. That exact scenario is a composite we built, but the rule it breaks is real. We'll walk through the four compatibility surfaces to check, then a six-step rollout you can run in your next code review.

1:04 And that story only works because production never has just one schema version. You have old writers, new writers, old readers, and new readers all running at once, because clients and servers do not update at exactly the same time. And in between them, bytes that nobody can redeploy. Take a message sitting in a queue, or a row in an archive... ...bytes written last year by a binary that no longer exists. Storage is the one writer you can't redeploy. Forever bytes. Yeah, nobody gets to go back and rewrite those.

1:37 So rollout order is part of the contract itself, right? It is. The change was never just the edit to the file. It's a choreography: who learns to read the new shape first, who starts writing it, and what happens to everything already sitting on disk. Which means a schema review is really a rollout review. So in that rollout review, everybody asks the same question. Is this change compatible? And that question is missing a word. Compatible where? Because there are four different answers. Okay, let's unpack this.

2:10 Four surfaces, and a change can pass one while failing another. Surface one, the binary wire format, the actual bytes protobuf writes when it serializes a message. Surface two, ProtoJSON, protobuf's official mapping to JSON, where field and enum names become part of the encoded message itself. Wait, the names are in the JSON payload? They are, yes. The name gets a second life inside the payload, and most teams never budget for it. Surface three, the generated code, the Python classes and Java builders the compiler writes for you.

2:43 And surface four, application behavior: defaults, fallbacks, what your business logic actually does when a value is missing. So wire-safe is not system-safe. A change can leave every byte parseable and still rename a Java method, or quietly starve a fallback. I think I would have called two of those "not my schema's problem" this morning. Now we can see why the field number is the first thing to protect. So, those four surfaces all trace back to one number. Why does the number itself carry all the history?

3:19 Because the number is what actually travels. On the binary wire, protobuf never writes the name customer note at all. It writes a tag built from the small integer seven, sitting right next to the value. We all read the names as the identity, but the wire never sees them. So renumbering a field really retires one badge and issues a different one. The badge is the number, not the name. And reusing a retired number hands that badge to a new employee while the old employee's records are still filed under it.

3:51 That's the whole shipping-address failure in one move. Which is why deleted numbers and names get reserved, the schema's way of saying that badge is never reissued. One line in the file, literally the keyword reserved with the number 7 on it, and the compiler refuses the reuse forever. Every future editor gets a guardrail nobody has to remember. With the number's identity clear, start with the gentlest change: adding a field. On the binary wire it genuinely is safe, as long as the number is fresh, because old binary readers don't choke on what they do not recognize.

4:26 And a reviewer will say exactly that. "It's additive, ship it." So what's left? Behavior and JSON. New code needs a valid fallback when the field is absent, which drags in presence, you know, whether a field was actually set versus just left at its default. Take a checkout service that adds an optional discount field: an old writer never sets it, and the new reader has to decide whether absent means no discount or a broken order. And that decision is application behavior, surface four. Then the JSON side has its own catch.

5:00 The official ProtoJSON guide is explicit about why: the field names themselves are serialized, and strict parsers reject fields they do not recognize unless they're configured to ignore unknowns. And you do not own every parser in that path. So one additive line in a diff gets three different verdicts. Wire-safe, JSON-conditional, and behavior-dependent. Now, that was JSON rejecting strangers. Binary parsers do something gentler automatically: unknown fields. When a binary parser meets a field its schema does not list, it does not throw the bytes away.

5:36 It keeps them, and it writes them back out when the message gets serialized again. Which sounds like a free pass. It holds for a normal parse-and-reserialize hop. That's the whole guarantee. Take a JSON gateway in the path, or any code that rebuilds the message field by field into a new object: those bytes are gone. Oh, man. I didn't realize it was that narrow, I assumed it held through anything. So every JSON conversion in the pipeline is quietly a filter. Right, and even when the bytes do survive, transport is not intent.

6:09 The proto3 guide draws exactly that line, between retaining unknown bytes and understanding them. An old service can carry your new field across the system without ever once acting on it. So "the parser didn't crash" only tells you the bytes arrived. The old service isn't processing your field, it's couriering it. Keep that distinction close; the whole migration leans on it. Same shape in reverse next: deletion. Unknown fields were bytes outliving understanding, and a deleted field stops being written long before it stops existing.

6:41 Because the queue still holds it, the archive still holds it, and a rollback can resurrect a writer you were sure was gone. So you stop the writers first, keep the readers tolerant while the stored copies age out or get migrated, and only then delete the field from the file... ...and reserve its number and its name in the same commit. That reservation is the tombstone, it lasts for the schema's lifetime, and it closes the exact hole we opened with. The result is a number that can never lie again.

7:12 Now, retirement covers removing a field outright. Renaming one looks gentler and crosses more boundaries, because on the binary wire a rename with the same number is completely invisible. The tag doesn't change, so the bytes don't change. And it's where teams get burned. Somebody renames a field, the binary side is flawless, and a JSON dashboard quietly goes blank that same afternoon. Ouch. Because the name still traveled, just on different surfaces. Right, that's the whole thing. Three of them. Generated accessors, so the Python attribute and the Java getter change and code stops compiling.

7:46 Documentation and tooling that referenced the old name. And ProtoJSON, where the field name is literally sitting in the payload. So a rename can be binary-invisible and still a breaking API change. If any of those three boundaries is exposed, treat it like a real migration: fresh field, fresh name, staged rollout. The same playbook we're about to run in full. Renames tempt people, and types tempt them harder. Let's say a counter outgrows int 32, and someone wants to widen it to int 64 in place, same field number. And here's where a staff engineer leans in and says, "the guide lists that pair as compatible, the parse works, just change it.

8:28 We are not adding a whole new field for one counter." The proto3 update guide does list a few pairs like that, and they're conditional. That particular one parses, but, the compatibility table has direction and edge cases, and some pairs quietly truncate. A 64 bit value crammed through an int 32 reader doesn't error, it truncates. Huh. Silently. Which is data loss with a green checkmark on it. So the production default stays: new field, new number, explicit conversion in code, controlled rollout.

8:59 Yeah. And the calm version of that rule is easy to remember: nobody ever got paged for adding one field too many. And types aren't the only trap that looks additive: two more warnings, I mean, before we run that migration: enums and oneofs. Adding an enum value is binary wire-safe. But add a refunded state to a payment enum, for instance, and every consumer with an exhaustive switch sends your new value straight into its default branch, or its crash. And oneofs, where only one field in the group can be set at a time.

9:33 The instinct is that it's just grouping, a tidier way to write the same fields. It is not. Membership changes how presence and clearing behave. Wait, setting one clears the others? Every time. And version skew makes it worse: a new binary enforces the exclusivity, an old binary doesn't, so the same bytes behave differently depending on who reads them. Moving fields you already shipped into an existing oneof breaks the wire format on top of that. So for both, ask which running code changes its behavior when this lands.

10:07 Okay, let's hear it. What's step one? So, step one, and it answers the opening failure directly: add, never repurpose. Shipping address becomes a new message on field 12. Customer note keeps field 7, untouched. Nothing old breaks on the wire, because nothing old changed. True for binary. Though imagine a strict JSON consumer tripping over the unfamiliar field. And readers have to handle both shapes before a single writer changes? That's the part I'd get wrong. Exactly that, and it's step 2: deploy tolerant readers first, before any writer moves.

10:41 Every reader that matters learns to check field 12, prefer it when it's present, and fall back to field 7 when it isn't. Readers first is honestly the whole trick, because a reader that understands both shapes is a lot harder to surprise. Right. And at this point the system looks unchanged from the outside. Every writer is still producing exactly what it produced yesterday. You've only expanded what the fleet can understand, which is why this phase is called expand. And it's the phase you can still undo for free.

11:11 With readers tolerant, phase 2 moves the writers, and it's deliberately the slowest part. Step 3: migrate writers to produce the new field. If both representations can be kept consistent, you can dual-write for a while, meaning each writer fills in field 7 and field 12 with the same information. Oh, right, and if they can disagree, don't dual-write at all. That isn't redundancy at that point, it's your next incident. That fear is exactly what step 4 is for. Move the old data and watch. Backfill where it's needed, meaning, a batch job that rewrites stored messages into the new representation.

11:48 Rate-limited, because you do not want to cause an outage while fixing tech debt. And then measure instead of guessing: who still reads field 7, who still writes it, what fails to parse, and which schema versions are actually live. Because the fleet you have is never quite the fleet you think you have. Once the numbers say the fleet has moved, phase 3 contracts. Step 5: stop the old path, in order. Remove the old writes first. Then wait, because your rollback window, your queues, and your archives are all still speaking the old shape.

12:21 Imagine a queue replaying Tuesday's backlog into a reader that already forgot the old field. Only when those are accounted for do you retire the old reads. So the field is sort of dead long before it's deleted. Deliberately. Dead, but still on the payroll until you reserve it. And then step 6: delete customer note from the file, reserve number 7 and its name, and set up a baseline in continuous integration, a known old copy of the schema that every future change gets compared against, automatically. And step 6 is where it stops being a rule and starts being a check.

12:56 The reservation protects the past, and the baseline protects the future. Now, that baseline idea deserves one more minute, because it turns everything we've said from memory into a check. There are tools that diff your schema against its history on every pull request. Buf, for example, groups its breaking rules into four categories, from strict to loose. Four categories. So you're picking what "breaking" gets to mean in your repo? Pretty much. It's a strictness dial. At one end, the file category, where even moving a message to a new file counts as breaking, because generated imports break in Python and Java while the bytes stay identical.

13:35 At the other end, the wire category, where only the binary bytes count. The two settings in between cover package moves and the JSON names. And that menu is kind of a confession: the tool's own authors didn't treat "compatible" as one property, so they shipped four definitions and made you choose. What no checker sees is rollout order, stored data, or your fallbacks. The history becomes executable; the judgment stays yours. And the people who designed the format drew that boundary for themselves.

14:06 Protobuf's own language is evolving through something called editions, and edition 2024 is the latest released one as of this recording. And the design goal says everything: editions change language defaults over time without ever changing what already-serialized bytes mean. So even the protobuf team, redesigning their own language, will not touch the identity of the forever bytes sitting in queues and archives. Field numbers get that level of respect from the people who invented them. So here's the whole thing in one breath: compatibility has four surfaces, the binary wire, ProtoJSON, generated code, and application behavior. And the migration has six moves.

14:45 Add the replacement with a fresh number, and deploy readers that understand both shapes. Then migrate the writers, backfill what's stored, and watch the fleet until the old bytes actually stop moving. Then shut off the old side in order, reserve the retired number and name, and let continuous integration hold the history. Run that list in a code review this week; the official docs we leaned on are linked in the description. Yeah. Nothing left on field 7 but the tombstone. Thanks for listening to Learning Podcasts.