Ch.10: Protobuf vs JSON: When Is a Schema Worth It?
Outline
- 0:00 One Payload, Three Contracts
- 1:07 Binary vs Text Is the Wrong Axis
- 1:57 What Raw JSON Guarantees
- 2:56 JSON Can Have a Real Contract
- 4:11 What Protobuf Bundles
- 5:13 Developer Contracts Came Second
- 6:04 The Hidden Costs on Both Sides
- 7:18 Case 1: Partner API
- 8:23 Case 2: Internal Event
- 9:37 A Hybrid Architecture
- 10:19 ProtoJSON Isn't an Escape Hatch
- 11:17 Performance Is Question 6
- 12:07 The Six-Question Checklist
Transcript
0:00 One order payload. An order ID, a status, a total in cents. I can show you that exact same data living as three different systems: raw JSON that everyone just agrees on, JSON backed by an OpenAPI contract a machine can check, and Protobuf compiled from a schema file. And all three can be the right call. That's what starts the arguments. Teams treat Protobuf versus JSON like a religion question, and picking on ideology means either a toolchain nobody needed or a quiet data bug a year from now, when the contract breaks and nothing crashes.
0:34 So today we walk through what each contract level actually buys you, then run six questions against two real-shaped boundaries, a partner API and an internal event, so you can defend the choice in a design review. The quick version of the 6: who reads the boundary, what contract already exists, how many versions coexist. How long the data lives, what the toolchain costs, and whether measured performance even changes the answer. Both cases are fictional composites, built to show the method. And no secret benchmark chart at the end, either.
1:05 That's a promise. Now, the way this debate usually goes, someone says binary beats text, and someone says schemas beat no schemas. And both of those axes are broken. Broken how? Binary versus text only tells you whether you can read a payload in your terminal. It's arguing cardboard box versus crate without asking who has to open the thing at the other end. And schema versus no schema pretends half the comparison doesn't exist, because JSON can have a real, enforced schema. We'll get there. So what's the variable that really predicts the right answer?
1:40 Coordination pressure. How many people, languages, services, and moments in time have to agree on what these bytes mean. Low pressure, conventions survive. High pressure, you start paying for contracts. Which means the format question is really a question about who has to stay in sync. So think of the three as a ladder. Bottom rung, raw JSON. According to RFC 8259, the official standard, JSON is a lightweight text format: objects, arrays, numbers, strings, booleans, null. That's the whole promise; links are in the description if you want to read it.
2:14 Wait, that's the entire standard? Yes. Which fields an order needs, allowed status values, what a missing total means, how a change rolls out, all of that lives outside the format, in your tests, your docs, your team's memory. It's a handshake agreement with a wire format attached. That explains some incidents I've watched, honestly. The payload was valid JSON the whole time. Valid just never meant correct. True, and to be fair to JSON, you can read it, diff it, and fix it with your eyes. Take a failed partner request: pasting the body into a ticket is genuinely useful debugging, and no schema-aware tooling stands between you and the data. Now, here's the correction, though: JSON can carry a serious, machine-checked contract.
3:01 Okay. So it's still the handshake, just with a contract drafted around it. Middle rung, then. Two specs do the heavy lifting. First, JSON Schema describes and validates JSON documents, required fields, value constraints, the works. Second, OpenAPI describes an entire HTTP API. Take an order endpoint: the same document that documents it can validate its requests, generate clients for it, and drive its tests. So when a teammate drops "we should switch to Protobuf because JSON has no schema" into a review thread, you can just say back: your JSON might already have one.
3:37 I mean, compare Protobuf against the contract you're actually running today. The sloppiest JSON you ever inherited is a strawman. I think that's the trap I would have walked into this morning. But a written spec doesn't mean anyone enforces it, right? Right, and that cuts both ways. Enforcement is assembled, not assumed: validators in the request path, generated clients, compatibility checks in your build. The document alone is paper; somebody still has to remember to lock the door. So check which of those three you actually have.
4:11 With that corrected, Protobuf's actual offer gets clearer, because it was never just "finally, a schema." Top rung. It's a bundle: one proto file feeds a compiler, the compiler writes your Python classes and Java builders, runtime libraries ship for each language, and the bytes follow one canonical binary encoding. Okay, but if the schema is that strict, don't the rules have to ride along in every message? That's what I'd assume. Feels right. It's the opposite, and that's why the binary can be so lean.
4:42 Take a single field: on the wire it travels as a field number and a wire type, a small code that tells a parser where the value ends. The name never travels. The catch is that nothing in the message tells you what field 7 means. The proto file is the decoder ring, and it lives outside every message. Which is what people mean when they say binary Protobuf isn't self-describing: the bytes alone can't explain themselves. Lose the schema, and you're holding compact gibberish. And the official history is really a story about that outside-the-message contract.
5:17 The earliest version, inside Google around 2001, didn't generate classes at all. They started with the version you would hate: callers manually added tag-and-value pairs, by hand, field by field. So somebody sat there typing out, this one is field 3? Exactly that, for every message. Generated classes only arrived as the design matured. Huh. I didn't realize the famous part came second. So that tells you what the design kept investing in. The compact bytes existed on day one; what it kept adding was the developer contract wrapped around them, the compiler doing the repetitive work for you.
5:56 That's the part everyone lives in day to day, and it's exactly the thing you're pricing when you ask whether the toolchain is worth it. Okay, before the two cases, let's take question five out of order, the toolchain cost: both sides' hidden invoices, because each one hides its costs somewhere different. Raw JSON's tax shows up as duplication. Consider a team running five services, each carrying its own copy of what an order "should" look like: validation here, defaults there, drifting apart quietly.
6:27 And nobody notices until two copies disagree in production. Though to be fair, that drift isn't inevitable. A very disciplined org could theoretically hold raw JSON together with reviews and immaculate docs. Sure, and theoretically I could commute to work on a unicycle. Possible; not a transit plan. While Protobuf's tax shows up as integration: the compiler in every build, runtimes in every service, generated code churning through review, schemas that have to reach every consumer, and the day you realize you can't just read a production payload with your eyes anymore. The official overview is blunt about the edges, too: binary messages aren't compressed, aren't self-describing, and aren't meant for data that doesn't fit in memory.
7:10 It's a tool with limits, not a default. So neither choice is free. Pick the invoice your team can afford to keep paying. First case, then. Your company opens an order API to partners. HTTP in, HTTP out. One partner is a TypeScript startup, one is a .NET shop, one is a PHP script somebody runs on a cron job. They all debug by reading requests, and imagine the API already ships an OpenAPI document: reference docs, generated clients, request validation wired into the gateway. Isn't this one already obvious, though?
7:42 Sure, and that's exactly why it's the right warm-up: watch how fast the questions close it. Question one, who owns and reads the boundary? Partners you don't control, reading payloads by hand. Question two, what contract already exists? A real one, already deployed. Question three, versions: every partner moves on its own clock, but they all already speak HTTP and JSON. So Protobuf here would mean exporting your integration tax to companies that never asked for it, to solve a problem the boundary doesn't have.
8:14 So JSON plus OpenAPI likely wins this one. The boundary already has a contract, in exactly the shape partners can use. Second case, same company, different boundary. An order-created event, published once, consumed by Python, Java, and Go services that deploy on their own schedules. And here's question four, how long the data lives: messages wait in queues, survive rollbacks, and get replayed, meaning re-read from an archive long after they were written. So the data outlives the code that wrote it.
8:47 Yeah, and now question three again, count the versions: three languages, a pile of services, no shared deploy clock. A convention can't hold that. Even a strong JSON Schema helps only if, Every team wires it in identically. Which nobody does identically. Same missing field, 3 services: one throws, one hands back a null that explodes 3 frames later, one quietly fills in a zero. Right. A single proto file generating everyone's classes, with the field number itself carried in the bytes, is built for exactly this pressure.
9:19 Let's say a consumer team rolls back a week. The older bytes still mean what they meant, because their field identity is pinned by the contract. So Protobuf can earn its whole toolchain here. Same order data as case 1, opposite answer, because the boundary is doing the deciding. Now put those two answers side by side and the architecture that falls out is a hybrid: OpenAPI JSON at the partner edge, Protobuf on the internal event path. And that split is a perfectly normal answer. With 1 honest caveat: the translation between them is not free.
9:55 Somebody owns that mapping code, and it's a 3rd contract. It can drift, and it needs tests. For example, the field you add internally but forget to expose at the edge, or the status value that exists in one world only. The hybrid is legitimate. You just went from owning two contracts to owning three. Three. That's the part nobody puts in the design doc. And if you're already writing that mapping code, there's a shortcut somebody will float in the design review: Protobuf ships its own JSON mapping, ProtoJSON, and it sounds like the best of both worlds.
10:30 Keep the proto schema, emit readable JSON. The ProtoJSON docs are explicit about the tradeoffs, though. You keep Protobuf's type system, but the payload is less efficient than binary, the field names get serialized into every message, and the evolution guarantees are weaker. Wait, back up. What gets weaker? In binary, when an old consumer meets a field it doesn't recognize, the bytes ride along untouched. In ProtoJSON they can simply vanish, and names sitting in the payload mean renames can bite.
11:01 So a rename that binary Protobuf would shrug off can break your ProtoJSON consumers. Exactly the kind of surprise. Reach for ProtoJSON when you have a concrete interoperation need inside a Protobuf system. It won't make today's decision go away. So we land on the argument everyone expected us to open with: speed. And here's where a staff engineer leans back and says, "this is all lovely, but Protobuf is smaller and faster. Ship it." Sometimes. For your payload, in your language, on your hot path.
11:33 Which is why a benchmark of somebody else's payload on somebody else's runtime is trivia, not evidence. That's why performance sits at question six. Measure the real event on the real runtime, allocations included, not a hello-world message with three fields in it. And let the number override the contract fit only when it genuinely moves your bottleneck, because if the request spends most of its life waiting on a database, a faster parser just makes it wait sooner. If it does, that's a fine reason.
12:03 A hypothetical byte win is not. So, the six questions, one breath each. Who owns and reads the boundary? What contract already exists there? How many languages and deploy clocks have to coexist? How long does the data live, queues and archives included? Is the toolchain worth its integration tax? And does measured performance change the answer at all? Run those against your boundary and the debate gets quiet. The best format is the one whose contract matches the boundary. Thanks for listening to Learning Podcasts.