Protobuf Ch.1: Schema and Serialization

Outline

Transcript

0:00 Right now, across the globe, there are just, I mean, millions of servers out there burning through massive amounts of CPU cycles, memory, bandwidth. Oh, yeah. Constantly. Just doing one mind-numbingly repetitive task. Literally just reading the letters in the string email address over and over again. It is wild when you actually think about it. It really is. So, welcome to this deep dive. Today, we are taking a massive stack of sources. We've got Google's original 2008 protocol buffers, white papers, some recent hardcore benchmarks on binary formats versus JSON.

0:33 Right. The really high throughput stuff. Exactly. And even a few internal engineering postmortems. And our mission today is to figure out why data needs a strict shape. Yeah, we're basically exploring why the comfortable, safe world of your program's memory turns into a complete, chaotic, wild west the second your data hits the network. Absolutely. A complete, absolute wild west. And if you are a software engineer, especially if you spend your days building microservices, jumping between, say, Python and Java, this is custom built for you.

1:02 Cool. Because we are officially kicking off chapter one of our 21-chapter series, protocol buffers for any software engineer. I am so excited for this. It is going to be such a foundational journey. Because before we ever look at a single line of protocol buffer syntax, we have to unpack the underlying physics of data serialization. Right. The physics of it. Yeah. Like, we need to explore what mechanically happens the exact microsecond your data leaves the safety of your local variables, right? Yeah.

1:34 And attempts to cross a network boundary to talk to a completely different service. Written in a completely different language, usually. Exactly. Let's actually start with that feeling of safety you get inside a single code base. Oh, it's so cozy. It is. When you're writing code in just Python or, you know, just Java, you define a class or a struct. Or even just a well-documented dictionary. You have this nice little walled garden. Right. Inside that single program, the structure of your data is obvious.

2:01 You don't have to wonder what type the age field is. It's right there in your class definition. And the Python interpreter or the Java compiler is running the show. It's keeping memory organized and enforcing your rules. But the problem, the real friction arises the moment the data needs to leave the garden. So let's look at a very real scenario you probably deal with daily. OK. Let's hear it. Say you have a user service. You have a service written in Python. Right. And it needs to talk to an order service written in Java.

2:27 The Python service needs to send a user object over the network. Right. And this is where we hit what I like to call the Python dictionary to Java object gap. Yes. That gap is everything. Because a Python dictionary is essentially a hash table. It's dynamically typed. It's scattered all over memory with various PyObject pointers under the hood. It's a bit of a mess structurally. A total mess to an outsider. And Java, meanwhile, expects this really good service. It's a rigidly structured object sitting neatly in the JVM heap.

2:57 You can't just, you know, scoop up the raw bytes of a Python dictionary from your RAM. Right. Hurl them over a TCP connection and expect Java to have any clue what to do with them. Right. Because they don't share a memory layout. They don't share a runtime. They speak entirely different structural languages. Exactly. They desperately need a translator in the middle. And we call that translator serialization. You take your in-memory Python structure, convert it into an intermediate format. Like a single file.

3:23 And you send those ones the bytes. Right. Send those bytes over the wire. And then the Java side deserializes it. It reconstructs it into a native Java object. And naturally, when engineers need that universal translator, they just reach for the one everyone already uses, which is JSON. Right. JSON became the status quo because it's just incredibly accessible. I mean, you take your Python dict, you call JSON.dumps, and boom, you get a text string. Super easy. You send that text. The Java program receives it, runs it through a library like Jackson, and builds the object.

3:53 Every language supports it. And to be fair, for low traffic applications, it works perfectly fine. It does. But, and this is a big, but the sources we pulled for this deep dive paint a very different picture once you move past those low traffic apps. Oh, yeah. The hidden tax of JSON. Exactly. Engineers know JSON is text, obviously. But we rarely calculate the compounding byte cost and CPU tax of that text at scale. When you push JSON across massive microservice architectures, we see four critical breaking points.

4:25 Let's dive into the first one, which is no schema. The bag of keys. Right. JSON's flexibility is honestly its greatest liability in distributed systems. A JSON payload is fundamentally just a bag of keys and values. There is absolutely zero enforced structure in the format itself. Nothing's stopping your Python service from throwing in a completely unexpected field. Or dropping a required one. Even worse, nothing stops it from sending the string 25 when the Java service is strictly expecting the integer.

4:52 Oh, wait, I was actually reading through one of the postmortem notes in our stack about exactly this. A major tech company suffered this enormous billing outage. Let me guess, type mismatch. Bingo. The Python front end passed a transaction amount as a string instead of an integer, and the JSON parser just accepted it because, well, it's valid JSON. Right, JSON doesn't care. Exactly. But the downstream Java billing logic threw a fatal exception trying to do math on a string, and the whole service cascaded into failure, so the receiver has to be incredibly paranoid.

5:24 That paranoia is so expensive, too. Because it translates into defensive coding. Your Java order service has to validate every single piece of data at runtime. Boilerplate everywhere. Exactly. You find your team writing tedious boilerplate simply to ask, hey, does the email field exist? Is it null? Is it a string? If the amount is a string, can I safely cast it to an integer before I process this payment? You're just pushing the responsibility. You're just pushing the responsibility of data integrity onto the application logic and duplicating that code across every service that touches that payload.

5:55 Yep. Which brings us to breaking point number two, bloated size. This one hurts my soul. Because JSON is purely text, we are shipping an enormous amount of redundant data. Think about it. If your payload has a field called email address that is 13 bytes of text just for the key. Just for the name of the field. Right. If you are processing a million user objects per second, you are transmitting those same 13 bytes a million times. It's massive egress costs, wasted network bandwidth, just to tell the receiving server the name of the data it's about to read.

6:27 And consider the values themselves too, not just the keys. If you want to send the number 150 in JSON, it takes three distinct text characters, three bytes. Right. In a binary format, that same number can be represented mathematically in a single byte. Right. When you multiply that inefficiency across thousands of fields and millions of requests, well, spelling everything out in human readable text hits your infrastructure limits directly. Okay. To visualize this, using JSON is like shipping a completely assembled unboxed house down the highway.

6:58 Oh, I love that analogy. Right. It takes up a huge amount of space, it catches the wind, it's just highly inefficient to transport. What you really want is to ship highly compressed IKEA flat packs along with a tiny set of machine readable instructions. Yes. Which leads perfectly into breaking point number three. Sluggish parsing. Right. Because all that unboxed text doesn't just take up space on the network, it absolutely punishes the CPU. Punishes it. This is where we see the mechanical difference between just-in-time text parsing and ahead-of-time compilation.

7:28 Think about what a JSON parser must physically do on the CPU level. It's receiving a stream of text. Right, a stream of text. It has to scan through it character by character, watch for escape sequences. It has to maintain this complex state machine pushing onto a stack every time it sees an opening curly bracket. And popping off when it sees a closing one. Exactly. And when it encounters those three characters for the number 150, it has to run a mathematical conversion to turn those text characters back into a usable integer in memory.

7:57 It is a massive amount of branching logic. Yeah. Dynamic memory allocation. The CPU is constantly just guessing and reacting. Now compare that to a well-designed binary format. A binary parser doesn't have to scan for quotation marks or colons. There are no colons. Exactly. The structure is predictable. The machine just reads a tiny tag that basically says, hey, the next four bytes are an integer. And it simply copies those four bytes directly into memory. The byte offsets are known at compile time.

8:25 It skips almost all of the busy work. Okay, wait. I have to jump in here and play devil's advocate for a second. Go for it. We're talking about the hidden tax of JSON, but I mean, JSON being human readable is a massive feature. It's true. Like if I'm debugging my Python or Java service, I can just open my network. I can just go to the network tab, look at the payload and literally read it. I can see the bug with my own eyes. I don't need special tools to decode some binary stream. Isn't that worth the trade off for developers?

8:53 So that is the most common defense of JSON. And honestly, it holds a lot of weight during local development. Right. However, when you look at production environments at scale, the math fundamentally changes. You are paying a continuous astronomical tax in CPU cycles, memory allocation and bandwidth, for the theoretical convenience of the data. And that's the most common convenience of a human being able to read a payload. Ah, I see. Because in a high throughput microservice architecture, humans are not reading those millions of payloads per second.

9:22 Machines are. Prioritizing human readability over machine efficiency becomes this really expensive bottleneck. Machine to machine scale does not care if you can read the matrix. Exactly. They don't care at all. Which brings us to the fourth and perhaps most dangerous breaking point. And this one isn't just a technical problem with CPU cycles. It is a cultural problem. Oh yes. No single source of truth. Exactly. When the Python team and the Java team need to agree on the shape of the data they're exchanging, how do they do it in a JSON world?

9:54 They write a wiki page. They document the expected JSON structure and confluence or notion. They list the fields. They say the amount should be an integer. And then they just rely entirely on hope. Just pure hope. They hope every engineer reads it and they hope everyone keeps their code bases perfectly synced to that documentation. And predictably, drift is inevitable. I mean, the Python team adds a new field to support a feature, but they forget to update the wiki. Or the Java team refactors a type in the documentation, but the Python team misses the Slack message and doesn't update their serializer.

10:28 Right. Because there's no mechanical enforcement. You're relying on fragile human communication to maintain a strict machine contract. And the application only finds out about the mismatch at runtime, usually in production. And the JSON parser throws a fit and the logic crashes. So if human readable text formats like JSON break down under the weight of massive scale and parsing taxes and this inevitable team drift, how did the industry actually solve this? Well, according to our sources to find out, we have to look back at how Google handled it internally over two decades ago.

11:00 The early 2000s. Yeah, around 2001. Their internal systems were experiencing a massive scaling crisis. They were moving toward an architecture with thousands of distinct microservices. All communication. And they were using text based request protocols at the time. They realized that their servers were spending a terrifying percentage of their total compute power just doing string manipulation and text parsing. The unboxed houses were clogging the highway. Exactly the problem. They realized the wiki and text payload approach was mathematically unsustainable for their growth.

11:34 So they needed a paradigm shift. And what did that look like? It required three specific innovations. First, they needed a way to define data structures exactly once in a single file and have that definition automatically generate native code in C++, Java, Python, whatever language. Okay. Write once, generate everywhere. Right. Second, they needed a compact binary encoding that stripped out all the field names and text bloat, allowing for lightning fast ahead of time serialization. So getting rid of the email address string repetition?

12:07 Exactly. And third, they needed a system where the shape of the code was the same as the code was written. So the shape of the data was a literal compiler contract, not just a suggestion on a web page. And the result of that massive internal engineering effort is what we now know as protocol buffers or protobuf. Yep. Google battle tested it internally for years. They used it to encode almost all of their internal data, and then they open sourced it in 2008. And today it's become one of the most widely used serialization frameworks in the whole industry.

12:33 I mean, it is the default wire format for gRPC, and it powers service to service communication at basically 100% of the time. Right. that enforces the contract happens in step two. The compiler. Yes. You take that .proto file and you run it through the protobuf compiler, which is called protoc. And this compiler reads your definition and generates native code for your specific languages. So it spits out a pb2.py file for the Python team and a generated Java class with a builder pattern for the Java team.

13:30 Exactly, and those generated files give you ready-to-use classes with strictly typed fields. But crucially, the compiler also generates the highly optimized methods to serialize those objects into that compact binary format we mentioned. And to deserialize them back out. Right. The generated code already knows the exact schema. It knows that field number one is a string and field number two is an integer. The byte offsets are baked in. Yes, and I want to frame this for you using our core analogy from earlier.

13:58 Think of the .proto file as a class definition that lives outside any single program. Oh, that's a great way to put it. It sits right in the middle of your architecture. It acts as an absolute unbreakable contract between your Python team and your Java team. You aren't checking a wiki to see if an integer is expected. Because if you try to pass a string to an integer field in your Java code, the generated code literally will not compile. And if you try it in Python, it throws a type error immediately.

14:27 You have replaced fragile human trust with guaranteed machine-generated code. This creates perfect alignment. Since both the sending service and the receiving service are generating their serialization logic from the exact same .proto file, they are mathematically guaranteed to agree on the shape of the data. The file is the single source of truth. Right. And the receiver doesn't need to write paranoid validation code checking if a field exists or if it's the right type. The generated parser handles all of that safely before the data ever reaches your application logic.

14:58 It is a massive paradigm shift from the JSON bag of keys approach. You are moving from just-in-time guessing to ahead-of-time certainty. It really is night and day. And this perfectly sets the stage for where we are going in this series. Now that we understand the why, why we need a contract, why scale breaks text formats, and the foundational concept of this .proto file, it's time to map out the roadmap ahead for the rest of our 21 chapters. Oh, there is so much good stuff coming. We are gonna get really deep into the mechanics of this ecosystem.

15:31 We'll uncover exactly how that code generation works under the hood, what the generated Python code actually looks like, and how Java safely constructs these objects. We're also going to dismantle the binary wire format itself. The fun stuff. Oh, yeah. We'll look at how protobuf actually encodes data onto the network. We'll explore the exact byte level mechanics of tags, wire types, and variant encoding. You'll get to see exactly how it manages to compress numbers so efficiently compared to text.

15:58 And then we'll tackle schema evolution, which is just, it's critical for long-lived systems. Well, absolutely. Because how do you add or remove fields from your .proto file, Yeah. without breaking the legacy services that are still running the old compiled code? Right. We'll dive into the reserve keyword, the add, migrate, remove pattern, and how to maintain backward and forward compatibility. We also won't ignore the broader ecosystem. We have dedicated chapters comparing protobuf directly against JSON, breaking down the exact differences in size, speed, and schema safety.

16:28 We'll even compare it to other binary formats like Avro, Thrift, and FlatBuffers, so you know exactly when to choose which tool for the job. Plus, managing the data, and managing these schemas across large teams, setting up CI pipelines with tools like the Buf CLI to detect breaking changes before they ever merge. It's going to be comprehensive. But before we run, we have to walk, which brings us to what is immediately next. In chapter two, we are moving from theory to practice. Time to get hands-on.

16:57 We're going to write our very first .proto file. We'll cover the syntax, how to define messages, understanding scalar types, and that crucial concept of field numbers, which are really the absolute secret sauce to how the whole binary encoding actually works. Get ready to write some actual proto code. So to quickly summarize the journey we've taken today, we started in the safety of in-memory data structures where everything is rigidly defined by a single compiler. Right. We then crossed the perilous network gap between a Python service and a Java service, exposing how JSON's lack of schema, its text-based network bloat, its sluggish just-in-time parsing.

17:32 And its reliance on human updated wikis. Exactly how all of that makes it crack. And finally, we discovered how protocol buffers solves this by introducing a single .proto contract file that generates safe, fast cross-language code for both sides of the network boundary. Perfect summary. Now before we wrap up, I want to leave you with a provocative thought to mull over, building on this whole idea of machine contracts. OK, let's hear it. Right now, we still write most of our microservice logic ourselves.

18:04 But looking at the trajectory of our industry, if AI agents eventually start writing, deploying, and managing our microservices for us, will human readable formats like JSON go completely extinct? Oh, wow. Right? Will strict, highly compressed binary compiler contracts become the universal native tongue of the future internet simply because machines have no use for human readability? That's a really interesting way to look at where we're headed. Something to think about as you look at your own architecture today.

18:32 Thank you so much for joining us on this deep dive. We will see you in chapter two, where we write our first proto file. Thank you.