Protobuf Ch.3: Code Generation
Outline
- 0:00 Opening Hook: One File, Eleven Languages
- 0:49 Chapter 3 Introduction
- 1:39 What is protoc?
- 1:53 Installing protoc (Homebrew, apt, GitHub)
- 2:32 The Compilation Workflow (Three Inputs)
- 4:00 Compile-Time Validation
- 5:13 The Scaling Challenge
- 5:42 Two-Tier Plugin System
- 6:00 Tier 1: Built-in Generators
- 6:24 Tier 2: External Plugin Architecture
- 7:17 How Plugins Communicate (stdin/stdout + Protobuf)
- 9:41 Generated Code: Python vs Java
- 10:01 Python: Dynamic, Mutable, _pb2.py
- 11:18 Java: Immutable Builder Pattern
- 13:07 The Golden Rule: Never Edit Generated Code
- 14:10 Composition Pattern
- 14:40 Build System Integration (Makefile, Gradle, Bazel)
- 16:52 Cross-Service Contract Enforcement
- 17:18 Chapter Summary
- 17:47 Next: Chapter 4 (Binary Wire Format)
Transcript
0:00 You write one single text file, just one. And from that solitary language agnostic blueprint, you can generate working, heavily optimized code in 11 officially supported programming languages. Right, with zero handwritten serialization code, like absolutely zero. Exactly. You define your data once, and suddenly your Python code has a method to serialize it. Your Java code gets this perfectly designed builder class, and your Go code is ready to marshal and unmarshal, all from the exact same source.
0:34 It's a massive paradigm shift, honestly. When you're dealing with traditional formats, you're always writing that bridge manually between the data on the wire and the objects in your computer's memory. Yeah, which is just tedious. Oh, it's awful. But with this approach, that bridge is built for you automatically and in the native idioms of whatever language you happen to be using. Well, welcome back to the Deep Dive, everyone. For our resident Python and Java engineer listening right now, consider this chapter three of your journey into protocol buffers.
1:00 We're glad to have you back. So in our previous sessions, we established why data needs a schema when it crosses boundaries. And you already know how to write the protofile itself. You know, the syntax declaration, the messages, the scalar types, and why field numbers are permanent wire identifiers. Right. You essentially hold the contract in your hands at this point. Yeah. But a contract on paper doesn't actually move data across a network, does it? No, it doesn't. A contract is just an intention.
1:29 So today, our mission is exploring the engine that turns that intention into reality. We need to look at the protocol buffer compiler. Right. We're talking about protoc. It rhymes with rock. Just say it as one smooth word. You don't spell it out. Exactly. And this compiler, protoc, is what elevates protobuf from just being a neat format into a really robust multi-language platform. Because if we think about JSON or XML, I mean, those are just formats, right? You still have to write the code yourself or import some heavy third-party library to parse them into your Python dictionaries or Java objects.
2:00 Yeah. Whereas protobuf, through the protoc compiler, just hands you the native objects directly. So to understand that platform paradigm, we have to look at the tool acting as the gatekeeper. Getting your hands on protoc is surprisingly lightweight, actually. It really is. It ships as a standalone binary maintained by the protocol buffers project on GitHub. So there are no massive IDE installations required here. My goodness. Yeah. You just download the precompiled release for your operating system, extract it, and add it to your path.
2:30 Or, you know, if you're on macOS, Homebrew makes it trivial. It's just a formula called protobuf. And on Ubuntu or Debian, you just use the apt package manager and grab the protobuf compiler package. Easy. Yep. And then a quick protoc dash dash version check in the terminal confirms you're ready to roll. Okay. So let's say I have a file called user.proto, and it's sitting in a directory literally just called protos. How do I actually invoke this thing? So running it follows a very strict but straightforward three-directive workflow.
2:59 The compiler essentially needs to know three specific things. First, you pass the double dash proto underscore path flag pointing to that protos directory. Okay. And what does that do? That tells the compiler where its root directory is for resolving any import statements across your files. Oh, right. Because in a large microservices architecture, you might have like hundreds of protofiles importing common types from each other. Exactly. So protoc needs this sort of virtual file system root to find everything.
3:29 Makes sense. Okay. So it knows where my files are, and it knows how to resolve the imports, but it still needs to know what language I want to generate. I assume there's a flag for that. This is where the magic happens. For Python, you pass a flag that is double dash pycon underscore out equals followed by your target directory. Got it. For Java, it's double dash Java underscore out. C plus plus is CPP underscore out. Go is go underscore out. Oh, I see. It's a perfectly consistent pattern. Just the language name, an underscore, and the word out.
4:00 You got it. And the final piece is just passing the path to the user dot protofile itself. But, you know, the most important part of this entire workflow is what happens the millisecond you hit enter. Right. Because it doesn't just blindly write code. It validates. I actually like to think of running Protalk like submitting a manuscript to an aggressively strict editor. That's a great way to put it. Yeah. It completely refuses to go to the printing press, meaning it won't generate a single line of code until every single typo and logic gap is fixed.
4:32 Exactly. It reads the protofile, validates the syntax, and resolves all the imports. So if you missed a semicolon, or if you accidentally reused a field number. Or if you referenced a type that doesn't exist anywhere in your import path, Protalk just halts everything. It prints an error message pointing exactly to the problem. Which is such a massive relief, honestly. It saves you from those disastrous runtime errors. Like, imagine finding out you have a typo in a JSON key while the code is running in production at 3 a.m.
5:00 Ah, that's a nightmare. Right. With Protalk, that entire class of error is just caught at compile time. The editor simply rejects the manuscript before it ever sees the light of day. And that strictness is really what enables the sheer scale of the platform. But thinking about that scale actually brings up a significant logical puzzle. Oh, like what? Well, we're talking about one standalone binary. How does one binary possibly keep up with the ever-changing landscapes of dozens of different programming languages?
5:30 Oh, wow, yeah. Language idioms change. New languages emerge. Standard libraries get updated. A single executable cannot possibly contain the generation logic for the entire software industry. Right. Keeping up with just Python and Java updates is a full-time job for entire engineering organizations. So how does Protalk pull it off? Through an incredibly elegant two-tier translation system. So tier one consists of the built-in generators. The compiler itself natively ships with code generators for a core set of historically significant languages at Google.
6:00 And what are those? That's C+++, Java, Python, C Sharp, Ruby, Kotlin, Objective-C, and PHP. Okay. So for those languages, which obviously includes our listeners, Python and Java stacks, there's zero extra setup. You just pass that out flag we talked about and you're done. Yep. Completely seamless. But what about a language like Go or TypeScript? Do we just have to sit around and wait for the core team to update the main Protalk binary to support them? Not at all. That is where tier two comes in.
6:28 The plugin system, Protalk uses external plugins that follow a really strict naming convention. It's Protalk Gen followed by the language name. Oh, I see. So for Go, the executable has to be named Protalk Gen Go, pronounced Protalk Gen Go. And for TypeScript, it's Protalk Gens, pronounced Protalk Gen T-S. Exactly. So when you pass the double dash do underscore out flag, Protalk searches your system's path for that specific executable. But wait, I want to dig into how they communicate. Does Protalk just hand the raw text of the user.proto file over to the Go plugin?
7:02 Like, why not just have the Go plugin read the file from the hard drive itself? Well, think about it. If Protalk just handed over the raw text file, the Go plugin would have to write its own complex text parser. Oh, right. It would have to scan for curly braces, semicolons, and comments. Exactly. That is a nightmare to maintain. You'd have dozens of plugins all trying to parse the same syntax independently, which inevitably leads to bugs. So how do they have? Instead, Protalk does all the heavy lifting.
7:30 It reads the text, validates it, and converts it into a clean, pre-digested data structure called an abstract syntax tree, or AST. Okay. Think of Protalk as the master architect who draws the master blueprint. Instead of trying to build the plumbing and the electricity themselves, the architect hands a highly detailed standardized 3D model, the AST, to the specialized contractors. That's spot on. Right. Because the Go plugin only looks at what it needs to write ghost trucks. Yeah. The Python plugin only looks for what it needs.
8:01 They don't have to interpret the messy raw sketches. That is the exact dynamic. And the way Protalk hands that 3D model over to the contractors is a beautiful piece of engineering design. It sends that parsed abstract syntax tree to the plugin over standard input. Okay. But the payload it sends is a protocol buffer message. Wait, really? Oh, I love how wonderfully meta that is. Protobuf is using protobuf to generate protobuf code. Exactly. The compiler serializes the parsed schema into a binary protobuf message, pipes it to the plugin's standard input, and the plugin deserializes it, generates the target code, and sends the text back over standard output.
8:41 That's the ultimate eat-your-own-dog-food scenario right there. It really is. And because this protocol between Protalk and its plugins is so well-defined, anyone can write a plugin for any language. You just need to be able to read a protobuf message from standard input and write text to standard output. That's kind of a classic Unix philosophy applied to code generation. Yeah. And this is the secret sauce allowing a relatively small core team to support a massive global ecosystem. They maintain the core compiler and the plugin protocol.
9:11 The community handles the rest. All right. So the manuscript is approved. The architect handed off the blueprint, and the plugins did their job. What do we actually get? Let's look at the actual output. Yeah. Let's tailor this to what our listener is staring at in their IDE right now, Python and Java. Okay. Let's start with Python. When you run Protalk on your user.proto file, you're going to get a file named user underscore pb2.py. Whoa. Hold on. User underscore pb2.py. We're writing Proto3 syntax right.
9:44 Why does it say pb2? Are we accidentally using an outdated version of the compiler? No, no, you're not. But that suffix throws everyone off the first time they see it. But that did. It causes a lot of confusion, but it's purely a historical artifact. Back in the day, Google had an internal version known as Proto1. When they open-sourced Proto2, they had to prevent naming collisions in their massive internal monorupo. Ah, so they couldn't have new Proto2 files clashing with legacy Proto1 files. Right.
10:10 So they forced the underscore pb2 suffix for the new generation. And then by the time Proto3 came around. By the time Proto3 was released, the Python ecosystem was so heavily dependent on that specific suffix and millions of import statements that removing it would have just broken the world. Oh, wow. So whether you're writing Proto2 or Proto3 syntax today, the generated Python file will always have that underscore pb2 suffix. Good to know. Don't panic. It's just a naming quirk, not a version mismatch.
10:36 Hmm. Now, inside that Python file, what does the developer actually interact with? It's very dynamic, fitting the Python ecosystem perfectly. You get a Python class for each message you defined. Okay. So for your user message with name, email, age, and eyes active fields, you get a user class with those exact fields as attributes. Exactly. You just instantiate it and you can set the fields directly. It's mutable. You can change the user's name whenever you want, just like a normal Python object.
11:04 And when it's time to send it over the wire? You just call serialized to string on the object and it hands you the binary bytes. That's awesome. And to go the other way, I'm guessing you create an empty instance and call parse from string, pass in the binary data, and the object populates itself. You got it. It feels very Pythonic. Quick, mutable, direct. Nice. Now, contrast that with the Java output. I imagine the story there is pretty different. Vastly different, yeah. And it reflects Java's underlying design philosophy.
11:32 In Java, the generated code is strictly immutable. Wait, I'm picturing this generated Java code in my head. And if there are no setter methods on the user object, how am I supposed to populate the data when it arrives in pieces over the network? Right. Or when I'm pulling fields from a database? Am I supposed to pass 50 arguments into a giant constructor? No. No giant constructors required. Yeah. You use a nested builder class. Okay. You call user.newBuilder. This gives you a mutable builder object.
12:01 Then you chain setter methods on that builder, set name, set email, setage. Ah, I see. And when you've populated all the fields, you call .build. That method returns the final immutable user object. It's exactly like a car assembly line. You use the builder to walk down the line. You specify the red paint job, the V8 engine, the leather interior. You can swap things around while it's on the line, change your mind about the paint color. But the moment you press .build, that car rolls off the line.
12:28 The doors are locked. The paint is dry. You cannot change the color anymore. It is a finished immutable object. That analogy captures the mechanism perfectly. Java developers highly value thread safety and predictability. In a highly concurrent server environment, passing a mutable object between threads is a race condition nightmare. Oh, for sure. So immutability solves this. Fluent interfaces, chaining methods, and immutable data transfer objects are standard idioms there. And once you have that immutable user object.
12:59 Then we'll call .toByteArray to serialize it or user.parseFrom to deserialize incoming bytes. It's incredible that a single protofile produces such philosophically different yet perfectly native feeling code for both languages. Python gets its dynamic mutability. Java gets its thread safe builders. It really is brilliant. But, you know, seeing all these beautifully crafted files appear in your project folder triggers a very dangerous developer instinct. Oh, yes. The overwhelming urge to tweak the generated code.
13:29 We all know that one engineer who thinks they can outsmart the compiler. They see that Python class and think, oh, I'll just hard code a quick rejects check into user underscore pb2.py to validate the email format. And then Friday at 5 p.m., someone updates the schema, protoc runs, wipes out the rejects, and brings down production. Yeah, we must violently disabuse you of this notion right now. This brings us to the golden rule of protocol buffers. Do not edit generated code. Never. Not even to add a helpful comment.
13:58 Really? Not even a comment? Nope. The generated files are completely ephemeral. Every single time you change your protofile and rerun protoc, the target files are overwritten entirely from scratch. Any manual edit, any custom logic you cleverly placed in there is instantly obliterated. I actually compare generated protobuf code to compiled byte code, like a javadot class file. Yeah. You would never open compiled by code in a hex editor to manually tweak a loop, right? You treat it as an untouchable intermediate artifact.
14:28 The same logic applies here. You own the protofile. The protoc compiler owns the generated code. If you need custom behavior, like the email validation regex you mentioned, you write a separate wrapper class or use composition in your own code base. You import the generated protobuf class into your custom code and you build your logic around it. You keep the data container entirely separate from your business logic. Which leads to a much cleaner architecture anyway. But speaking of avoiding manual work, running protoc by hand in the terminal every time you update a schema gets tedious fast.
15:02 No one actually does that in a real production environment, do they? No. Running it manually is great for learning, but in practice, teams integrate protoc directly into their build systems so the code generation happens completely automatically. How does that look in practice for our Java and Python setups? Well, it depends on your tooling, but the pattern is universal. If you're using something classic like makefiles, you might have a target that watches for changes to .protofiles and triggers protoc.
15:28 What about Java specifically? In the Java world, Gradle has fantastic, officially supported plugins that hook protobuf compilation directly into the Java build lifecycle. Oh, nice. Yeah. Before your Java code even attempts to compile, Gradle ensures the protofiles are compiled first so the generated classes are always available in the class path. And we absolutely have to mention, Bazel pronounced Bazel Google's open source build system. Since protobuf and Bazel both came out of Google, they work together flawlessly.
15:59 Bazel has first class protobuf support. It really does. Bazel handles this with incredible efficiency through its build graph. You define a protolibrary rule, and Bazel automatically knows how to generate code for whatever target languages depend on it. So it's smart about caching too, right? Extremely smart. It caches the generated abstract syntax tree. If you change a downstream Java file, Bazel knows the proto hasn't changed, so it uses the cache generated files. But if you update the schema?
16:26 It instantly knows to rebuild both the Java and Python stubs before compiling the rest of your application. The build system ensures you are never out of sync with your schema. Wow. So to summarize our journey today, we started with a single source of truth, that one text file. We passed it through the strict, unforgiving validation of the Protoc compiler. Right. We saw how it uses a brilliant two-tier system, communicating via its own protobuf messages over standard input, to farm out work to language plugins without having to write a million text parsers.
16:58 Exactly. And the result is perfectly tailored, untouchable, idiomatic code for both Python and Java, all wired up to run automatically in your build system. It's the machinery that makes the whole ecosystem viable. But up until now, we've only been looking at the surface, we've defined the schema, and we've generated the native code. Which means it's time to actually push some data. Next time in Chapter 4, we are diving below the surface of the generated code to look at the actual wire. Oh, that's going to be fun.
17:26 Yeah. We're going to see how those field numbers we talked about translate into binary bytes. Yeah. And we will finally unpack why the mysterious variant encoding makes protobufs so much dramatically more compact than text formats like JSON. That is where the massive performance claims of protobuf actually get proven. I can't wait. But before we go, I want to leave you with a final lingering thought based on that plugin architecture we discussed. Think about the power of what protoc is actually doing.
17:54 Okay. It parses your entire schema, figures out all the types, the fields, the relationships, and then simply streams that rich abstract syntax tree over standard input to any plugin that asks for it. The architect is handing out the 3D model to anyone who wants it. Exactly. Who says a plugin has to generate code? If you have all that structured data about your system's entities, could you write a custom plugin that reads your protofiles and automatically generates beautiful interactive API documentation?
18:25 Oh, that's an interesting idea. Or maybe a plugin that reads your messages and automatically generates the exact SQL schemas needed to store them in your database. You have the ultimate blueprint and protoc is handing it to you on a silver platter. What else could you build? Think about it. We'll see you next time.