gRPC Ch.3: Protobuf Types

Outline

Transcript

0:00 Picture this right. You're writing a simple status update for a user account, and you make a typo in a string field. You type A-C-T-V-E instead of active. Just, you know, completely drop a vowel. Oh yeah. I already know exactly where this is going. Right. It slips past code review. It passes all your unit tests, mainly because the tests are just checking if a string is present. So you push the curd, and it silently breaks production. Yep. The database updates with total garbage data. Exactly. Downstream services panic.

0:27 Suddenly users are locked out, and it's literally all because of one dropped letter. It's a chilling scenario, honestly. And it happens constantly because, well, that right there is a fundamentally fragile data contract. Welcome to the deep dive, everyone. Today, we are exploring chapter three of the 14-part gRPC for any software engineer series. And for you listening, we know you're a software engineer. You write Python. You write Java. Right. So you already understand things like functions, classes, and, you know, data structures.

0:57 Exactly. And you already know those basic protobuf scaler types we covered in chapter two, your standard strings, your int 32 types, your booleans. But our mission today is to entirely kill that active typo. Yes. We need to eradicate it. We are going to explore how advanced protobuf types can model complex real-world data structures. And more importantly, how they make entire categories of these silent data bugs completely impossible at compile time. Which is really the holy grail, right? Impossible before the code even runs. Yeah. But to get there, I mean, we really have to recognize the core problem with that active typo. That a standard string field is just infinitely permissive.

1:40 Exactly. It will accept literally any combination of characters. To fix it, we have to constrain the choices. We have to lock down what the developer is even allowed to send. Right. So if we bring back our bookstore example from the last chapter, let's say we need to track the format of a book. A book might be a hardcover, it might be a paperback or an ebook. Yeah. Which is a very standard enum use case. Yeah. But if I define that as just a string field in my message, I am essentially crossing my fingers.

2:05 I'm hoping that every single Java and Python developer across the entire company spells hardcover perfectly with the exact same casing forever. Which is a terrible architecture strategy. Hope is not a strategy. Yeah. So instead, protobuf gives us enums. And an enum allows you to define a type with a strictly fixed set of named values. So you just define a book format enum? Right. You define a book format enum, and inside it, you assign a hardcover to one, paperback to two, and ebook to three. And for you listening, that maps directly to a native Python enum or a Java enum.

2:40 So the developer experience, it feels totally standard. You're just using your language's native features. Exactly. You don't even have to think about it when you're writing the backend logic. But I do want to talk about how this actually behaves under the hood. Because on the network wire, these aren't being sent as the heavy string hardcover, right? They are being efficiently transmitted as those assigned integers. That's the beauty of it. It's highly compressed on the wire. But there is a very specific, very strict rule in protobuf when you define these enums.

3:08 The very first value in your enum must always be assigned to zero. Okay. Let's unpack this. Why force the zero? I'm trying to think through the mechanics here. Think about default values. Right, right. We have to remember how proto3 handles defaults. If a field is unset by the client, the protocol defaults it to the zero value. So does this zero value rule essentially force engineers to handle edge cases explicitly, like rather than blindly trusting system defaults? You hit the nail on the head. By convention, that zero value must represent unspecified.

3:43 So your zero value should be named something like book format unspecified. Ah, I see. Imagine if you didn't do this. Like imagine if you put hardcover at zero. If a Python client creates a book message, but the developer completely forgets to set the format field, protobuf defaults it to zero. Oh, wow. And suddenly every single forgotten field silently becomes a physical hardcover book in the system database. Right. You'd be shipping physical books to people who just wanted digital downloads. That is a wildly expensive bug.

4:10 Precisely. By forcing zero to be unspecified, an unset format safely defaults to an unknown state. The receiving Java server reads it, sees unspecified, and the code is forced to say, wait, I don't know what format this is. I need to throw an error. Rather than silently corrupting the business logic, that makes a lot of sense. So once you define that book format enum, you just use it as a field type in any of your messages, and you physically cannot compile the code if you misspell hardcover. You literally can't. The compiler catches it immediately.

4:41 But let me push back a bit, or rather point out a vulnerability. Enums lock down single choices. But what happens when the data itself has layers? What do you mean? Because a flat string field for, say, a customer's shipping address is just as dangerous as our string status typo. I mean, I have my own rejects horror story of trying to parse a messy flat address string from a single text box. Oh no. Let me guess. Parsing apartment numbers. Yes. Users were typing apt4 or apartment4 or literally just hashtag4.

5:13 It was absolute chaos trying to extract the actual street address from that. I think every senior engineer has the scars from a rejects address parsing disaster. And that brings us to the exact reason we need nested messages. So how does Protobuf solve the nested data problem natively? Because obviously we aren't writing rejects here. Well, in object-oriented programming, you use object composition, right? Classes containing other classes. Protobuf does the exact same thing. You don't just have an order message with a massive list of flat string fields. Okay.

5:45 Your order message can explicitly contain a customer message as a field type. And then that customer message can contain an address message. And the address message could contain a country message. Yes. You nest them as deeply as your domain model requires. You're building a graph of types. This completely eliminates my rejects headache. The address isn't an arbitrary string anymore. It's a structured object with strict typed fields for street, city, and zip code. Exactly. You're guaranteeing the structure before it ever leaves the client.

6:13 But I want to talk about how this looks in the Protobuf itself because the sources mention you can define a message entirely inside another message's definition. You can. It's called scoping. If a structure only really makes sense in one specific context, you just scope it right inside the parent message. Ah, that is actually brilliant for namespace protection. It prevents namespace pollution at the top level of the Protobuf. Oh, absolutely. Because that is a nightmare when you have 50 different generic types named something like result item or metadata across a massive microservice architecture.

6:48 So if I have a search books response message, I can define my result item message nested directly inside it. And that nested result item is completely isolated from any other result item defined in other files. It keeps your contracts incredibly clean. Okay. But let's keep building this out because my order message still has a problem. If I have an order, I rarely just have one book. I have a shopping cart. I need collections. I need lists. And this is where we introduce the repeated keyword. In Protobuf, there is no list type or array type.

7:18 You simply write the keyword repeated before a field type to indicate that this field can hold zero, one, or many values. So the syntax is just repeated books equals three. Exactly. And just to remind everyone from our previous chapter that equals three is just the unique field number for the binary wire. It's not the amount of books. Right, right. That's just the tag. And repeated works with any type, I assume. Anything. You can have repeated string for a list of tags. You can have repeated in 32 for page numbers. You can even have repeated book format if a title is available in multiple formats.

7:52 And it maintains the order. Yes. The protocol guarantees that the order of the items is preserved exactly across the network. Okay. So repeated lists are easy enough. But what about key value pairs? In both Python and Java, we rely heavily on dynamic metadata where we don't necessarily know the exact keys ahead of time. I need a dictionary. Protobuf handles that with a map type. The syntax is simply the word map followed by angle brackets containing the key type and the value type. So map angle bracket string comma string angle bracket that creates a field that maps string keys to string values. Okay. So that functions exactly like a Python dict or Java hash map.

8:29 But wait, the sources mention a strict rule here. Map keys must always be a scalar type. Why? Why can't I use a complex nested message as a map key? I can do that in Java if I write a custom hash function. You can do it in Java, yes. But remember, Protobuf isn't just running in Java. It's a cross-language serialization format. Okay. So it has to work everywhere. Exactly. Under the hood, Protobuf actually serializes maps as a repeated list of hidden key value message pairs. If you allowed complex nested objects as keys, the wire format would become massively complicated. That sounds messy.

9:04 It is. And more importantly, hashing those complex objects consistently across entirely different languages like Python, Java, Go, and C++Z would be a deterministic nightmare. That makes total sense. So restricting keys to scalers like strings or integers guarantees that every language hashes and deserializes the map identically. But what about the map values? Are they restricted to? Not at all. The map values can be absolutely anything. Wait, let me make sure I have this straight. I can have a repeated list of nested messages.

9:34 Yes. And those nested messages can contain maps. Yep. And the values inside those maps could be enums. What's fascinating here is how seamlessly these Protobuf constructs map directly to the native collections you use every day. You aren't learning a strange esoteric paradigm. No, it really just feels like standard programming. Exactly. You are using plain text syntax to define the exact same nested hash maps and lists you write in Python or Java. But now they are baked into a strict cross-language typing system.

10:04 It is so powerful, but I do have a complaint. Uh-oh. Let's hear it. As an engineer, there are certain concepts that I have to implement in literally every single software project I touch. Time. Durations. If I have to build a custom nested message for a timestamp every single time I start a new microservice, I am going to lose my mind. Well, you shouldn't have to reinvent the wheel for universal concepts. And the creators of GRPC entirely agree with you. That is why Protobuf provides a suite of what are called well-known types.

10:37 Here's where it gets really interesting. I, uh, I actually didn't realize this initially, but Protobuf isn't just a syntax. It actually comes with its own standard library. It does. These well-known types are provided by Google in the google.protobuf package. You just import them at the top of your proto files exactly like you would import a standard library module in your Python or Java code. So let's talk about timestamp because using a standardized timestamp is huge. If I don't use it, I am guaranteed to get into a drawn-out debate with the front-end team over how to format time. Oh, the classic time debate.

11:09 Yes. Should it be an integer of seconds since the Unix epoch? Should it be milliseconds? Should it be an iso string? And those debates usually end in parsing bugs. Think about the internal mechanics of a timestamp. If you just use an int32 for seconds, you hit the year 2038 problem, where the integer simply overflows. Yeah, the systems just crash. But if you use a string, you have time zone ambiguity and massive parsing overhead. The well-known timestamp type solves this by internally modeling time as two fields.

11:39 An int64 for seconds and an int32 for nanoseconds. It gives you massive range and incredible precision. And more importantly, the debate is just over. The contract explicitly says this is a Google protobuf timestamp. I don't have to write custom conversion logic in my Java service to handle whatever weird time format the Python service decided to spit out. Exactly. It just works natively. And it's not just timestamp. There's duration, which represents an interval of time, like five seconds or 10 milliseconds.

12:06 You use that exact same internal seconds in nano's structure. Now, I saw another well-known type in the sources that actually confused me a bit. The any type. The documentation says it allows you to hold arbitrary message types. Ah, yes. The any type. Wait a minute, though. We just spent this entire deep dive talking about how strict typing is the holy grail. Why would a strict typing system offer an any escape hatch? That just sounds like a back door to bad architecture. It does sound counterintuitive at first.

12:35 I totally agree. But there are rare, very specific architectural patterns where you genuinely need polymorphism. Imagine you are building a generic logging service. Okay. It needs to accept log events from 50 different microservices. And each event might have a totally different custom data payload attached to it. Okay, I see the problem. The logging service can't import the proto definitions for every single service in the company. That would be an insane dependency web. Exactly. So how do you handle that safely?

13:05 The any type isn't just a blind byte array. Under the hood, it packs the binary serialized data, yes, but it also attaches a type URL. A type URL? Yeah. It's a string that explicitly identifies what the original message type was. Ah. So it's dynamic, but it's still self-describing. The receiving service can look at the type URL, check if it actually knows how to process that specific type, and safely unpack it without just guessing what the bytes mean. Right. If we connect this to the bigger picture, using these well-known types, even the complex ones like any, guarantees data consistency across entirely different teams and languages.

13:43 If your Python backend team and your Java Android team both use the timestamp type, they literally cannot misinterpret the time data. The serialization and deserialization are handled flawlessly by the generated code. So what does this all mean? Let's zoom out and summarize. You, listening to this, you now have a highly mature vocabulary for defining data. You're not just writing generic strings anymore. Right. You have enums complete with the zero value rule to lock down specific choices and avoid those string typos.

14:11 You have nested messages with proper scoping to build deep, complex objects instead of relying on fragile rejects parsing. You have repeated fields and maps, constrained safely for the wire to handle your lists and dictionaries. And finally, you have a built-in standard library of well-known types to standardly tackle universal headaches like time and duration. And the most important part of everything we discussed today is where this all lives. These constructs are defined in plain text proto files.

14:39 You check them in diversion control. Both sides of a network call, the client and the server must agree on this exact structure. It catches the bugs at compile time. That silent active typo we started the show with, it can never make it to production now. A developer physically cannot compile their Java or Python code if they assign a string to an enum field. Or, you know, if they try to pass a flat string where a nested address object is required. The protocol simply won't allow it. But as rigorous and robust as these plain text proto files are, well, right now they are still just blueprints.

15:13 They are architectural diagrams. Which perfectly sets up where we are heading next. In chapter four, we're going to dive into the compiler itself. We'll look at how Protok takes these plain text blueprint files and magically generates the actual, usable, fully typed classes in Python and Java. We'll look at the generated code itself. I really can't wait for that one. Seeing the generated code is where it all finally clicks. But before we wrap up, we want to leave you with something to chew on. We've spent a lot of time focusing on how these rigorous data contracts save the machine from bugs.

15:46 But think about the human element for a second. If your data contracts are this rigorously typed, heavily version controlled, and completely language agnostic, how might that completely change the way your front end and back end teams negotiate features before a single line of Python or Java is even written? Does it turn what is usually a chaotic API negotiation into an actual collaborative design session? It's definitely something to think about as you start modeling your own data. Until next time, keep your data tight, use those enums, and watch out for those silent string typos.