Protobuf Ch.4: 87% Smaller Than JSON

Outline

Transcript

0:00 A Boolean true costs exactly two bytes total, the number 150, two bytes. An entire field tag plus a small integer value fits perfectly into three bytes. Which is just wildly small. Right. Let me show you exactly what happens to your data byte by byte. Because when you compare that to JSON, like just the string key is active alone eats up 11 characters. Yeah, 11 bytes. Exactly. And that is before your parser even reaches the colon, the whitespace,ace, or the actual value. Which, you know, translates directly to CPU cycles, memory allocation, network bandwidth.

0:34 Yeah. When you multiply that text parsing overhead across, say, millions of messages a second. The waste becomes astronomical. It really does. I mean, this is the core reason protocol buffers absolutely dominate in high performance distributed systems. Today, our mission is to look completely under the hood. We are, like, stripping away the generated code, ignoring the getters and setters, and looking at the exact binary encoding that travels over the wire. Right. We're moving from the human-readable contract right straight to the machine -readable reality.

1:04 Because it's fascinating down there. It is. Protobuf achieves this unbelievable compactness through basically three intertwined, highly optimized concepts. Tags, wire types, and variance. Okay. So let's start with the tag, because this is the fundamental mechanism that lets us eliminate those bulky JSON keys entirely. Right. Like, you have a field-like email address. In JSON, every single time you send a message, you are sending the characters E-M-A-I-L underscore A-D-R-E-S. That is 13 bytes of redundant text just to tell the receiving server what the next piece of data actually is.

1:39 Exactly. But the wire format for Protobuf, it does not care about your field names. The schema your .proto file holds the names for human convenience. The wire holds only numbers. Yeah. Every serialized field is simply a pag followed by a value. That's it. If you assigned email address as field number two in your schema, the wire only transmits a small integer tag representing field two. And then the receiver's generated code knows that field two maps to email address, so it just, like, stitches the human meaning back together on the other side.

2:10 Exactly. But the tag itself isn't just the field number. It's actually, well, it's a bitwise masterpiece, honestly. Totally. It packs two distinct pieces of information into a single integer. It tells the decoder the field number, and it also includes the wire type. Right, which is a tiny number indicating the shape of the data that follows. Hmm. So the encoder takes your field number and shifts it left by three bits. Like, if you visualize a row of binary slots, it literally slides the field number to the left, leaving the bottom three bits empty.

2:42 Okay. Then it uses a bitwise or R operation to slot the wire type directly into those lowest three empty slots. Think of it like packing two pieces of information onto a single tiny shipping label. I like that analogy. Right. You physically slide the destination address to the left to make a little blank space on the right margin, and then you stamp a tiny cod for the shipping method like ground or air right into that empty space. One label, two critical pieces of data. Exactly. And the decoder just reverses the process when the package arrives.

3:15 It reads the tag, looks at the bottom three bits to find the shipping method, the wire type, and then shifts the rest of the bits back to the right to get the actual field number. But the size of the tag itself creates a really important optimization rule that, you know, every engineer designing an API needs to memorize. Yes, the rule of 15. Right. Because of the shifting works, tags for field numbers 1 through 15 fit perfectly into a single byte. And the moment you hit field number 16 and all the way up to 2047, I think the tag suddenly requires two bytes.

3:44 It overflows the single byte capacity. Which means you need to put your most frequently used fields, your core identifiers, in the 1 to 15 range in your .proto files. Yeah. If you put a rarely used or, like, a deprecated field in slot 2, you are burning a premium one byte tag for absolutely no reason. So we know the tag is an integer, and we know it's compact, but how is that integer actually encoded on the wire? Right. And for that matter, how do we make sure small numbers don't waste massive amounts of space?

4:15 Like, if you declare an int64 field in your schema, but it only holds the number 5. Sending 64 bits of mostly zeros is a tragic waste of network capacity. Which brings us to the wider types, those tiny codes stamped into the bottom three bits of the tag we just talked about. Yeah. There are only four active wire types you will see in practice today. Okay. Wire type 0 is variant. Wire type 1 is exactly 8 bytes, used for fixed 64-bit types. Wire type 2 is length delimited. And wire type 5 is exactly 4 bytes, for fixed 32-bit types.

4:47 Wait, you skipped 3 and 4? Good catch. They belong to a deprecated feature called groups from the very early days of protobuf. You won't see them used today, but those numbers remain permanently reserved, so modern parsers don't break if they, you know, encounter legacy data. Makes sense. Let's focus on wire type 0, the variant. Variable length integers are the magic trick that shrinks our numbers down. They really are. Instead of a fixed 64-bit box for every number, the box grows or shrinks depending on the size of the number inside it.

5:16 The rule for varints is elegantly simple. You take the binary representation of a number and break it into chunks of 7 bits. So 7 bits, okay. And store those chunks least significant first. But a byte has 8 bits, not 7. That 8th bit, the highest bit in every byte, is repurposed as a continuation flag. Oh, I see. Yeah, so if that 8th bit is a 1, it tells the decoder, keep reading, there are more bytes coming. If it's a 0, it means stop, this is the last byte of the number. Let's do a literal step-by-step walkthrough of this.

5:42 Let's trace the number 150. Okay, let's do it. So in binary, 150 is 100, 0, 1, 0, 1, 1, 0. Right. And we need to split that from the bottom. Yep. So the lower 7 bits are 0, 0, 1, 0, 1, 1, 0. And the upper chunk is just 0, 0, 0, 0, 0, 0, 1. So to make this concrete for you listening, imagine your number is a giant stack of coins. And you are packing them into standard coin sleeves. I love this. But the catch is, each sleeve can only hold 7 coins. You can't fit 150 into a single 7-coin sleeve. You have to split the stack.

6:15 So you take the first 7 coins, the lower chunk we just talked about, and put them in the first sleeve. Right. And because you still have more coins left over, you take a bright red sticker, the continuation flag, and slap it on the top of that first sleeve. So that first byte gets the lower chunk with the 1 flag. So it looks like 100, 1, 0, 1, 0. Exactly. The decoder sees that sticker and knows it has to grab the next sleeve, too. Then you take your remaining coins, the upper chunk, which fits easily into the second sleeve.

6:40 And since there are no more coins left in your stack, you do not put a red sticker on this one. You set that highest bit to 0. So the second byte gets the upper chunk with a 0 flag, making it 0, 0, 0, 0, 0, 0, 0, 1. Total size, 2 bytes. The decoder reads the first sleeve, sees the sticker, reads the second sleeve, sees no sticker, and stops. Then it just dumps the coins out and recombines them into the number 150. Better yet, any number from 0 to 127 fits perfectly in a single sleeve, right? A single byte.

7:09 Because 7 bits cover exactly that range. Which means a Boolean true, which evaluates to 1 under the hood, costs 1 byte for its tag, and 1 byte for its varint value of 1. 2 bytes for the whole field. It's incredibly tight. But varints are brilliant for mathematical numbers. Data doesn't always have a fixed mathematical size, though. What happens when you are sending a string of text or a whole nested message where the size could be, well, anything? Right. A varint can't handle a string of characters.

7:33 That is where wire type 2 comes in. Length Delimited Encoding, or LEN for short. You're LEN. Got it. It's elegantly simple. It consists of three parts. The tag, followed by a varint that explicitly states the byte length of the payload, followed by the raw bytes themselves. So if I want to send the string hello, the wire format is just the tag, then a varint with the value 5, because hello is 5 bytes long, and then the actual 5 bytes of text. Exactly. And embedded messages work the exact same way.

8:03 If you have a complex inner message, it is serialized into raw bytes and then prefixed with its total length. Wait, the structure here is actually a massive superpower for the decoder. How so? Well, because it's length delimited, a decoder doesn't have to parse the payload to know where it ends. It can just look at the length varint and jump completely over the payload without needing to parse it. Yes. You just touched on one of Protobuf's most powerful features, unknown field skipping. And it works for all the wire types, not just strings.

8:30 The wire type itself tells the decoder exactly how to skip data it doesn't recognize. So if the tag says wire type 0, the decoder just keeps reading sleeves until it hits one without a red sticker, a continuation flag of 0, and then it just skips it. Right. If it's wire type 1, it blindly skips 8 bytes. If it's wire type 5, it skips 4 bytes. And if it's wire type 2, it reads the length varint and jumps ahead exactly that many bytes. Exactly. In every single case, the decoder knows how to safely bypass data it doesn't understand without crashing.

9:01 This is the absolute foundation of schema evolution. If an older microservice receives a message with a brand new field it has never seen before, it doesn't throw a parsing error like JSON often does. No. It checks the wire type, skips the correct number of bytes, and processes the rest of the message perfectly. It is future-proofing built directly into the lowest level of the binary format. Which allows you to decouple deployments so service A and service B don't have to be updated at the exact same time.

9:29 Exactly. Let's put this all together and build a whole message to see the final weight. Imagine a simple message with three fields. Okay. Laid on me. Field 1 is an int32 user ID set to 42. Field 2 is a string name set to eta. Field 3 is a Boolean active set to true. Let's do the math on the wire. Okay. The user ID is field 1 wire type 0 for a varint. The tag computes to 8, which fits in one byte. The value 42 also fits in one varint byte. That is two bytes total. Two bytes. Next, the name string, eta.

10:02 Field 2 wire type 2 for length delimited. The tag computes to 18, which is one byte. The length of Ada is 3, which fits in one varint byte, and the string itself is three bytes. 1 plus 1 plus 3 gives us five bytes total for the string field. Finally, the active Boolean. Field 3 wire type 0. The tag computes to 24, one byte. The value true is encoded as a varint 1, which is one byte. Two bytes total. So two bytes for the ID, five bytes for the name, two bytes for the Boolean, nine bytes. A nine byte masterpiece.

10:30 Honestly, it is. If you wrote that equivalent JSON braces, quotes around user ID, name, active, colons, commas, the string Ada, the word true, you are looking at roughly 60 to 70 bytes. We're talking about a 7 to 1 compression ratio for a tiny everyday message. And that efficiency gap only widens as messages carry more fields and longer, more descriptive field names in the schema. Because, again, the wire format never includes those long names. But wait, what happens if a field isn't explicitly set by the engineer?

11:03 Say the active Boolean wasn't true. Say it was false, which is its default state. Does it still take two bytes? Ah, this is a crucial optimization known as the zero byte trick. If a field is set to its default value false for a Boolean, zero for an integer, an empty string for text protobuf does not encode it on the wire at all. It just completely drops it from the payload. It costs zero bytes. The serialized message omits it entirely. When the decoder is reading the stream on the other side and finishes, it simply notices that field number three never appeared.

11:32 And it checks its local schema, sees the field exists, and automatically fills in the default value of false for you in memory. You got it. That is incredibly clever. A message could have 50 fields, but if 40 of them are sitting at their default values, you only pay the network cost for the 10 that actually contain unique data. It is a massive architectural advantage over JSON, where every key usually appears in the payload, regardless of whether it's holding a default value or not. Hold on, though.

12:02 I'm thinking about the edge cases with the varint logic. If the highest bit of every varint byte is always reserved as a continuation flag, what happens to negative numbers? Ah. Because negative numbers in binary use the highest bit, the sign bit, to indicate they are negative. Doesn't splitting them into chunks completely break the math? If I send negative one in a standard int32 field, what happens? You've spotted the trap. It absolutely wreaks havoc on the varint logic. I knew it. If you declare a field as a standard N32 and send negative one, protobuf actually sine extends that number all the way out to 64 bits to preserve the negative value across different systems.

12:41 Because negative numbers in twos complement binary are mostly ones. Exactly. A negative one is a solid wall of ones. When you break 64 bits of ones into 7-bit chunks, you end up with 10 chunks. Wait, really? Yes. Yeah. That single negative one encodes as a massive 10-byte varint. It is the absolute worst-case scenario. How do we fix that? By changing your schema declaration. Instead of N32, you declare the field is sint32. With an S. The S stands for signed. When the compiler sees sint32, it triggers a mathematical transformation called zigzag encoding before it applies the variant chunks.

13:16 Zigzag encoding? How does that stop the 10-byte explosion? Think of a standard number line. With zero in the middle, positives to the right, negatives to the left. Zigzag encoding folds that number line in half right at the zero. Okay, I'm picturing it. It interleaves the negative and positive numbers, like the teeth of his number. Zero stays zero. Negative one becomes positive one. Positive one becomes two. Negative two becomes three. Positive two becomes four. It zigzags back and forth. I see.

13:44 It maps every negative number to a small positive integer. So negative one gets mapped to positive one. And positive one fits perfectly into a single varint byte. It turns a disastrous 10-byte payload back into a single byte. That's genius. But there is a catch. The decoder needs to know whether to apply zigzag decoding or standard varint decoding when it reads that byte. But the wire type for both is just zero. The tag just says this is a varint. Exactly. The wire format only tells you the shape of the container.

14:14 It tells you it's a varint. It doesn't tell you if it's signed, unsigned, a boolean, or an enum. This brings us to a fundamental philosophical trade-off in protobuf. Semantic opacity. Meaning the structure is visible, but the meaning is hidden. Right. If I intercept that nine-byte stream without having your .proto file, I can parse it structurally. I can see that field one has a varint value of 42. But I have absolutely no idea if field one is a user ID, a temperature reading, or an age. Right. Structure without the schema tells you the shapes.

14:48 The schema tells you what they mean. Which is a complete departure from JSON, where usrid 42 tells anyone intercepting the payload exactly what the data represents right there in plain text. So protobuf sacrifices human-readable keys and context on the wire to achieve this blistering, fast, dense numerical encoding. It is a deliberate, calculated trade-off. And because those numerical tags uniquely identify everything on the wire, there's another fascinating side effect. What's that? Field ordering does not matter at all.

15:18 Really? Yeah. The decoder just processes tags in whatever order they arrive. Yeah. The producer could serialize field three, then field one, then field two. The decoder doesn't care. It just maps the numbers back to the schema. So, wrapping up chapter four. Protobuf's insane compactness comes down to stripping out field names, whitespace, and punctuation, replacing them with packed integer tags. It shrinks small numbers using varints or coin sleeves. And it uses length-delimited wire types to allow decoders to gracefully skip unknown fields.

15:49 And crucially for your day-to-day engineering, this low-level wire format is incredibly stable, a vital piece of trivia. The wire encoding didn't change between proto two and proto three. Even with all the syntax updates and default behavior changes, a message serialized with proto two and one serialized with proto three, using the same field numbers, produce identical bytes on the wire. Which is why teams can migrate incrementally without simultaneous rollouts. It's a massive operational benefit.

16:15 Definitely. Well, now that we know exactly how the bytes flow over the network, we are ready to build more complex payloads. In chapter five, we are going to start modeling richer data structures. We'll dive into enums, nested messages, and repeated fields. It's where the schema really starts to reflect complex, real-world application logic. I can't wait. We will see you in chapter five.