Ch.5: Protobuf Maps Are Just Repeated Messages

Outline

Transcript

0:00 Welcome to Learning Podcasts. Protobuf, chapter five: Maps Are Just Repeated Messages. By the end of this chapter, one of the four features we cover is going to look strangely familiar. You are modeling a user account. The status field can be active, suspended, or deleted. So you reach for a string. One service writes lowercase active. Another writes capital A Active. A third gets enthusiastic and ships ACTIVE in all caps. All three are valid Python. All three pass code review. All three deploy.

0:32 And somewhere downstream, a switch statement only matches one of them. Three valid spellings, three potential bugs, zero compile-time protection. Protobuf gives you the enum. You declare it with the keyword enum, you give it a CamelCase type name like AccountStatus, and inside curly braces you list each value with a name in upper snake case and an integer assignment. So AccountStatus contains ACCOUNT_STATUS_UNSPECIFIED equals zero, ACCOUNT_STATUS_ACTIVE equals one, ACCOUNT_STATUS_SUSPENDED equals two, ACCOUNT_STATUS_DELETED equals three.

1:08 A named set of integer constants. The string is gone. Typos become compiler errors. One source of truth, one valid set of values. Now look at the value names. Every one of them starts with the type name. Why all that ceremony when the values already live inside an enum block. The answer is a quirk inherited from C plus plus. Protobuf uses C plus plus scoping rules for enums, which means enum values are not scoped to their containing type. They live at the same level as the enum itself. So if you defined two enums in one file, and each one had a bare value called ACTIVE, the compiler would refuse the file.

1:49 Two definitions of ACTIVE at the same scope. Prefixing every value with the type name in upper snake case sidesteps the collision entirely. Strict-looking convention, real-world reason. The first value in a proto3 enum has to be zero. That is a hard rule, and it is not arbitrary. In proto3, when a field is not explicitly set, it deserializes to the zero value. For an enum, that is whichever name you put at zero. So picture the alternative for a second. You make ACTIVE the zero. A sender that explicitly sets the status to ACTIVE looks identical on the wire to a sender that never set the field at all.

2:30 You have erased the difference between known and unknown. The convention, baked into protobuf style guides, is to make the zero value an UNSPECIFIED entry. A meaningful business value never sits at zero. Otherwise unset and active mean the same thing. On the wire, an enum is encoded exactly like an int32 using the varint encoding from chapter four. Same bytes, same wire type. No overhead compared to a plain integer, but with compile-time safety you do not get from a raw number. Cheap on bytes, expensive on bugs avoided.

3:05 Beyond scalars and enums, real messages contain other messages. A user has an address. An order contains line items. Protobuf lets any message reference any other message as a field type. And it lets you define a message inside another message, which is called a nested message. Picture a User message that needs a home address. You could define Address at the top level of your proto file. That works if Address is shared across many messages. But if Address only ever appears inside User, defining it at the top level just clutters the namespace.

3:40 So you define Address inside the User block and reference it directly. One organizational choice, no wire-format consequence. What does that nested type look like in generated code. In Java, the nested Address class is exposed as User dot Address. From outside User, you reference it with the qualified name. From inside, the short name is sufficient. In Python, it appears as a nested class accessible through the outer class, same idea. And on the wire, there is no difference between a top-level message used as a field type and a nested one.

4:13 Both are encoded as length-delimited values, wire type two: one varint byte count followed by the serialized inner message. Nesting is a schema-level organizational tool, nothing more. Nest when the inner type is tightly coupled to its parent. Define at the top level when multiple messages share the same structure. A user might have one email address or five. An order might contain a single line item or dozens. To model collections, you put the keyword repeated before the field type. So repeated string emails equals five means field number five holds zero or more strings.

4:38 The generated code exposes this as a list in Python and a repeated field accessor in Java. The sequence is ordered, and elements come back in the order they were added. Same field number, many values. On the wire, the encoding depends on what the elements are. For messages and for strings, each element gets its own tag, length, value entry, and they all share the same field number. The decoder collects every entry it sees with that number and assembles a list, preserving the order they appeared in.

5:08 There has to be a length prefix on each one, because a string or a message has no fixed size. Without the per-element length, the decoder would not know where one element ends and the next one begins. One tag per element, lengths everywhere. Repeated numeric scalars get a smarter treatment. By default in proto3, they use packed encoding. Instead of a separate tag per element, the encoder writes one tag with wire type two, one varint byte count, and then all the values concatenated back to back. So a repeated int32 holding ten, twenty, thirty becomes one tag, one length, three varints.

5:48 Each value is still a varint from chapter four, just stored back to back without per-element tags. That is a real saving when a list contains hundreds of small integers. The reason strings and messages cannot be packed is the one we just covered. Each one needs its own length prefix. Tags on the outside, varints on the inside. The fourth feature is the map. You declare it as map angle bracket key type comma value type angle bracket, then the field name and number. So map angle bracket string comma int32 angle bracket word_counts equals one defines a mapping from strings to thirty-two-bit integers.

6:29 The generated code exposes it as a dict in Python and a map accessor in Java. Key plus value plus field name plus field number. Familiar shape, new keyword. There are constraints. The key type has to be an integral or string scalar. So int32, int64, uint32, uint64, sint32, sint64, fixed32, fixed64, sfixed32, sfixed64, bool, or string. Floating-point types, bytes, enums, and messages cannot be keys. The value type can be anything except another map. No nesting maps inside maps. And there is one more thing the docs are explicit about.

7:08 Maps are unordered. Serializing and deserializing a map may reorder the entries. If insertion order matters, you do not want a map. If order matters, reach for repeated, not map. And here is the title of this chapter. A map field is syntactic sugar. The compiler does not generate a special wire format for it. Under the hood, it treats map angle bracket key comma value angle bracket as a repeated nested message with two fields. A key field at number one and a value field at number two. So a map with three entries on the wire looks identical to three instances of that little nested message. Same length-delimited entries, same field numbers, same bytes.

7:51 The fourth feature in this chapter is the first three composed. So old code that predates map syntax can still decode it as a repeated field. That is how the format stays backward compatible. Each feature fills a distinct role. Enums for a closed set of states known at schema time. Nested messages for sub-structure that belongs to one specific parent. Repeated fields for ordered collections of any size. Maps for unordered key-value lookups, even though under the hood the wire format is reusing repeated nested messages.

8:23 Four tools, one wire format underneath all of them. You now have the protobuf vocabulary for modeling real data. Closed sets, sub-structures, ordered lists, key-value lookups, all with compile-time safety and predictable wire encoding. Next chapter we look at oneof and optional fields, two features that let you express mutually exclusive choices and explicit presence inside a single message. Thanks for listening to Learning Podcasts.