Ch.8: Protobuf: Stop Inventing Timestamp Formats

Outline

Transcript

0:00 One team stores time as Unix epoch seconds. Another stores an ISO string. A 3rd splits seconds and nanos into separate fields. And then those teams need to exchange data. In production, the field is called created at in all three systems, but the meaning is not portable anymore. Yeah, we've all inherited that field. Every choice reads fine in isolation, and then a boundary matters, and suddenly created at means three different moments depending on who wrote it. Right, and that disagreement becomes an incident.

0:30 Protobuf's well-known types exist to stop it: regular proto messages every runtime already understands, so common shapes stop being local policy. Imagine a billing service sending epoch seconds into an analytics pipeline that expects ISO strings. The dashboard sorts events wrong, and honestly, the postmortem starts as a formatting argument. So the fix ships with the protobuf toolchain itself. Well-known types are ordinary protobuf messages that live under the google.protobuf package, and the compiler already knows where to find them.

1:00 Yup. They're, I mean, plain message definitions written in proto files, same as anything you'd write yourself. For example, a Python worker and a Java API can pass a Timestamp without agreeing on a private JSON date format first. And that is what a standard library buys you: a type that gets to be deliberately unremarkable, so no organization has to design its own timestamp. Now, the first type teams pull out of that library is Timestamp. Use it when a field means one absolute instant, full stop.

1:30 Not a local calendar appointment, and not three PM in whatever zone the user is sitting in. A real moment. Internally Timestamp has two fields. Seconds is an int 64 count of seconds since January first, 1970, UTC. Nanos, an int 32, holds the fractional part at nanosecond resolution. Huh. Just two integers. And the contract is kind of strict on purpose. The receiver never has to guess whether you meant milliseconds, local time, or a formatted string. It really is. The schema promises one point in time, and your runtime maps that into the language's native time type.

2:10 So if Timestamp is always UTC, where does the user's time zone go? Nowhere, by design. Timestamp stores the instant; the presentation layer decides how to display it for a user or a locale. Take a delivery app: it records payment settled as one instant, then renders the customer's local time on the receipt. Sure. So if I store meeting starts at 9 AM every Monday in New York, Timestamp alone is not the whole model. A recurring local schedule needs calendar rules plus a zone identifier. Timestamp wants, I mean, one specific instant: message created, payment settled, job started.

2:46 Which saves a lot of quiet time bugs. Store instants as instants, store calendar intent as calendar intent. And the next time shape is Duration, Timestamp's sibling. It reuses seconds and nanos, with a completely different meaning. Duration is not a point on the timeline. It's an amount of time. Take a retry scheduler: after a downstream request fails, a Duration can hold 5 seconds, 10.5 seconds, even negative 3 seconds, without pretending any of those are calendar instants. And the arithmetic sort of comes free with the types.

3:20 Subtract one Timestamp from another and you get a Duration. Then add a Duration onto a Timestamp and it lands on a new Timestamp. Huh, the types do the calendar math for you. In Python, that maps cleanly to datetime and timedelta. In Java, the natural targets are Instant and java.time.Duration. Same math on both ends of the API. Then using them is refreshingly mundane, you know. You import the proto file from google/protobuf and use the type as a field. For Timestamp, the import path is google/protobuf/timestamp.proto.

3:55 For Duration, it is duration.proto. Then the field type is google.protobuf.Timestamp or google.protobuf.Duration. You don't vendor a random copy into your repo, and you never paste the message definition into your own schema. The include files ship with protoc, and the language libraries understand them. If your production build can't find them, fix the proto path or the toolchain setup. Do not fork the standard type. Yeah. A forked copy drifts quietly, and the shared type you counted on becomes, honestly, one more local convention to argue about.

4:31 So that covers the fixed time shapes. But the next boundary is trickier: what happens when you cannot know, until runtime, which message will arrive? Then you reach for Any. It wraps an arbitrary serialized protobuf message plus a type URL that names what is inside, usually type.googleapis.com/the fully qualified message name. And the receiver then checks that type URL, picks the matching message class, and unpacks the bytes into it. For example, an extension registry can publish a BillingCharged message after a plugin rollout, while older consumers keep routing unknown events instead of crashing.

5:06 Still a contract, though. Producers and consumers agree up front on which message types are allowed, and on how to react when an unknown one shows up. And I can already hear a reviewer on this one: "oneof handles a field with several possible shapes. Why not always use that?" Fair question. A oneof lists every possible field inside the schema file, the compiler sees the whole set, and the generated API gives you a clear case check. Sure, and when the variants are known at design time, that's exactly what you want.

5:38 Any is for a set that grows outside the file. Imagine a plugin that adds a new message type later, or an event bus routing event types owned by different teams. Hold on, though. That flexibility costs readability. If I open a schema and see Any, I cannot tell from the field alone what might arrive. That tradeoff is the clean line: oneof declares the set in the schema, while Any moves part of the contract out of the field declaration. Then there is the one my instinct says to handle with gloves: Struct. Struct represents arbitrary JSON shaped data.

6:14 It lets you smuggle an entire JSON document into a typed system, no questions asked. Right. Under the hood it's a map from string keys to Value objects, and a Value holds null, a number, a string, a boolean, another Struct, or a list of Values. Bags inside the bag. And customs, meaning your schema review, waves the whole suitcase through every time. Stack those layers, Struct, Value, and ListValue, and you can carry a full JSON document without storing raw JSON text. And a staff engineer will push back right here: "schemas are overhead, just put Struct everywhere and move on."

6:50 That's a worse trade than it looks, though. The schema gets weaker, the generated API gets less precise, and the binary encoding is less compact than a real typed message. So reach for it only when the data is genuinely unstructured. Consider a feature-flag service holding arbitrary vendor settings, where no two integrations share a stable shape. Struct is an escape hatch, not the normal design path. Now let's say an older profile API needs an age field where 0 and missing mean two different things.

7:22 Those are the wrapper types. Each scalar wrapper is really just a tiny message, like StringValue or Int32Value. Each one holds a single field named value, carrying the scalar it wraps. So if the wrapper is absent, the user never gave an age. If it's present holding 0, the age is explicitly 0. Hmm, so presence itself is doing the work? No wrapper means no answer, and a wrapper around 0 means a real 0? Right. That gave older proto3 schemas a nullable scalar shape, the same presence problem as the 0-value trap.

7:55 And it is exactly where optional comes back into the picture. But today, wrappers are mostly compatibility history for ordinary scalar fields. Because proto3 optional came back. If you are designing a new scalar field that needs presence, optional int 32 is usually clearer than the old wrapper shape. They still show up in real APIs, though, especially older Google-style surfaces and schemas that needed nullable scalars before optional returned. You should understand them because you will read them, and you should kind of resist copying them into new fields.

8:31 So in review, the question is just: is this presence need new, or inherited? That gives reviewers a clean migration line. Now, wait till you hear what these types look like once they hit JSON. It is half the reason they feel built in. Proto JSON gives them special treatment, each type with its own mapping. In JSON, Timestamp comes out as an RFC 30 3:30 9 string, a UTC stamp ending in Z, rather than an object holding seconds and nanos fields. Oh, wow. So a JSON client just sees a normal date string, with no idea protobuf was ever involved?

9:08 Exactly. Duration becomes a string with an s suffix, like three hundred seconds or 10.5 seconds. And Any becomes a JSON object with an at type field, so the type URL stays visible. And wrappers usually serialize like the scalar value they wrap, which is much easier for JSON clients to consume. The binary side stays normal protobuf. So now the wire format and the human-facing form both behave. And after that JSON detour, the homegrown timestamp comes back to bite. It always starts out looking simpler.

9:41 Maybe you only need created at seconds today, so you store an int 64. Then someone needs fractional seconds. Or you store ISO strings because they are readable, and then one producer includes an offset, another always sends Z, and a 3rd leaves the zone off because the database column was, you know, local time. Yeah. I've shipped this exact pattern, more than once. Every custom time field starts as convenience and turns into policy. And that bug ages badly in production: the schema says string, the comment says UTC, and the data says mostly UTC until 1 day it doesn't.

10:17 Timestamp doesn't solve every time problem, but it removes format negotiation from the schema instead of hiding policy inside every integration. So let's land the rule of thumb. First, Timestamp for a concrete instant, and Duration for a span of time. Then Any, reserved for when the incoming type really must be discovered at runtime. Struct when you are intentionally carrying JSON shaped data and you accept the weaker guarantees. Wrappers when an existing schema already uses them, or when compatibility demands that shape.

10:48 For new nullable scalars, prefer optional. And above all, do not define your own common type unless the standard one genuinely does not fit. In the end this is an architecture call: a schema is not just a place to store bytes. It is a place to remove arguments between systems. So after that whole tour, the payoff: well-known types give protobuf a shared vocabulary for the unglamorous shapes every system needs. Time, duration, dynamic messages, JSON shaped blobs, and a little nullable scalar history.

11:18 Next, we look at packages, imports, and how proto files stay organized once one schema becomes a project with many teams. Thanks for listening to Learning Podcasts.