Ch.1: OpenTelemetry: How Engineers Debug Distributed Systems
Transcript
0:00 Two AM. Your phone is buzzing. The page says checkout latency just jumped to 4 seconds, the error rate is climbing, and something is broken in production. And of course, you are the one on call tonight. That is the part nobody warned you about. By the end of this chapter, you'll see how that same page resolves in 6 minutes instead of 6 hours. We'll walk through why grep stops working at this scale, and how engineers fix the slow request without searching 40 machines by hand. So you get up, you crack open the laptop, you SSH into the checkout service, and you do what you have always done.
0:36 You tail the logs. And you see nothing useful. Right. The logs are scrolling. The timestamps are moving. But there is no error. No stack trace. Just a quiet stream of normal-looking lines while the dashboard says many requests are failing. Yeah. So you try a different machine. Same story. You try a 3rd one. Same. For example, you grep for the user ID. You grep for the order ID. You grep for the word `error`. An hour in, you still do not know whether the problem is in checkout itself, or in something checkout calls, or in something that calls checkout.
1:11 The logs cannot tell you, because each machine only sees its own slice. The result is an hour gone and 0 answers. And the system is not one box anymore. Take a real checkout request. It lands on the API gateway, then hits auth, cart, inventory, pricing, recommendations, and payments, and triggers notifications on the way out. 8 services for a pair of socks. And you wrote maybe two of them. Sure. And every one of them is logging into its own file. 8 log streams. None of them carry a thread that ties the same user request together. So even if every service is logging perfectly, you have no way to ask the obvious question.
1:51 Which service is making this one request slow? I think the first time I owned a service like this, I spent more time finding the right repo than reading the actual error. After that first hour, the part most engineers, Do not see coming starts to land. The debugging tools you learned in school still work. They just do not scale. Print statements work when there is one process. Sure. Print works on your laptop. Grep works when there is one log file. They keep working all the way through your first internship and into your first real job. And then 1 day the system has 40 machines, the requests cross half of them, and you cannot grep across all 40 at once.
2:33 Honestly, my instinct was to grep harder. That did not work. For example, I once piped grep across 40 hosts and watched the SSH connections time out one by one. The whole approach is a nightmare. You reach for the next tool and it is just not there. From the grep-doesn't-scale moment, the next move is the one nobody handed you in school. You need to follow one specific request as it travels across every service in the call path. With timing at every step. So you can look at one slow checkout and see, immediately, that the request spent 40 milliseconds in cart, 50 in pricing, and 3.6 seconds inside one downstream call from inventory. Here's where it gets interesting.
3:17 That is a different shape of data than logs. Completely different. Logs are a stream per machine. What you need is a tree per request. Right. It is the difference between reading the notebooks of 40 separate security guards and watching one camera follow that person through the whole building. Exactly. One trace is that single camera. Now you can see the whole request as one timeline. And the industry has names for the pieces of that tree. The whole thing, the full record of one request as it moved through your fleet, is called a trace. Each step inside that trace, every function call or database query or outbound request, is called a span.
3:58 And every span in the same trace shares one identifier. A trace ID. It rides along inside the request itself, gets passed from service to service in an HTTP header, and any service that wants to log or measure something attaches that trace ID to whatever it produces. Now your 8 log streams can be reassembled, after the fact, into one timeline. Oh, wow. A timeline that shows you exactly where the 3.6 seconds went. Let's hear it. Imagine each service drawing a horizontal bar for the time it spent on this one request. The bars stack into a waterfall.
4:34 The widest bar is the slowest step. That is the picture. Same trace ID idea opens a bigger frame. Traces are one of three kinds of telemetry that modern systems emit. The other two are metrics and logs, and each one answers a different kind of question. Metrics are the smoke alarm. They tell you the house is on fire, and roughly how big. Right. The dashboard showing p 99 latency jumped from two hundred milliseconds to 4 seconds. That is a metric. It sizes the fire, but not the room. Sure. Traces walk you to the exact room.
5:07 Which path is slow, and at which step. That is the waterfall you just heard about. Same request, every service it touched, with timings. Right. And honestly, that is the click for me. But the widest bar only says inventory was slow. Not why. Right, so are logs obsolete now? The trace already found the room. That is the trap. The trace finds the room. Logs tell you what actually sparked the fire, the exact thing that happened in that step. The connection pool was exhausted. The retry fired three times. The downstream replica was lagging.
5:41 So once you are in the right room, you read 10 log lines instead of 10 million. So somebody has to produce all three of those things, in a way that ties them together with a shared trace ID, in a format that the storage and dashboard tools downstream can read. That somebody is your application code, and the standard it speaks is called OpenTelemetry. One vocabulary for traces, metrics, and logs. One wire format on the way out, called OTLP. One set of SDKs across most of the popular languages. And dozens of backends that all know how to ingest it.
6:14 Which means you instrument once and stay portable. From that one-vocabulary framing, three things OpenTelemetry is not. New engineers mix these up regularly. Sure. Walk me through what those three things are, one at a time. It is not a backend. It does not store your traces. It does not draw your dashboards. It does not run your alerts. It is the thing that produces the telemetry, not the thing that holds it. Right. And vendors do not own it either, even though most people first hear about it from one.
6:43 Yeah. OpenTelemetry is an open standard, governed by the Cloud Native Computing Foundation, with contributors from across the industry. And it is not a dashboard product. The dashboard is whatever backend you point your telemetry at. Same backend-versus-standard split, one analogy that makes the value land. OpenTelemetry is the language your application speaks about itself. The backend you choose is the dictionary that turns that language into queries, alerts, and graphs. I love that. And the analogy survives the practical question that always comes next.
7:17 Sure. What happens when you want to change backends. With OpenTelemetry, your application code does not care. It still speaks the same language. You just swap the dictionary, and the application keeps going. Wait, really? Even the SDKs stay the same? They do. Which is roughly the opposite of what observability looked like 5 years ago, when every vendor had its own proprietary agent and switching meant rewriting your instrumentation in every codebase. That gives you the freedom to change backends without touching application code.
7:46 So the timing matters. Consider three snapshots. In two thousand 5, most production systems were one monolith on one box, with 1 log file, and grep was enough. In two thousand 15, the world moved to microservices, and the observability tools split along vendor lines while the architecture was getting harder to debug. And in 2026? By now the request you are chasing might also be calling a language model, which calls a tool, which calls another service, which streams tokens back. The trace tree got deeper, the latency got more variable, and the cost of a slow path got more visible.
8:22 So the demand for one shared way to see across all of that finally caught up. Oh, right. And that is the moment OpenTelemetry has been built for. Exactly the moment. And the adoption picture matches the demand. For example, Datadog, New Relic, and Honeycomb all ingest OpenTelemetry natively now. So do open-source platforms like Jaeger and Grafana. Every major observability backend is on the same standard. Exactly. So picking OpenTelemetry is not picking a side. Source links and the full vendor inventory are in the description.
8:55 Right. It is the opposite. It is the one decision that does not lock you into a side. So here is what the series gives you, end to end. For example, following one request through every machine it touches. The three signals in production and when to reach for each one. Instrumenting a real service from scratch. Right. Walk me through the harder ones. Sure. Sampling, because nobody stores every trace, and the strategy matters more than people think. The collector, which is the middleman every production deployment ends up running.
9:26 Choosing a backend. And the production-culture pieces? SLOs and error budgets, which is the language production teams use to negotiate reliability. Debugging an actual incident end to end. What observability costs, and how teams keep it from getting expensive. And the newest corner, observability for AI agents. That last one is the corner most engineers have not touched yet. True. And it might be the first place a new grad ever has to instrument something. That sets up everything else the series builds on.
9:57 From that whole topic list, back to the two AM page. After the series, the same page resolves differently. Imagine you go straight to the dashboard, you spot the failing endpoint, you open one slow trace, you see exactly which downstream call is eating the latency, you pivot from that call to its logs, you find the connection pool error, and you roll the change back. Right. 6 minutes, not 6 hours. Oh, man. One on-call shift, two different physics. Next time, we walk through one real production debug end to end, and let traces, metrics, and logs show up as three different questions about the same problem.
10:36 Production systems are going to keep getting more distributed, year after year. The engineers who can see across them are the ones who do not get stuck at two AM. So pick the standard early. Thanks for listening to Learning Podcasts.