DNS Explained: Why Answers Disagree

Outline

Transcript

0:00 Today, we are focusing on DNS, and we're going to start directly with the fundamental purpose of the system. I mean, software needs stable names like api.example.com because the infrastructure underneath is always, always changing. Right, exactly. I mean, you build a system and you absolutely cannot hard code a specific machine's IP address because, you know, next month that machine is going to fail over or get replaced or get tucked behind a new load balancer. Yeah, the IP address is move. So DNS exists to separate the stable name your application calls from the destination machine address that is, well, constantly shifting.

0:35 And that brings us to the core thesis for this deep dive. If you're a software engineer, you need a clean mental model here. DNS is not a single lookup against one globally shared database. It's just not. Right, that's the trap. It is. It's actually a hierarchical naming system that sits in front of multiple independent cache layers, and that's the key. The hierarchy explains how an answer is found, but the caches explain why answers diverge. And once you separate that naming path from the caching path, things start to click.

1:05 I mean, the familiar failure pattern of getting mixed answers across different devices stops being this mystery and becomes a really highly specific debugging hypothesis. Yeah, let's actually, let's look at that familiar failure pattern. It's a rite of passage. You know, you're in the middle of a SEV1 outage. Your primary database cluster just failed. Oh yeah, panic mode. Exactly. To stop the bleeding, you jump into your configuration dashboard, you update a DNS record to point your traffic to the backup cluster, and you hit save.

1:36 And then you clear your local terminal cache, you ping the API, and verify the new address is responding. Right. You announce in Slack that the fix is live, but immediately customer support says users are still getting errors. You check the logs, and half your application servers are successfully routing to the backup, but the other half are, you know, stubbornly trying to connect to the dead cluster. Which is incredibly frustrating. But to understand why your system breaks like that, you first need to trace the exact path an application takes when it operates correctly.

2:06 Okay, so let's follow the lookup path. First, we follow the naming path. My application wants to connect to a host name, but it can't connect to text, right? It needs an IP address. Right. It needs a number. So it asks the operating system's stub resolver. The stub resolver is just a lightweight client on your machine. It checks local caches first, like your OS cache or the local hosts file. Even some browsers have their own internal caches, right? Yeah, they do. But if there's no local answer, the query has to leave your machine.

2:36 It goes out to a recursive resolver. And the recursive resolver is the real workhorse here. It absolutely is. This is usually run by your ISP, your corporate network, or public service like Cloudflare. Its entire job is to take your question, track down the answer, and then crucially cache it for the next person. Okay, but let's say the recursive resolver's cache is empty. It has no idea what the out key address is. What happens then? Well, then it walks the hierarchy downward. It starts at the very top.

3:02 It asks a root server. Now, the root server doesn't know the IP for your specific API, but it refers the resolver to the top-level domain server, the TLD server. Like the servers running the .com layer. Exactly. So the recursive resolver asks the .com TLD server. That server says, I don't know the IP, but I know exactly which authoritative name server handles your specific domain. It's basically a strict referral system. Yeah, exactly. Like, I don't know the answer, but I know exactly which server is responsible for the next layer down.

3:33 That's the perfect way to think about it. And finally, the recursive resolver asks the authoritative name server, which is the actual source of truth for the zone, and that server hands back the actual record. So the authoritative server finally hands back an answer. Yeah. Now, before we look at what that answer actually contains, we should probably separate out some roles that engineers commonly confuse. Oh, absolutely. This is a huge source of debugging pain. People mix up four very distinct roles.

4:00 First, you have the registrar. That's just where you bought the domain. They handle billing and ownership. Right. Ownership is not authority. Exactly. Second is the DNS host. That's the dashboard where you actually log in to manage your records. Third is the authoritative name server itself, the actual internet servers that serve as the source of truth for your zone. Which might be run by the DNS host, but conceptually, it's the server responding to queries. Right. And fourth is the recursive resolver.

4:29 That's the workhorse the user's laptop actually talks to. Mixing these up is a disaster during an incident. You log into the registrar, see the correct IP, and assume the internet is fixed. But the internet is talking to recursive resolvers, not your registrar dashboard. Exactly. Okay, so the authoritative server hands back an answer. Let's talk about the practical record types that software engineers need to care about. What's actually in that answer? Well, the most direct answers are A records and quad A records.

4:56 Right. A records mapped to IPv4 and quad A records mapped to IPv6. Yeah, that's the final destination. You might also see NS records for name servers, MX records for mail routing, and TXT records for, you know, metadata and verification. But the big one, the one that really matters operationally, is the CNAM. Yes. The CNAME, it stands for canonical name. And this is vital because a CNAME maps one name to another name, not to an IP address. Which is essential if you're pointing an internal name to a provider-managed host name.

5:29 Like, you know, pointing to an AWS load balancer without hard-coding the IP. Right. Because AWS controls that IP, and they might change it. But the operational trap of a CNAME is that it adds an extra layer of indirection. Right. It's not just a flat alias. It means one more resolution step. Exactly. One more cache boundary. If a hostname follows another hostname first, the real failure you're seeing might actually be one step further down the chain. Because the recursive resolver gets the CNAME, and then it has to start the lookup process all over again for the new target name.

6:01 Exactly. And that extra cache boundary introduces the most painful operational aspect of the whole system, TTLs and caching layers. Which brings us right back to that SEV1 outage where half the users are seeing the new database and half are seeing the old one. Every cached answer comes with a TTL, right? A time to live. Yes. And this is where the biggest myth exists. I hear engineers say, well, we lowered the TTL so the change should be live in five minutes. Right. They treat it like a global deployment timer.

6:29 Exactly. But it's not a global deployment timer. A TTL is just a local caching policy telling a specific cache how long an answer is valid before it becomes stale. So setting a five-minute TTL right now does absolutely nothing to erase an older one -hour TTL that a resolver cached, say, 59 minutes ago. Precisely. Let's say a recursive resolver in London asked for your record 59 minutes ago when the TTL was an hour. It's going to hold on to that dead IP address for another minute. While resolver in Tokyo that asks right now gets the new IP immediately.

7:04 Right. And this is why the phrase DNS propagation is so misleading. I mean, people talk about it like it's a weather front moving across the Internet. Right. Like this global wave of updates. But there is no global distribution way of pushing new answers outward. It's just a chaotic patchwork of independent caches browsers, local networks, ISP resolvers, all expiring and refreshing on their own uncoordinated schedules. Okay. So caching old IP addresses makes intuitive sense. You want to save the answer so you don't have to look it up again.

7:32 But the story gets significantly more confusing when the system caches the absence of an answer. Oh, negative caching. This one really catches software engineers off guard. Because engineers naturally expect positive data to be cached. But we're consistently surprised that not found has its own cache lifetime. Yeah. If a resolver asks for a name and receives an NX domain response, meaning the domain doesn't exist, it caches that failure. Wait. Why would it save a failure? It has to. To prevent the authoritative servers from being hammered.

8:06 If there's a broken script frantically requesting a typoed hostname, you don't want every single one of those queries going all the way up to the root server. Ah. Okay. So it's a protection mechanism. Exactly. But the operational impact is huge. Let's say you're in an incident and you create a brand new record to fix a broken route. You test it and users are still getting a failure. Even though the record definitely exists now. Right. It doesn't mean the authoritative zone is broken. It means a resolver cached the previous absence of the record, that NX domain, and it just hasn't retried yet.

8:36 I mean, that is brutal for debugging if you don't know what's happening. It really is. Now, beyond timing and caching, there's one final scenario where the system will actively return different answers. Not because of stale caches, but by intentional design. You're talking about split horizon DNS. Right. Network dependent resolution. This is when the exact same hostname intentionally returns different answers depending on where the query originates. Yeah. Split horizon or split view. A classic symptom is an internal database named like db.internal.example.com.

9:11 And it resolves to a private IP when you're connected to the corporate VPN. Exactly. But if you disconnect and go on the public internet, it fails to resolve entirely. Or maybe an API resolves to an internal load balancer inside the office, but a public load balancer from home. And the core lesson here is that the system isn't malfunctioning. It's doing exactly what it was configured to do. Right. Context is everything. Which is the exact trap engineers fall into. You jump into Slack and say, hey, what does this name resolve to for you?

9:40 But asking, what does this name resolve to? Is it incomplete debugging question? It's completely incomplete. You're not defining the network context, the resolver or the cache state. So we've mapped the hierarchy, the caches, the aliases and the network context. Let's synthesize this into a concrete six step diagnostic checklist. Something a software engineer can use the next time they hit a DNS issue. I love this. Okay. Step one. What is the authoritative answer right now at the source of truth?

10:07 Right. Don't trust your local cache. Query the authoritative server directly to see what the actual record says. Exactly. Step two. Which specific recursive resolver is the failing client actually talking to? You and a customer might be hitting completely different resolvers. Okay. Step three. What is the TTL? And when was the old answer likely cached? Remember, lowering the TPL now doesn't rewrite the past. Yes. Step four. Is there a synonym-y alias chain hiding a failure one step down? If you're pointing to an AWS load balancer, is the load balancer's host name actually resolving correctly?

10:43 That extra hop is crucial. Okay. Step five. Are you looking at a negative cache in NX domain instead of a positive one? Did the resolver cache the absence of the record before you created it? Right. And finally, step six. Is there a split horizon configuration intentionally serving different internal versus external results? Are you testing from the right network context? And equipped with this six-step checklist, that classic exasperated phrase, it's probably DNS, is no longer a vague complaint about internet magic.

11:14 No. It becomes a narrow, highly specific hypothesis about where two components of the system disagree. Because DNS is not a single lookup against a database. It is a hierarchical naming system layered behind independent caches. Most weird behavior is simply those two systems interacting exactly as designed. Once you understand the mechanics of authority, recursion, caching, aliasing, and context, you can systematically interrogate the system, and it ceases to be a mystery.