ChatGPT Ch.2: Why It Makes Things Up

Outline

Transcript

0:00 So in 2023, two attorneys in New York City were handling a personal injury lawsuit, that Mata versus Avianca case. And they needed case law to back up their argument. So they did what, you know, millions of people were just starting to do at that time. They turned to ChatGPT to do the heavy lifting for their research. Right. And the system gave them a list of legal precedents, but it didn't just give them a list. It provided full case names, specific docket numbers, and even these incredibly authoritative sounding quotes from judicial opinions.

0:31 Which they then took, formatted perfectly, and submitted directly to the federal court. But the unbelievable twist here is that none of it existed. Yeah. Not a single case. Every single citation was just fabricated out of thin air. Yeah. Federal Judge P. Kevin Castel ended up slapping the attorneys with a $5,000 penalty for submitting made-up case law. But, you know, we're not sharing this story to tell you that the tech is dangerous or that you should avoid it. No, not at all. The point of the Mata versus Avianca story is that understanding why ChatGPT makes things up is arguably the single most important lesson for you as a user to learn.

1:10 I want you to picture a jazz musician improvising over a chord progression. Think about what that musician is doing on stage. They've internalized thousands of melodic patterns from years of practicing and studying theory. Right. There, that sounds good together. Yeah. When it's their turn to solo, they aren't pulling up a specific sheet of music in their head. They aren't recalling one specific recording from 1998. They're generating something new moment by moment. Exactly. Based on what sounds right in the context of the song, they feel the flow and predict where the melody should naturally go based on the rules of jazz.

1:48 And most of the time, that's beautiful and coherent. But occasionally, they hit a note that perfectly fits the rhythm of their solo, but it clashes completely with the underlying harmony of the band. In the rapid flow of the solo, it sounded like it belonged there, but musically speaking, it was just wrong. Oh, wow. So ChatGPT works like that with language. Yes. It's predicting the next word or part of a word based on internalized patterns from a massive amount of training data. So it's not lying to us with, like, malicious intent.

2:18 It's essentially playing a linguistic solo that just happens to be factually off -key. It's optimizing for flow, not facts. That distinction is the key to everything. It doesn't pause to check a fact database to see if a docket number is real. It just plays the next note. It's calculating probabilities. Right. It calculates the mathematical probability of what string of characters should follow the phrase according to the ruling in. And if the most probable next words happen to be factually incorrect, it produces them anyway.

2:47 Because for a text prediction engine, probability and objective truth are not the same thing. But, you know, if the machine is essentially just predicting words like an improvising musician, why is it so incredibly easy for smart, highly educated people to get completely fooled? I mean, these were federal attorneys. What's fascinating here is what we call the confidence problem. When ChatGPT delivers a wrong answer, it does so with the exact same confident authoritative tone as when it delivers a right answer.

3:16 Ah, right. It doesn't hesitate. No, it provides zero subtle clues that it might be guessing. Human communication is totally different. We are full of micro signals of doubt. Yeah. If you ask me for directions and I'm not totally sure, I'd be like, um, I think it's on the left, but I haven't been there in years. Exactly. You give reliability signals. Even in digital spaces, we have visual warnings. Think about Wikipedia. If a claim isn't verified, you get that little citation needed tag. Which is like a flashing yellow light telling you to be suspicious.

3:46 But an LM has absolutely no equivalent to that hesitation. Every single response sounds equally certain. And the reason for that is deeply ironic. Why is that? Because it learned to write by studying human language. The text it trained on books, articles, websites was mostly written by humans stating things with absolute conviction. Oh, man. So it just absorbed our unwarranted confidence as a linguistic pattern. Exactly. It's just mirroring us. And this problem goes way beyond the legal field. The 2023 study published in Scientific Reports looked at academic citations generated by these models.

4:23 Right. They asked the system to generate references for various scientific claims. And they found a massive portion of the citations were entirely fabricated. But here's the crazy part. The fake references included real author names, plausible journals, and even correctly formatted digital object identifiers, or DOIs. Because the structure of those citations was completely flawless. The system knows what an academic citation is supposed to look like. It understands the rhythm and the punctuation.

4:53 It just perfectly mimics the structure and fills in the blanks with fictional details. Yeah. And we have that incredible anecdotal report from a researcher at the University of Southern California, too. Oh, right. They brought 35 different article citations to a university librarian for help. The librarian spent hours digging through databases. And out of 35 highly specific citations, exactly zero of them were real. Not a single paper existed. It's the jazz solo playing out in an academic context.

5:21 It played the notes perfectly, but the song itself was a phantom. But wait, hold on. That was 2023. This tech moves at light speed. You can't tell me the newer models like GPT-40 are still making up fake cases. They must have patched that by now, right? Well, newer models definitely hallucinate less often. Developers are constantly tuning them. But it is impossible to eliminate the problem entirely. Really? Impossible. Yes. Hallucinations are not a bug in the code that you can just squash. They are a direct, unresolvable consequence of how the technology fundamentally works.

5:56 Because it's still just predicting text. Right. As long as you're using a prediction engine instead of a retrieval engine, the risk of it generating plausible fiction will always be there. So what does this all mean for how you should actually use it every day? You have to divide your tasks into two categories. Safe zones and danger zones. You need an intuitive sense for where a hallucination doesn't matter and where it matters immensely. Okay, let's start with the safe zones. Where can we just let the jazz musician riff?

6:22 The safe zones are all about creative, generative, and structural tasks. This is where an LLM absolutely shines. Brainstorming marketing themes or drafting a difficult email you're going to edit anyway. Or like asking it to explain a complex concept using simple vocabulary. Yes. It's also incredibly safe when you provide the boundaries yourself, like pasting in a long report and asking it to summarize the three main arguments. Because in those scenarios, there isn't an objective, verifiable, single correct answer.

6:54 If it's just an unusual angle, that's not a hallucination. It's just a creative idea. Exactly. The danger zones, on the other hand, are whenever you're dealing with specific factual claims. That is high-risk territory. Historical dates, statistical data, named individuals, legal citations. Medical dosages, scientific papers. Right. The rule of thumb is straightforward. The more specific the claim and the higher the real world stakes, the more cautious you must be. I like to treat the system like a highly enthusiastic, incredibly fast intern who occasionally gets critical details wrong.

7:26 You value their input immensely. But if they hand you a spreadsheet right before a board presentation, you double-check those numbers yourself. You don't fire the intern, but you don't blindly trust their math either. And that mindset brings us to the most practical takeaway of this entire deep dive. We know you can't ask the machine, are you sure? So how do you verify? Right. What's the actual method? It's a framework we can call the air gap method. You go to a traditional search engine or an academic database, and you search for that exact title to verify it exists in the real world.

7:58 That one 30-second habit, just opening a new tab and searching for Mata versus Avianca, would have saved those New York attorneys five grand and their professional reputations. It is an incredibly small step, but it completely neutralizes the danger. You don't need to verify every sentence when brainstorming. But if a specific stat is going into a slide deck, you jump the air gap and cross-reference it. It's so important to emphasize that hallucination isn't a reason to stop using this technology.

8:27 It's just a reason to use it with your eyes wide open. Think about the software we use every day. A spreadsheet calculates complex data with perfect precision, but only if your formulas are correct. It won't warn you if your underlying logic is flawed. Right. An LLM is incredibly powerful, but it simply cannot guarantee that every factual claim represents objective reality. Knowing those boundaries is the difference between mastering the tool and using it blindly. Well, now that we thoroughly understand the quirks of the prediction engine and why the jazz musician occasionally plays off keynotes, it's time to get hands-on.

9:01 In our next deep dive, we're moving from theory to practice. We'll walk through having your very first real conversation with the system, show you how to provide the right context, and teach you how to steer the output when that first answer isn't quite right. But until then, remember the attorneys, remember the jazz musician, and always, always double-check the sheet music.