gRPC Ch.1: Remote Procedure Calls

Outline

Transcript

0:00 Welcome to the Deep Dive. So if you're listening to this, you're exactly the kind of person we designed this for. You know, you're a software engineer. Right. You probably spend your days in the trenches, writing Python, maybe building out some architectures in Java. Yeah, exactly. You're incredibly comfortable with functions, classes, data structures. I mean, you know how to build logic. But today, our mission is to take those programming paradigms that you already know intuitively and, well, push them to their absolute physical limits.

0:28 Because today we are unpacking chapter one of a pretty massive 22-chapter series titled gRPC for Any Software Engineer. And this first chapter is really about setting the stage. Before we can even touch the modern frameworks, we have to look at the fundamental physics of distributed computing. We're basically taking the environment you completely control and throwing it into a place where, honestly, chaos is the default state. Perfectly said. Okay, let's unpack this. Because the best way to understand where we're going to be in the future is to anchor ourselves to where you currently are.

1:03 Think about the absolute bedrock of your day-to-day programming, the local function call. Right, the foundation of everything we build. We don't need to explain how a function works to you. I mean, you write the code, you pass an argument, it executes, you get a return value. And what's really fascinating here is the illusion that it creates. It feels entirely seamless. You pass a Python dictionary or, you know, you instantiate a Java object, and because that function is operative, it's operating inside the exact same memory space as the rest of your program.

1:31 It just works. Yeah, the CPU just reads the RAM, processes the logic, and hands the result back in literally nanoseconds. Exactly. You never have to think about the physical mechanics of your motherboard when you type out a function call. Right. But this raises a pretty important question, and it's really the pivot for our entire deep dive today. What happens when that incredibly useful function you just wrote, maybe it runs a really heavy machine learning model or, like, handles a complex inventory lookup?

2:00 What if it lives on a completely different computer? That is the exact moment the illusion shatters. Yeah. Because other teams in your organization, they need your machine learning model. But they're running their own programs on their own physical machines, maybe in a data center halfway across the country. They can't just type import your module into their code. No, they can't. Your function simply does not exist in their computer's memory. Because you can't just, you know, reach across the void and get a code.

2:27 You can't just grab code from a motherboard you don't own. So we have to figure out how to bridge the physical gap between two separate machines. Which is the classic distributed systems problem. Right. Now, my naive instinct here, and I'm sure a lot of engineers think this way initially, is to ask, well, why is this actually difficult? I mean, why can't I just open a standard network connection, send over a few bytes that say, hey, run this specific function with these arguments, wait for a second, and let the other computers send the result back over the wire?

2:56 It's an entirely different question. It's an entirely different philosophical question. And honestly, at the absolute highest level, that is the goal. You want to send bytes, have them do the work, and get bytes back. Sounds easy enough. But the execution of that idea introduces a massive amount of friction. Creating a system that makes a remote network call feel as easy and as reliable as a local function call is the fundamental historical dilemma of distributed computing. So let's get into the mechanics of why that just send some bytes approach falls apart in practice.

3:26 Let's do it. So when you actually try to build this, you run into what we can basically call the three hidden nightmares of remote calls. These are three invisible mechanisms your local machine handles perfectly for you. But the second you introduce a network cable, they just completely break. They really do. So let's start with nightmare number one, the data. Right. The actual information you're attempting to pass from one machine to another. Wait, if memory is ultimately just, you know, ones and zeros under the hood, why can't I just take the ones and zeros that make up my Python dictionary and just push them over the wall?

3:57 I mean, I can't just take the ones and zeros and just push them over the network to a Java program. Well, if we connect this to the bigger picture of language runtimes, it's because those ones and zeros are structured entirely differently depending on the ecosystem you're in. Oh, because of how the language is actually managed memory. Exactly. A Python dictionary isn't just raw data. It's a very specific CPython hash table structure. It's managed by Python's internal memory allocator. It's got pointer references and garbage collection reference counts baked into it.

4:27 And Java has absolutely no idea what to do with the Python reference count. None whatsoever. Java manages memory through the Java virtual machine heap. Its objects are laid out according to strict JVM specifications. They simply do not speak the same memory language. Right. If you just blasted raw Python memory bytes into a Java program, it would look like complete garbage data to Java and the program would just crash. So we need some sort of translator. We have to strip away all the language-specific memory management stuff.

4:57 And extract just the pure data. Yes. And that concept is called serialization. Right. Serialization. You take your structured language-bound data and you translate it into a universal sequence of bytes. And actually, if you're listening to this, you've probably already done this. If you've ever taken a Python dictionary and run like JSON.dumps to turn it into a JSON string, you've serialized it. You packed it up for transit. Yep. And parsing that JSON string back into an object on the other side, that's deserialization.

5:26 Which, neatly, solves the data problem. But once you have those serialized bytes, you immediately crash into nightmare number two, which is transit. The networking itself. Because you can't just dump a raw stream of bytes into an Ethernet port and hope the other side understands what to do with them. But, I mean, I know we have TCP IP to guarantee the packets arrive. So isn't the transit part kind of already solved by the operating system? TCP guarantees delivery of the packets, yes. But it treats everything as an infinite, continuous stream of bytes.

5:56 It does not know what a file is. It doesn't know what a function call is. Oh, right. It just sees data moving. Exactly. So these two separate computers must agree on application-level conversation rules. When a stream of bytes starts arriving, how does the receiving computer know where your serialized data actually begins? Or how does it know when the message is completely finished? Ah, I see. Because the network wire doesn't just stop transmitting. It's an open pipe. So we have to invent rules to frame the messages.

6:26 Precisely. You need a robust network protocol to handle establishing the connections, the framing of the actual messages, and the rules of engagement for how to handle concurrent requests so you don't jumble different function calls together in the same pipe. Which feels like it naturally leads us into nightmare number three. And to me, honestly, this is the most terrifying one because it requires a complete paradigm shift for a developer. Nightmare number three is unpredictable failures. It's a big one.

6:54 Because if my local Python script throws an error, it is going to be a failure. It's almost always my fault. Right. It's a bug in my code. I missed a null check or I have a syntax error. But over a network, the causes of failure shift from the logical to the physical. And what's fascinating here is that chaos is the default state of a network. A remote machine might be experiencing a massive spike in traffic and its CPU is just too overloaded to answer you. Or like a physical cable gets unplugged mid transfer.

7:23 Let me push back on this, though. If the network drops, wouldn't my program get stuck? I can just throw a network disconnected exception. I can just catch that error and retry the function call. Why is that considered a nightmare? Because of the unknowable state. Unknowable state. Yeah. Consider this scenario. Your local machine sends a request over the network to process a user's credit card. The request reaches the remote server perfectly. The remote server processes the payment, charges the card, and packages up a success response.

7:51 It puts that response onto the network. Then the network router fails and drops the response. But it's not a failure. It's a packet. Oh, wow. OK. So my local machine is just sitting there waiting for an answer. The connection eventually times out and throws an error. But from my perspective, I have absolutely no idea if the remote server actually did the work or if the request never even made there in the first place. Exactly. So do you retry the function call? If you do, you might charge the user's credit card a second time.

8:17 Nice. But if you don't retry, the payment might not have actually gone through at all. These aren't software bugs you can just fix with a pull request. They are fundamental properties of physical networks. Any system designed for calling functions across computers must have a robust, built-in way to handle that chaos. Man, that is stressful. So we have the three nightmares. Serialization of the data, the transit network protocol, and handling unpredictable physical failures. Right. And because these are rooted in the physical reality of networks, this isn't a new problem.

8:48 Programmers have been trying to solve this since, what, the 1980s. Let's look at the ghosts of distributed systems past. It really begins in earnest with Sun Microsystems. In the 1980s, they created something called SunRPC. And the core idea they had was undeniably elegant. How did it work? Like what was their breakthrough? They recognized that writing all this boilerplate networking code, you know, the socket connections, the byte framing, the error handling was just tedious and prone to bugs. So they said, let's define the functions we want to expose.

9:22 Then we will build a tool that automatically generates all the messy networking that we want to do. code for you. Oh, that's clever. Callers could just invoke those remote functions in their code exactly as if they were local functions, and the generated code would handle the serialization and transit completely under the hood. So they basically wrote a program that writes the network program for you. And this is actually where the terminology itself comes from, right? Yes. They literally coined the term remote procedure call, or RPC.

9:47 Remote meaning the function lives on another machine, procedure being an older kind of academic term for a function. And call is the action you are taking, remote procedure call. Okay. But if Sun Microsystems solved this in the 80s, and we even still use their terminology today, why are we sitting here decades later starting a 22 chapter series on a modern framework? I mean, why did the industry keep churning out new versions? Because every generation of RPC systems that followed made painful and honestly sometimes fatal trade-offs between speed, language, flexibility, and complexity.

10:21 SunRPC was a brilliant foundation, but it was really just the beginning. Let's trace that evolution then. Because knowing why the older systems failed is so crucial to understanding why modern systems are built the way they are. What came after Sun? Well in the 1990s, the industry tried to build the ultimate universal standard. It was called CORBA. And the ambition of CORBA was just staggering. It promised that a program written in literally any language running on any operating system, could seamlessly call functions on any other system.

10:54 That sounds like an engineer's utopia. You write your front end in one language, your back end in another, your database in a third, and they all just magically communicate. Why didn't that win? Because of the crushing weight of abstraction. To accommodate every possible feature of every possible language, from like C++ memory pointers to small talk objects, Corba became a design by committee nightmare. Oh, I can only imagine. It required these heavy, complicated software layers called object request brokers just to function.

11:24 Developers ended up spending way more time configuring the Corba system than actually writing their business logic. It just collapsed under its own complexity. So the pendulum naturally swings the other way. The industry says, OK, a universal standard for every single language is just too hard. Let's simplify. What did the Java ecosystem do? Java introduced its own solution called Java RMI, Remote Method Invocation. And if you were working entirely in Java, RMI worked brilliantly. It was deeply integrated into the language.

11:51 You could pass complex Java objects back and forth over the network. And the JVM handled the serialization and garbage collection beautifully. But there is a massive catch there, isn't there? A huge catch. It only worked if both sides of the network were locked into Java. If your company decided to write a new lightning fast microservice in Python or, you know, a front end in JavaScript, Java RMI was totally useless. It sacrificed multi-language flexibility for a perfect single ecosystem experience.

12:18 So, yeah. It was a huge catch. Which brings us to the era of XML. Because the industry realized that locking into one language was just a dead end. We needed a universal language that any machine could easily read, regardless of the operating system or the programming language. So they turned to plain text, things like XML RPC and later SOAP. They decided to use XML as the serialization format. And from a purely human perspective, what was the advantage of that? Well, as a developer, the advantage was transparency.

12:47 XML is human readable. If a network call broke, you could just intercept the payload, print it out, and physically read the XML text with your own eyes to see exactly what data was being sent. It was highly debuggable. But this raises a really important question about performance. Because humans reading text is very different from computers reading text. Wait, I always assumed plain text was the easiest thing for a computer to process. I mean, it's just ASCII characters. Why would that be a performance bottleneck?

13:13 Because CPUs fundamentally hate parsing text. It's a huge problem. But when a CPU receives a binary integer, it just loads it into a register and does the math. Boom. Done. But when it receives an XML string, say, age 34 age, it has to do an enormous amount of work. Like what? Well, it has to scan the string character by character looking for the opening angle bracket. It has to allocate memory for the tag name. It has to do strm comparisons to figure out what field it's even looking at. Then it has to read the ASCII characters three and four and run an algorithm to convert those text characters back into an actual integer in memory.

13:48 Oh, I see. So for a single data point, the CPU is executing hundreds or even thousands of instructions just to translate it. And because XML wraps every single piece of data in opening and closing text tags, the payload size just balloons. You're sending huge text files over the network just to pass a few variables. Right. It bloated the network transit, and the CPU cycles required to parse those giant XML strings back into memory objects became a massive bottleneck. But if you look at the historical pattern, Java RMI was fast but locked to one language.

14:22 Corvo was universal but impossibly complex. XML was language agnostic but painfully slow. Nobody had cracked the perfect balance to solve all three nightmares at once. Not yet. Until we zoom in on what was happening behind the scenes while the rest of the industry was wrestling with these trade-offs, we have to talk about Google. Because in the early 2000s, Google was dealing with the RPC problem at a scale that literally no one else on earth had ever experienced. Yeah. If we look inside Google during that era, we're going to see a lot of things that are happening.

14:48 They had already embraced what we now call a microservices architecture. They weren't building massive monolithic applications. Their internal reality consisted of thousands of separate services constantly calling each other over the network. Right. Like the search index calling the ad server, the ad server calling the user profile database, the profile database checking permissions. It's just endless calls. Endless. And to survive that kind of internal traffic, they absolutely couldn't rely on massive XML payloads or the clunky frameworks of that era.

15:18 They needed something custom. So they built an internal framework from scratch and they named it Stubby. And Stubby was an absolute engineering workhorse. It handled the serialization of data efficiently. It managed the networking protocols. And crucially, it has sophisticated failure detection and load balancing built right in to handle the chaos of Google's data centers. And it wasn't just a prototype. For over a decade, Stubby quietly ran Google. It processed billions of RPC calls per second.

16:09 Google looked at this decade of hard-won lessons from running Stubby, and they decided to untangle it from their internal systems. They rewrote it, standardized it, and released it to the world as an open-source framework. And they called it gRPC. Which is, of course, the subject of our deep dive. Yes. And before we get into the mechanics of it, here's my absolute favorite piece of trivia. Everyone, understandably, assumes the G in gRPC stands for Google. But it officially doesn't. It really doesn't.

16:38 The development team changes the meaning of the G with every single release version. So on one version, the G stands for good. And the next release, it stands for green. Then groovy. They maintain a whole list of them in their release notes. It's hilarious. It is a fun quirk. But, you know, the technology behind it is completely serious. gRPC took those decades of historical trade-offs and made very specific modern choices to definitively address the three nightmares of remote calls. Let's summarize how it did that at a high level.

17:07 Keeping in mind, we're going to dive deeper. We're going to dive deep into the specific technologies in the upcoming chapters. Right. First, for the data nightmare serialization. Instead of using slow text-based XML or JSON, gRPC utilizes a highly fast binary serialization format. It strips away the bloat and turns data into tight CPU-efficient bytes. Which solves the CPU parsing issue XML had. Exactly. Second, for the transit nightmare, it runs on a modern, robust network protocol that handles multiplexing, meaning it can shoot thousands of messages.

17:38 It can also use back and forth over a single connection rather than opening and closing a new pipe every single time. What about the language lock-in problem? Does it force me to use Java like RMI did? Not at all. It is polyglot by design. gRPC natively supports multiple programming languages right out of the gate. You can have a Python script seamlessly call a function on a Java server, and the framework auto-generates the necessary code for both sides, just like SunRPC envisioned back in the 80s.

18:05 That's incredible. And finally, to address those unpredictable issues, we have a new program called gRPC. It has built-in mechanisms baked right into the framework to handle network timeouts, dropped connections, and retry logic. It really is the culmination of everything the industry learned from the 1980s onward. So what does this all mean for you, the listener? Let's look at the journey we've just taken. We started in your comfort zone — the nanosecond-fast, perfectly reliable world of a local Python or Java function.

18:32 And we shattered that comfort zone by throwing that function across a network, exposing the chaotic reality of differing memories, spaces, complex network protocols, and unpredictable physical failures. We saw how SunRPC, Corba, and SOAP tried to tame that chaos, and how Google's need to process billions of calls, a second-birth stubby, which eventually evolved into the open-source powerhouse that is gRPC. And this is just chapter one. Exactly. Now that the history and the fundamental problem are crystal clear, we have a roadmap for the rest of this series.

19:03 In the next chapter, we're going to dive into exactly how gRPC defines the data contracts between those services. We'll explore how it actually ensures Python and Java agree on what the binary data means. And as we move later in the series, we'll get incredibly hands-on. You will learn how to write the servers and clients, how to properly configure those failure mechanisms, how to secure your remote calls, and how to tune the performance for production deployments. It is going to be fantastic. It is going to be a fantastic ride.

19:34 But before we close out, we want to leave you with a final thought to mull over. Something that pushes beyond the software and into the, well, the philosophy of what we're building. Yeah, consider the sheer audacity of the remote procedure call. Every time a remote call succeeds, two entirely different computer universes with different memory structures running different languages, separated by miles of unpredictable physical cables, have managed to perfectly synchronize their reality for a fraction of a second.