The Leaked Claude Code: Architecture Tour
Outline
- 0:00 A leaked source and a paper
- 0:47 The design-space map
- 1:47 Five guiding values
- 2:48 Permission system as a design choice
- 4:28 Context pipeline, five layers
- 6:05 Four extensibility surfaces
- 7:30 Subagents and sessions
- 8:47 The OpenClaw contrast
- 10:06 Six open design questions
- 11:14 The design-space map, payoff
- 12:03 Closing and link
Transcript
0:00 Welcome to Learning Podcasts. Inside Claude Code. On March 30-first, Anthropic shipped a Claude Code build to npm with a packaging mistake. They forgot to exclude the source map files. And the source maps pointed at unobfuscated TypeScript in an R2 bucket. So when people followed the references, nineteen hundred files, five hundred 12 thousand lines of Claude Code's actual source were sitting there. Public. Two weeks later, four researchers had read the leaked source and written this paper. The paper says the tool you use every day is a while loop.
0:35 They say it out loud. A while loop. Yeah. They went into the leaked TypeScript and that is what they found at the center. Call the model, run the tools, repeat. And then they spent the rest of the paper mapping everything else. The paper is *Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems*, by Liu, Zhao, Shang, and Shen, MBZUAI VILA Lab. arXiv 26 oh 4.14 two two 8. The link is in the description. And the rule for the episode is simple. Every load-bearing claim is the paper's claim.
1:10 The paper's claims come from the leaked source. We are walking through both. Right. So the map. The center of the diagram is the loop we just named. Call the model. Run the tools. Repeat. That is the agent. Everything else is the layer around it. The paper organizes those layers into 5 regions. Permissions. The context pipeline. Extensibility. Subagents and sessions. And the comparison with OpenClaw, the open alternative the authors hold up as the contrast point. So one diagram. The loop at the center, 5 layers around it, one contrasting harness pinned at the edge.
1:42 That is the whole episode. Right. And the loop is honestly the least interesting part. The design space is the layers. Before we visit the layers, the paper anchors them on 5 guiding values. These are the values the authors say drive the design. And they read them out verbatim. Right. The 5 are: Human Decision Authority. Safety, Security, and Privacy. Reliable Execution. Capability Amplification. And Contextual Adaptability. Huh. Reliable Execution feels like the load-bearing one to me. The others are guardrails on that.
2:15 Reading these out loud, they sound suspiciously like a corporate mission statement. They do. It is easy to dismiss them as buzzwords. But the paper is making a stronger claim. The values are architectural blueprints. Think of it like a city's founding document. If the document says we value public health, that is not just sentiment. It dictates how the plumbing is laid out and where the water filtration plants go. The philosophy dictates the infrastructure. So the values are not decoration. They map to engineering.
2:45 That is the claim. The rest of the episode is those values made concrete. First layer is the permission system. The paper identifies 7 modes in the source. plan, default, accept edits, auto, don't ask, bypass permissions, and a 7th called bubble. 7 modes is more than I expected. Are they all surfaced to the user? Not all of them. Most are user-facing settings; bubble is internal. But the framing the paper gives us is more interesting than the list. The headline design choice is per-action approval.
3:18 Not perimeter access control. Per action. Right. And the most interesting mode is auto. Auto runs an ML-based classifier on each action and decides whether to ask. The safety check is itself a model, not a rule. Wait, really? The safety check is a model? Right. Per call. Hold on. If I told my infosec friends that the firewall is basically vibes-based, they would lose their minds. Yeah, they absolutely would. How is an ML classifier not just a worse, less predictable rule engine? Hmm. The paper acknowledges that exact tension.
3:52 It is a valid critique from the infosec community. The paper does not claim to fully resolve it. It just says this is the deliberate architectural direction. Auto is feature-gated for now, but per-action ML approval is where the architecture is heading. So instead of a static allowlist that an operator wrote, the decision is delegated to a classifier per call. And bubble is the other interesting one. Bubble is how a subagent escalates a permission decision back to the parent, instead of resolving it inside its own context.
4:22 So the parent stays the decision-maker even when the work is delegated. Human Decision Authority, applied through the harness. Second layer, the context pipeline. This is the part every Claude Code user has the most tacit feel for, because every user has watched a session compact mid-task. Yep. Trading fidelity for headroom is the whole game. I think about it like a chef trying to cook a 10-course meal on a tiny cutting board. As the meal gets more complex, you run out of physical space. You chop things smaller.
4:54 You sweep things away. Just to keep cooking. The 5 layers are the exact compromises the system makes. Oh, I like that. The paper unpacks it into 5 layers that run in order before every model call. Layer one, budget reduction. Per-message size limits on tool results. Layer two, snip. A lightweight temporal trim that drops older history. Layer three, microcompact. Fine-grained compression keyed by tool use ID. Layer four, context collapse. A read-time projection over the whole conversation. And layer 5, auto-compact. The last resort.
5:29 A model-generated semantic summary of the session. I have watched auto-compact eat my plan three times in one afternoon. It is doing exactly what it should. It is still painful to watch. Yeah. Eaten my plan more than three times. And the paper is honest about it. Each layer trades fidelity for headroom, but auto-compact is the one with no take-backs. It is the only thing standing between you and a hard context error, though. Exactly. And the reason there are 5 layers and not one is that the cheap layers run first.
5:59 Auto-compact only fires when the others were not enough. Reliable Execution, from the values list, is what the whole stack is paying for. 3rd layer, extensibility. The paper names four mechanisms, and the distinction between them is sharper than it looks at first. MCP. Plugins. Skills. Hooks. MCP is external tool integration. A server somewhere exposes a tool surface that the harness can call. Plugins are packaged bundles of components for distribution. Skills are domain-specific instruction injection, loaded from the dot claude skills directory. And hooks? Hooks are lifecycle interception.
6:36 Think of a hook like a developer tripwire. You tell the system, the exact millisecond the agent realizes a tool failed, pause everything and run this specific script. The paper counts 27 distinct event types you can hook into. 27 tripwires the harness lets developers inject custom logic into. 27? That is a lot of foot-guns. We have all written that hook that turned out to fire on a partial tool retry too. Do those 27 differentiate tool error from tool retry? They do. PreToolUse and PostToolUse fire on every call.
7:10 There are separate events for tool error and for tool denial. So error versus retry is a hook-level decision, not a single-event one. Got it. Four different ways to bend the harness, cleanly separated. Capability Amplification, from the values list. Each one amplifies a different surface. MCP for tools. Plugins for distribution. Skills for instructions. Hooks for the lifecycle. 4th layer, subagents and sessions. The paper treats these as separate subsystems, but the design point is the same. Same design point? They look like different problems to me. One is parallelism, one is history. How are they the same?
7:46 Both subsystems trade more storage and more bookkeeping for stronger isolation and recovery guarantees. Subagents first. The worktree functions like an alternate timeline. The parent agent stays in the prime timeline. The subagent goes into the alternate copy and runs. It might introduce a bug. It might delete the wrong directory. It might fail entirely. But the parent does not see any of that intermediate chaos. Its context stays clean. If the subagent succeeds, the parent pulls the finished result back.
8:16 If it fails, the parent incinerates the alternate timeline like it never happened. Blast radius contained. Right. Sessions. The session log is append-oriented. Every event is appended. Nothing is edited in place. Which means recovery is the same shape as the log. Replay forward to wherever you want to be. Two subsystems, one design tradeoff. Exactly. The paper does not say it in these words, but the pattern is the same. Pay more in storage and bookkeeping. Earn stronger isolation and replay guarantees.
8:45 Here's where it gets really interesting. The 5th region is not really a layer. It is the contrast. The paper compares Claude Code to OpenClaw, an open alternative agent harness that the authors hold up as a real point of contrast. And the framing is what makes the comparison earn its place. The paper's headline contrast is: Claude Code prioritizes per-action safety classification. OpenClaw emphasizes perimeter-level access control. Right. Same design space. Different choice. I think I have read every paper on agent harnesses, and I did not have this axis labeled until now.
9:21 Every harness I have shipped lives somewhere on it. Yeah. That is the gift of the paper. Per-action approval means every call asks the harness, "Should I do this." Perimeter approval means the harness sets a boundary at the start of the session, and anything inside the boundary runs. Both are valid. The deployment context decides which one is right. Claude Code is a developer's daily driver, where the user is present and per-action is acceptable. And the operators running Claude Code in CI are already saying per-action is the wrong fit for batch.
9:55 They are right for that deployment. The paper says so without saying so. Two points in the same design space. Both are real engineering. Then the paper closes with 6 open design questions. The authors think the next generation of agent systems has to answer them. 6 is a lot. Are they ranked? Let's hear it. They are paired more than ranked. First pair. Silent failure and observability gap. Detecting when an agent action produced no visible feedback. Paired with cross-session persistence. Memory and longitudinal context across sessions.
10:28 Both about what the harness can see over time. Right. Second pair. Harness boundary evolution. Where, when, what, and with whom the agent acts. Paired with horizon scaling. Expanding from a single session to a scientific program. So harness boundary evolution is the question of whether the agent gets to act in your repo, your inbox, or your bank account. Got it. 3rd pair. Governance and oversight at scale. Managing autonomy as deployments expand. And long-term human capability preservation. Balancing amplification against skill atrophy.
11:00 Both about what humans hold onto as the agents do more. Right. Three pairs. What the harness sees. How far it reaches. What humans keep. The paper names them, the paper does not editorialize, and neither will we. Back to the map. The center is still a while loop. The model is called, the tools run, the loop repeats. That is the agent. Around it, 5 regions. Permissions. The 7-mode taxonomy that lets per-action approval do the work. Context pipeline. 5 layers that pay for Reliable Execution by trading fidelity for headroom.
11:33 Extensibility. Four mechanisms, each amplifying a different surface. And subagents and sessions, paying the same price for the same guarantee. Storage and bookkeeping for isolation and replay. And then the OpenClaw contrast. Per-action versus perimeter. And the 5 guiding values still anchor the whole map. The values are not decoration; they are why the layers look the way they do. Right. The paper's real gift is the axes. Next harness you read about, those 5 regions are the questions to ask. One last thing. The paper is short. Go read it.
12:04 The link is in the description. arXiv 26 oh 4.14 two two 8. It is worth an hour. The framing pays off well beyond Claude Code itself. One last thought, building on the last question. We are pouring engineering into isolated worktrees and context pipelines so the machine can keep iterating. Meanwhile, the human's role is shrinking to sitting at the perimeter as a permission node, clicking approve. Yeah. That follows from the design space the paper lays out. So if we are engineering the machine to do all the iterating, what happens to our own cognitive loop?
12:39 Do we lose the intuition for problem solving when we are not the ones standing in the arena making the mistakes? That is the question the paper opens up without answering. Right. Thanks for listening to Learning Podcasts.