Pi: A Coding Agent That Writes Its Own Tools
Transcript
0:00 You are deep in a coding flow. The agent knows the data flow. You know the next move. You're cruising. It just feels like magic. Right, until it doesn't. Exactly. The agent hits a wall. It needs a tool it does not have. Most harnesses give up at that point. They tell you to install a plugin, restart, and try again. And the problem is wasted turns and wasted context. A whole session that ends with, "I cannot do that yet." Today we'll look at Pi a coding agent that can extend its own runtime. We'll see how it closes that loop.
0:33 The agent writes a TypeScript extension, runs `pi install`, reloads the runtime, and the next LLM turn has the new tool. That is the payoff. The mechanism is what we are here to read. That phrase, self-extensible. Honestly, when I first read it I almost rolled my eyes. Oh, wow, another marketing word, right? Same. It sounds like marketing. What matters is the mechanism underneath, and whether the loop closes the way the tagline implies. Because every agent harness lets the model call tools. Calling tools is not extension.
1:08 The line worth drawing is between using tools you already have, and the agent reaching into its own runtime and adding a new one. Before we dig in, the source pack. This episode is built from the public `earendil-works` slash `pi` GitHub repository and the `pi.dev` site. Links are in the description. We read the code, we did not interview the maintainers. The thesis is small. Self-extension is, Four primitives wired correctly. Runtime TypeScript loading. A package manager at the front of the CLI.
1:39 Lifecycle hooks that touch the LLM call. And a `bash` tool that drives the package manager. Right. What makes Pi worth reading is that all four primitives are visible in code. Not just architecture diagrams that never shipped. Anyone who has built a plugin system knows the difference. Something that works, versus something that just looks good on a slide. The way these compose is portable. Even if you never use Pi, the design is worth borrowing. Now the distinction worth nailing first. Registering a tool at startup is the standard plugin shape.
2:12 You define the tool, you ship it, the harness loads it on boot. That is not extension that is just tool registration with a settings file. Exactly. Every coding agent today calls tools. The agent emits a tool call, the harness runs it, the result comes back, the model sees it. Standard. So extension means something stronger. New code that did not exist when the process started becomes part of the runtime new tools, new behaviors, new providers, no restart. From that first primitive the package manager.
2:46 The first place it shows up is the CLI itself. `pi install` is not buried in a settings menu. It is dispatched before the chat REPL even parses arguments. In `packages/coding-agent/src/main.ts`, the first thing the entry point does after migration is check for a package subcommand. Install, remove, update, list. By the way, only four tools ship built-in. Read, write, edit, bash. Wait, four? That's it? Painfully small, on purpose. Everything else is delivered as an installable extension. And the install surface is broad.
3:21 For example, the 5 source kinds are `npm:`, `git:`, `https:`, `ssh:`, and a local path. No proprietary Pi registry. A package is either an npm install or a `git clone`, on disk, under `~/.pi` or the project's `.pi` directory. Deliberately boring. From the install side to the package side. A directory becomes a Pi package by adding a `pi` field to its `package.json`. That field declares which resources the package contributes. Four resource kinds. Extensions, which are TypeScript modules. Skills, which are markdown files with frontmatter.
3:56 Prompt templates. And themes. Most of the weight is on the first two. One file, no manifest that's almost too convenient. For example, a tiny extension can be a single TypeScript file with no manifest at all. An `index.ts` at the package root is loaded as an extension by default. One level deep, no recursion. From the manifest into the loader here's where it gets really interesting. This is the primitive I think genuinely distinguishes Pi from every harness with a plugin folder. Most "drop a file in the plugins directory" systems expect pre-compiled JavaScript.
4:33 You ship `.js`, the harness requires it, that is the contract. Pi loads TypeScript directly at runtime. No build step. Inside `extensions/loader.ts`, the loader creates a `jiti` instance and imports the extension module. `jiti` is the on-the-fly TypeScript transpiler that Nuxt uses. Same idea here. Hand it a `.ts` path, get back an executed module. That changes the contract. Pi also ships as a Bun-compiled standalone binary. Self-contained. No `node_modules` at runtime. So how does a user's extension import from the Pi SDK?
5:07 This is the clever bit. The loader has a dual mode. When Pi runs as a Bun binary, it statically imports its own SDK packages at build time and exposes them as `virtualModules` inside `jiti`. Think of it as sailing a ship across an ocean and 3D-printing a new pump for the engine room mid-voyage the new code never has to look for a physical `node_modules` folder because the runtime is satisfying every dependency virtually. So when an extension does `import` from `@earendil-works/pi-agent-core`, `jiti` resolves that import against the bundled-in module, not against the filesystem. So `jiti` does the runtime resolve and the SDK bridge in one path. That's the clever part.
5:49 From the loader, then, to what the extension actually gets. When `jiti` loads an extension module, the call returns a factory function. The factory takes one argument, which the codebase calls `pi`. That `pi` argument is the `ExtensionAPI`. Everything an extension can do, it does through this object. The most obvious method is `registerTool`. It takes a tool definition with a name, a description, a TypeBox parameter schema, and an `execute` function. For example, imagine wrapping your team's internal deploy CLI you know the kind, the one with three required flags and one positional argument nobody can remember.
6:24 You describe the schema, you write `execute`, you ship it. Right. TypeBox is like an incredibly strict bouncer at the door the schema gets validated before the function ever runs. Same machinery the four built-in tools use. That alone would be a respectable plugin system. Same idea, extended. If Pi stopped at `registerTool` it would already be ahead of most plugin systems. But `registerTool` is just the door. The interesting weight is in the events. Roughly 24 lifecycle hooks an extension can subscribe to.
6:55 24 is a lot of surface area. Some are obvious the session-start hook, the before-agent-start hook. Standard plugin lifecycle. The ones I want to surface are the four that fire around the model round-trip. The context hook lets the extension mutate messages before they are sent to the model. The before-provider-request hook rewrites the raw HTTP payload, after the framework assembled it but before it hits the wire. The tool-call hook intercepts any tool invocation, including the built-in ones, and patches or blocks the arguments.
7:28 The tool-result hook mutates the result before the LLM sees it. That is a much higher ceiling than a plugin system. Closer to middleware. And it lands in-process, not out-of-band. So an extension can rewrite the LLM payload before it goes out and mutate the response after it comes back, against the actual runtime types. No IPC boundary, no JSON serialization tax. That is what `jiti` buys you upstream. That leads to the other extension method worth surfacing `registerProvider`. An extension can register a brand-new LLM provider at runtime.
8:02 Not just a base URL? No. A full provider record. Models, capabilities, transports, and, optionally, an OAuth flow. That is uncommon. Most harnesses let you configure a provider in a settings file. Pi lets you ship a provider as an npm package, complete with `login`, `refreshToken`, and `getApiKey` callbacks. For example, install one package, and your corporate SSO-backed LLM shows up next to the built-in providers. Right. And it works at the same layer as the built-ins, so the abstraction is honest.
8:36 Let's hear it. Can the agent really write a tool, install it, and use it in the same session? Code-wise, yes. Step one: the agent uses the `write` tool to put a TypeScript file on disk. Let's say a small extension wrapping an internal HTTP API your team uses. OK. The file is on disk. Walk through what the install actually does at that point. Step two: the agent runs `pi install ./that-path` through the `bash` tool. Same `bash` tool you would use yourself. The package manager resolves the local source, optionally runs `npm install`, and registers the path. Step three: the agent calls `/reload`, which and this is the important part re-runs the resource loader without restarting the process. So the loader sees a new package on its next pass.
9:24 What does it actually do with it? Step four: it walks the registered packages, finds the new one, calls `jiti` on the entry point. The factory runs, `registerTool` fires, the runtime knows the tool exists. Step 5: next LLM turn, the tool is in the model's list. That is real. The primitives exist. The code is straightforward to follow. If you wanted to demo "the agent built itself a new capability," you could do it inside Pi today. Now the honest question, given that the loop closes. The runtime is set up to extend itself, but the prompt does not push the agent to use it?
10:02 Right. The primitive is there. The autonomy is not, at least not in the default system prompt. Yeah. The machinery is built, the engine is running, but the agent has not been told it holds the keys. Hm. Strong image. We searched. No built-in instruction tells the model, "if you need a tool you do not have, write it and install it." The capability lives in the runtime. The default behavior does not nudge the agent toward it. So when the README says self-extensible, that is a description of the runtime, not a description of agent behavior.
10:34 Right. A user can drive the loop. A skill the user installs can drive the loop. The agent on its own, with the shipping prompt, will not. I cobbled together an autonomous-tool-author skill in my head while reading this code, and it was, like, 5 lines of markdown. That is the size of the gap. Tiny. But the maintainers did not ship it as the default, which I think is a real design call, not an oversight. Same idea on the prompt side. The second extension surface in Pi is skills, and these are not TypeScript at all.
11:05 A skill is a `SKILL.md` file with YAML frontmatter. Name, description, maybe a flag to disable model invocation. The body is markdown. Pi loads it through the same package system, scoped uh, the same way as extensions, basically but no code runs. Skills get injected into the system prompt as procedural guidance. For example, a skill could say, "when the user asks for a deploy status, follow these steps." Take that one example and you can see why skills and extensions are a clean split. Extensions extend the runtime.
11:38 Skills extend the behavior. Both ship through the same `pi install`. Stepping sideways once, the AI layer. It shows up everywhere upstream, but it is not the main act today. `packages/ai` is Pi's multi-provider abstraction across Anthropic, OpenAI, Gemini, and a long tail of self-hosted providers. The Pi-level abstraction is one knob, `ThinkingLevel`. The enum has 5 steps from minimal to xhigh. Each model maps its own translation to the provider's native shape, so one knob in Pi covers budget tokens on Anthropic and reasoning effort on OpenAI and thinking on or off for Gemini.
12:16 Application code never has to know which provider it talks to. After the AI layer, here's the piece worth porting. Pi sessions are JSONL files. One entry per line. Append-only. Every entry carries a `parentId`. So the session is a tree, not a flat transcript. Wait, really? So you never overwrite history? Right. Every change adds a new entry and the old one stays. Branching is moving a leaf pointer. Forking across projects copies history but does not rewrite. Compaction is just a new entry whose parent points back, not destructive.
12:52 I love that. For example, there are entry types beyond plain messages. `ThinkingLevelChangeEntry`. `ModelChangeEntry`. `CompactionEntry`. `CustomEntry` for extensions to persist private state the LLM never sees. `CustomMessageEntry` for state that gets fed back into the LLM next turn. If you were going to steal one design from Pi for your own agent harness, I would argue this is the one. Append-only JSONL with parent pointers beats "rewrite the whole transcript every turn" on every axis. After the big stuff, one infra footnote because there is something funny about a coding agent shipping a fix for a Node fetch timeout, and we are obligated to admire it.
13:35 In `cli.ts`, on the very first lines after the process title is set, Pi configures the `undici` HTTP dispatcher. It disables the default body and headers timeouts entirely. The kind of thing you only write once you have been burned by it. Picture a vLLM instance buffering a big tool call. The response can stall for minutes before the next byte. Node's bundled fetch gives up after 5 minutes. Without that fix, a long agent turn would be a deterministic nightmare on self-hosted endpoints. Three-line fix.
14:07 Real scar tissue. Let us close the way the series likes to close. Imagine you are building your own agent harness. Which Pi primitive do you take first? For me, the session model, with append-only JSONL and parent links. Small, universal, and it makes branching, forking, compaction, and custom entry types easier. That design carries its weight. I would take the lifecycle event surface. The four lifecycle events that intercept the model call. Context, before-provider-request, tool-call, tool-result. They are what turn extensions into middleware instead of plugins. And honestly, you cannot have those without the session model first, because extensions need somewhere to persist state across reloads.
14:50 Fair. Session substrate first, then the middleware surface. And the failure mode you inherit with both is the honest gap. You have built a runtime that can extend itself. Whether the agent does extend itself is a prompt question. Do not ship that confused. Now the close. Source links for the `earendil-works` slash `pi` repository, the `pi.dev` site, and the specific files we walked through are in the description. The takeaway is the one we opened with. Self-extensible is not a single feature. Right. It is `jiti` plus an install command at the top of the CLI plus a middleware-shaped extension API plus a session model.
15:31 Four primitives. Each one portable. That is how a marketing phrase decomposes into engineering. Thanks for listening to Learning Podcasts.