RTK (Rust Token Killer): Up To 90% Token Savings For Claude Code And Codex
Outline
- 0:00 The command-output boundary problem
- 1:23 Why users should care first
- 2:28 Basic RTK usage
- 3:38 What users should feel
- 4:41 What not to filter
- 5:39 Repo structure
- 6:32 Router as a contract map
- 7:35 Runner layer
- 8:35 Stream layer
- 9:41 Filtering moves
- 10:40 Filter contracts as code
- 11:34 Test output
- 12:29 Git output
- 13:31 Search, read, and logs
- 14:33 Claude Code hook integration
- 15:42 Rewrite internals
- 16:33 Codex `AGENTS.md` guidance
- 17:38 Built-in tool bypass
- 18:29 Custom TOML filters
- 19:21 TOML filters as an extension point
- 20:31 Raw-output recovery
- 21:35 Tracking and feedback loops
- 22:25 Discover and learn
- 23:08 Privacy and telemetry
- 24:10 The lossy contract
- 25:03 Rollout guidance
- 26:01 Rebuild order
- 26:57 Designed output for coding agents
Transcript
0:00 A coding agent asks the shell for facts. The shell answers with a transcript meant for humans. The problem is that noisy output becomes model input, and the next action can go wrong. Today: RTK, inside the output proxy for coding agents. We'll see how it turns shell noise into a designed input boundary. You already know the workflow. Claude Code runs tests. Codex checks a diff. An agent shells out to Git, grep, Docker, Kubernetes, package managers, and whatever internal CLI your team invented. And every one of those commands talks like a terminal tool.
0:35 It prints progress bars, ceremonial headers, repeated logs, passing-test reassurance, and 5 useful lines hiding inside two hundred boring ones. The model reads all of it. That is the weird part. Shell output used to be a human interface. In an agent session, it becomes model input. This episode is built from the public `rtk-ai` slash `rtk` GitHub repository, a pinned local source snapshot, and the official RTK guide. Links are in the description. Good. Because the interesting question is not "does this save tokens?" The interesting question is "what should a command result look like when an LLM is the reader?" Exactly. RTK turns that question into an inspectable local system, not just a slogan about smaller logs.
1:22 Let us start with users before internals. Yes. The user-facing promise is context hygiene. Bigger context windows give capacity. RTK tries to stop low-value terminal noise from entering the conversation in the first place. Right. That matters even if your model has a huge window, because context is not just storage. It is attention, latency, cost, compaction pressure, and future confusion. Exactly. Future confusion is the part engineers underestimate. A raw line that mattered 5 turns ago sits beside 40 lines that did not matter, and later the agent has to remember which one was signal.
2:02 Right. For example, imagine a test run where one assertion explains the bug, but the transcript also includes setup logs, coverage tables, and every passing test name. Exactly. The model can technically read all of that. The problem is that you made the reasoning object worse. Right. So RTK is not only a token saver. It is a way to make command output less distracting for the agent. From that user promise, the direct usage model is simple. Instead of `git status`, an agent or human can run `rtk git status`.
2:37 RTK runs the real Git command, captures the output, filters it, preserves the exit code, and prints a compact result. Same idea for tests. For example, `rtk pytest` should keep the failing test names, tracebacks, assertion details, and final summary, while suppressing the parade of passing noise. Yeah. And for logs, the value is obvious. If a service prints the same error a hundred times, an agent does not need to read the scream a hundred times. Right. The official docs list many optimized command classes: Git, test runners, package managers, Docker, Kubernetes, cloud CLIs, search, file reading, and more.
3:20 Do not memorize the catalog. The category is the point: high-frequency, semi-structured command output where the useful facts are predictable. The usage rule is not "always make output smaller." It is "give the model the output shape the task actually needs." So what should it feel like in day-to-day agent work? First, less repeated orientation. The agent should not spend a 3rd of a turn rediscovering that Git has staging instructions or that pytest prints passing tests. Second, sharper follow-up actions.
3:53 For example, if the output says one test failed in one file, the next tool call should probably inspect that file, not scroll through terminal ceremony. 3rd, better long-session behavior. If a session runs 20 commands, filtering each command can reduce the amount of stale junk that later compaction has to summarize. Right. So the benefit is not only the one command. It is the whole chain of decisions after that command. Exactly. A cleaner command result changes what the agent thinks the task is.
4:25 Right. That is the user-facing pitch I like: fewer detours, fewer rereads, fewer giant transcripts sitting in the model's working memory. Right. And because exit codes still matter, the agent should still know whether the underlying command passed or failed. That makes the next warning easier to say: not every command deserves compression. The other user-facing lesson is restraint. Some output is the artifact. For example, if you ask a command to print a generated config file, a legal notice, a migration SQL script, or a complete JSON payload you plan to paste somewhere, do not casually summarize it away.
5:02 Same for audits. If the task is "review every warning," then a filter that hides warnings is not a helper. It is a bug. That is the senior-engineer rule. Filter commands where the information shape is predictable. Keep exact output when exact output is the point. I have shipped enough "helpful" wrappers to be suspicious here. The failure mode is always the same: the wrapper looks great until the one omitted line was the whole incident. Yes. RTK is strongest when it is treated as an output contract, not a universal compression spell.
5:34 Exactly. Smaller is not the goal. Task-relevant is the goal. Now the internals. The repo shape matches the product story pretty cleanly. Here's where it gets really interesting. `src/main.rs` is the front door. It is a Clap command router with direct command families, generic proxy modes, hook setup, rewrite checks, verification, gain reporting, discovery, and learning commands. So RTK is not one giant summarizer sitting at the end of stdout. It is a dispatcher plus ecosystem-specific command modules.
6:07 Git has its own module. System read behavior has its own module. Other tool families get their own handlers. That matters because command output is not one language. Take a Git status, a pytest traceback, and Kubernetes logs. They have different contracts. Right. The architecture you should picture is parse, route, execute, filter, print, track. Nice. That is the whole episode in 6 verbs. Let's make that router idea concrete, because it tells you where RTK's confidence lives. The router is not just plumbing.
6:39 It is a map of which tools RTK believes have stable enough shapes to optimize. A command family gets a handler because its output has recognizable semantics. Git status has states. Test runners have failures, tracebacks, counts, and summaries. Logs have repetition. Package managers have progress and dependency summaries. If a command is unknown, the safe behavior is passthrough or generic proxy tracking. RTK should not pretend it understands a tool it does not understand. That is important. The absence of a filter is also a design decision.
7:15 Exactly. In source terms, the router is where the product boundary starts. It decides whether RTK has a command-specific contract or is just a wrapper around the real command. Right. And for engineers evaluating the repo, that is a useful audit path. Look at whether your favorite command has a real module or mostly falls back. Once the router says this command has a supported shape, execution mostly goes through core slash runner. That layer matters because RTK has to preserve command semantics while changing the view. Right. There are filtered, streamed, and passthrough paths. Unsupported commands can pass through.
7:53 Some commands need line-by-line streaming. Some can be collected and filtered after exit. Right. And the exit code is not decoration. If the real command fails, the wrapper must still communicate failure to the agent and shell. Right. That is the difference between an output proxy and a cute pretty-printer. A pretty-printer can lie accidentally. A proxy has to preserve the operational contract. I like that. The output is lossy, but the command result cannot become fake. Exactly. For example, if a test command exits nonzero, the filtered output can be shorter, but the wrapper still has to behave like a failed test command.
8:31 Right. That is the invisible promise the user depends on. The runner explains command semantics; the stream layer explains the live data path. The next layer is where the proxy stops being abstract and starts looking like process plumbing. `core::stream` is the plumbing layer I would copy early if I were rebuilding this. Same. It spawns the child process, captures stdout and stderr, preserves exit behavior, and can feed lines through a streaming filter while the command is still running. Right. That is important for long-running commands.
9:06 You do not always want the agent to wait for a full blob before seeing anything useful. Exactly. A build can stream the meaningful failure lines and still emit a final compact summary. Right. The tricky part is that streaming filters have less hindsight. A final filter can see the whole output. A streaming filter has to make decisions as lines arrive. Which is why command-specific knowledge matters. If you know the shape of a tool, you can stream a useful subset without pretending every line is equal.
9:35 That is the architectural theme again: output contracts beat generic vibes. So once output is captured, what does a filter actually do? Oh right. The filters mostly win through four moves. First, remove ritual: progress bars, repeated headers, ANSI decoration, passing-test confidence theater. Second, group repeated facts. For example, 10 similar lint errors can become a count plus representative locations. 3rd, truncate safely. Keep the head, the tail, the failure region, or a representative slice instead of dumping the entire stream.
10:11 4th, deduplicate. Repeated logs become one line plus a count, not a context-window denial-of-service attack. The word safely is doing work there. Bad truncation is worse than no truncation. Yes. The filter has to preserve the facts that drive the next action. File path, line number, assertion, exit status, changed files, failing command, affected resource. Exactly. Smaller output is only useful if the model can still decide what to do next. After the safety rule, the internals start to look like product judgment.
10:44 I want to sit on that word contract for a second. Please do. A filter is not just "delete boring lines." It is a promise about what information survives. In a test runner, the contract is: keep failing tests, error messages, file paths, line numbers, traceback context, and final counts. With Git, the surviving facts are branch, staged changes, unstaged changes, untracked files, and push or commit outcome. For logs, I want unique errors, counts, timestamps when they matter, and enough surrounding context to identify the component.
11:20 That is why deterministic filters make sense here. An LLM summary could sound plausible and still miss the one field the next tool call needs. RTK's bet is that many command outputs are structured enough to filter with code, not vibes. The clearest version of that judgment is still test output. The easiest proof point is a test failure because the useful shape is so familiar. A human may want reassurance that a thousand tests passed. An agent fixing a failure usually needs the failing test, the assertion, the traceback, and the summary.
11:55 Passing tests are emotional support for humans, especially when the agent already has the failure it needs. I like that phrase. It is useful sometimes, but often not the thing the model needs. Yeah. I have watched agents read pages of passing output, then miss the one assertion that actually explained the bug. Same. And once that junk enters context, it competes with the real signal in later turns. Right. RTK's test-runner story is basically: keep the failure surface, collapse the celebration. Show the failing assertion and final count, not every green check.
12:27 Exactly. The model does not need a parade. It needs the fire exit. The other high-frequency case is Git. Right. Status, diff, log, show, stash, branch checks, add, commit, push. It is the plumbing agents keep touching. Yeah. And Git output is semi-structured, stateful, and full of ceremony. A human terminal session benefits from reminders and progress. An agent mostly needs what changed, where it happened, whether it worked, and what branch it is on. Concrete case: push output, enumerating objects, counting objects, compressing objects, writing objects. Very real work, usually not the important model input.
13:09 Right. And `git status` can repeat staging instructions every time. Useful for a beginner. Less useful for an agent that has seen it a hundred times. Right. So Git is not just a savings example; it is a frequency example. A small bad shape repeated all session becomes a large context tax. Right. That is why output hygiene matters more as agents become more autonomous. After Git, search, read, and logs have the same shape, but the failure modes differ. The same idea shows up in search, read, and logs, but the risk changes.
13:44 Search is tempting because `rg` can return a lot. The useful result may be file names, line numbers, and a few matches, not every surrounding line. Reading files is trickier. Sometimes the agent needs the whole file. Sometimes it needs a compact preview. The consumer matters. Logs are the clearest case. Repetition is useful to count, not useful to read line by line forever. Imagine a container log where the same connection error appears two hundred times with slightly different timestamps. Right, the agent needs a gauge and a sample bottle, not the whole river in the prompt.
14:23 Exactly. Count the flood, preserve the error, and keep raw recovery nearby when the compact view misses the clue. Right. And when the flood itself matters? Raw recovery has to be close. With that usage model in mind, the strongest integration path is the hook. Wait, really? Why is the hook stronger than just remembering to type `rtk`? Because the environment can rewrite the shell call before it runs. For Claude Code, the documented path is a PreToolUse hook. The hook sees Bash tool calls before execution and asks RTK what the command should become. So Claude can request `git status`, and the environment can turn that into an RTK-filtered Git status before the shell actually runs it.
15:10 The key code detail is `rtk rewrite`. The hook adapter is not supposed to reinvent command policy. It calls the binary. And the rewrite command goes through the discovery registry, so rewrite logic is centralized instead of copy-pasted across hook integrations. It also has permission-aware outcomes: allow the rewrite, pass through, deny, or ask. That boundary is important because hidden command rewriting is powerful. A proxy changing shell commands needs a real boundary. "Trust me, bro" would be a nightmare architecture.
15:44 The rewrite path is worth separating from the filtering path. How so? Filtering asks, "given this command output, what compact result should we show?" Rewrite asks, "before execution, should this shell command become an RTK command?" Ah, right. One happens before the command runs. The other happens after or during execution. Exactly. The discovery registry is where RTK can recognize supported commands and decide whether a rewrite is appropriate. Right. And that is where permission outcomes matter.
16:15 Allow, passthrough, deny, and ask are different operational states. Right. This is why putting rewrite policy in the binary matters. If every hook had its own version, the Claude integration, Gemini integration, and future integrations could drift. I love that. The hook should translate the envelope. The binary should own the policy. Codex is the contrast case, and it changes how you should evaluate adoption. Here, the integration boundary moves up into instructions. RTK's Codex story is mostly `AGENTS.md` guidance.
16:48 The repo can tell Codex: prefer RTK for noisy commands, use RTK for Git or tests, avoid raw dumps when filtered commands exist. That helps, but it is not the same as a pre-execution rewrite. Right. Adoption is instruction-level. The model has to remember the preference. Right. For a Codex user, the evaluation should be behavioral. Look at the command log. Did it actually use `rtk git status`? Did it use RTK for test output? Did it fall back to raw shell commands? My first thought with Codex would be: prove it in the log. I would not trust the instruction file until I saw the commands it actually ran.
17:30 Yeah. The tool can still help Codex users without those same hook-level guarantees. So this is not anti-Codex. It only marks where the integration boundary sits right now. That contrast is useful, but it can create one bad expectation. Wait, what? Are we saying RTK does not compress everything Claude Code sees? Claude Code built-in tools like Read, Grep, and Glob do not pass through the Bash hook. Correct. RTK mainly applies to supported shell-command output, or explicit RTK commands such as `rtk read` and `rtk grep`.
18:04 Right. That is not a failure. It is a channel boundary. Built-in agent tools are a separate input path. The practical rollout language should be precise: RTK shapes command output. It does not magically shape every model-visible artifact. That precision matters when people start debugging why a huge file read still filled the context. Exactly. If the path is not shell output, do not assume a shell proxy touched it. Here is the team-level version of the idea. The part I like for teams is the custom filter story.
18:37 Same. Every serious engineering org has internal CLIs, weird test wrappers, release tools, and one script named `do-everything`. Always `do-everything`. RTK supports TOML filters so project-specific tools can get maintained output contracts without writing Rust for every case. For example, project-local `.rtk/filters.toml` can override user-global filters, and user-global filters can override built-ins. Right. That priority order is powerful, but it needs a trust gate because a repo-local filter changes what the agent sees.
19:09 Right. The `rtk verify` path is important too. If a filter becomes part of your agent workflow, test the contract. Exactly. That is the grown-up version: do not just delete lines. Specify what information the command is supposed to carry. The TOML filter path is a nice architectural compromise. For example, not every team wants to write Rust just to clean up one internal command. Exactly. A declarative filter can strip ANSI, replace text, keep or drop matching lines, take head and tail slices, cap line counts, and provide an on-empty message.
19:44 Right. That is enough for a lot of boring but valuable cleanup, but it also creates a governance question. Right. If a repo can define how command output is shown to the agent, the repo can hide things. Yeah. I can hear a platform lead saying, "Hold on. I do not want a repo-local filter hiding a production warning from the agent." Right. That is where the trust gate stops being configuration trivia. Project-local filters are powerful because they travel with the repo, so the trust decision has to travel with them too.
20:16 Right. Otherwise a helpful filter becomes an invisible policy layer. Exactly. The design lesson is broader than RTK. Any local agent runtime that lets repos modify tool output needs a trust boundary. Exactly. That boundary is the point. The same safety logic shows up again when filtered output hides raw detail. Lossy output needs an escape hatch. I love this part because the recovery path is what makes the lossy part tolerable. Non-negotiable. The tee layer exists so RTK can save raw output locally and print a path the agent can inspect.
20:53 That is the right default. Filtered output should be the first view, not the only view. Debugging often starts with the summary and then turns on the weird detail. The weird detail is usually where production lives. For example, a compact failure summary is useful until one omitted line is the real clue. Right. RTK also supports disabling or excluding commands, and verbosity can show what the filter is doing. Right. So the safe posture is not blind trust. It is filter by default on predictable noisy commands, keep recovery cheap, and teach the agent when to open the raw output.
21:28 Exactly. If recovery is awkward, people stop recovering. The raw path has to be obvious, not buried behind a debug ritual. I love that the repo does not stop at filtering. Yeah. The local tracking system is another architecture clue. RTK records command labels, estimated original and filtered tokens, saved tokens, execution time, timestamp, and project path in local SQLite. That powers `gain`, so savings are connected to real command execution instead of only benchmark claims. Then `discover` looks at Claude Code session logs to find missed RTK opportunities.
22:05 And `learn` mines command-error history for correction patterns. So the system is not just a filter library. It is trying to close the loop around agent terminal behavior: what got filtered, what was missed, and what mistakes repeat. That is the product move: output shaping stops being a one-time wrapper. It becomes something you can inspect and improve over time. That feedback loop is why the `discover` and `learn` commands matter. Because a pure wrapper would stop at `rtk command`. Exactly. Discovery asks, "where did agents use raw commands that RTK could have helped with?"
22:41 Right. And learning asks, "what command mistakes keep happening?" Exactly. That turns real agent sessions into feedback for the output layer. Right. It also gives teams a rollout loop: start with defaults, inspect missed opportunities, add custom filters where the pain is real, and verify the filters. Exactly. Much better than installing a proxy and hoping it magically helps everything. The useful lifecycle is measure, filter, recover, improve. After the adoption loop, policy enters the architecture.
23:13 There is also a governance layer hiding under the workflow. Before rollout advice there is one policy distinction teams should keep separate. Quick privacy separation, because teams will ask. Local tracking is one thing. External telemetry is another thing. The official RTK telemetry docs say anonymous aggregate telemetry is disabled by default and requires explicit consent. They also document opt-out and forget commands. Even with that, teams should read the policy before rolling it into work repos.
23:44 The engineering distinction is, Simple. Local usage tracking helps you understand savings and missed opportunities. External telemetry should be an explicit policy decision. Yes. A tool that sits in every agent's shell path becomes part of your development environment, not just a little CLI utility. Boring governance, but real architecture. Right. Especially when the thing being shaped is command output from private repos and infrastructure tools. After the telemetry point, one more design contract matters: RTK is lossy by design. And that is not automatically bad.
24:19 Every useful abstraction is lossy. A test summary is lossy. A stack trace is lossy. A Git status is lossy compared with the whole working tree. Right. The question is whether the loss matches the job. Exactly. For example, if the job is "fix this failing test," hiding passing-test noise is healthy. If the job is "audit every emitted warning," hiding warnings may be wrong. Smaller output is a hypothesis. Production use is the test. Right. That means token savings are not the final metric. Does the agent fix bugs faster?
24:50 Does it ask for raw output at the right moments? Does it miss fewer important details? My rule is simple: when a human engineer needs the full output, the agent needs the full output too. A summary is not enough. The rollout should be boring on purpose. So how would you actually roll this out? Start narrow. Git status, Git log, push output, passing-test noise, repeated logs, dependency lists, and maybe noisy package-manager output. Keep tee recovery on. Exclude commands where exact stdout is the artifact.
25:24 Watch the command logs. For example, in Claude Code, verify the hook is installed and watch whether Bash commands are being rewritten. Right. In Codex, assume the instruction is advisory until you see consistent RTK usage. Right. For internal tools, wait until the raw output annoys you in a predictable way. Then write a TOML filter and verify it. Right. And measure task outcomes. Let's say the agent fixes a failing test. Did it complete the edit with fewer detours? Did it stop rereading giant logs?
25:55 Did it recover raw output when needed? If all you have is a savings chart, you do not yet know whether the filter is good. Let's turn that into a rebuild plan. Start with the spine. If I had to rebuild the core idea, I would not start with every command category. First, a CLI router that recognizes a few high-value commands. Second, a runner that preserves exit codes and supports passthrough. Then you need the live-output layer, right? Exactly. 3rd, a stream layer that captures standard output and standard error and can filter line by line. 4th, deterministic filters for one test runner and one Git workflow.
26:31 Right. And the safety rails are not optional. 5th, raw-output tee recovery. 6th, a rewrite command with clear allow, deny, ask, and passthrough outcomes. 7th, local tracking so you know what actually happened. Then custom filters and discovery. That order is useful because it shows what RTK really is: not a pile of abbreviations, but a command-output runtime for agents. The rebuild plan points back to the main design lesson. The durable lesson is bigger than RTK. Coding agents need better input boundaries.
27:06 Every shell command is a little API call into your repo, test suite, build system, or cloud account. RTK is one public implementation of that boundary: a Rust proxy, command-specific filters, hook rewrite logic, raw-output recovery, TOML custom filters, and local feedback loops. Claude Code gets the stronger integration story because hooks can rewrite Bash calls before execution. Codex gets a useful but weaker story through `AGENTS.md` guidance. The engineering judgment is the same in both. Use filtered output where the signal is predictable.
27:38 Keep raw recovery nearby. Do not confuse smaller with more correct. Source links for the RTK repo, official guide, and inspected code paths are in the description. Maybe the future is every developer tool having an agent-output mode next to verbose and JSON. Until then, systems like RTK are how terminal output becomes a designed interface for coding agents. Thanks for listening to Learning Podcasts.