Reactor: The AI Runtime That Skips Unchanged Work

Transcript

0:00 Picture an agent that wakes up every morning on a cron schedule. A digest, an inbox triage, a risk dashboard. It re-runs the model over the whole world. Full price, every single tick, even on a dead-quiet day where nothing changed. So the clock is driving the bill, not the world. That is the whole problem. This is Systems in the Open, where we rebuild real systems from primary sources. And today, frankly, we are going to kill the AI cron job. We walk through Reactor, a runtime that fixes this by re-running a node only when its inputs actually moved.

0:37 And our source is the public `openprose` slash `prose` repository at a pinned commit, plus the project's language spec and harness spec. Links are in the description. And the ground rule here is, we read the code, we did not interview the maintainers, which means every claim points straight at a file you can open. So let's sit with the pain for a second, because it is one most of us are already paying for. Picture a scheduled summary agent that fires at 9, at 10, at 11. And is each fire just a full model pass over everything?

1:09 Every time. It rescans the inbox, re-summarizes the same threads, re-scores the same risks. Three identical reports, three identical bills, even when literally nothing moved between them. Yeah, I have shipped exactly that. A nightly summary job that cost the same on a holiday weekend, when 0 new data landed, as it did mid-quarter. It took me an embarrassingly long time to ask why. Right, so the trap is sitting in plain sight: you are paying for the schedule firing. Now you can see why the cost should track change, not the clock.

1:45 Now here is the thing, and I think you will recognize it instantly. Every other layer of software already solved this. You declare the state you want, a reconciler makes reality match, and work only happens on drift. Yeah. SQL, Terraform, Kubernetes, React, all the same shape. Totally. Think about React for a second. It does not repaint the whole page every frame just because the clock ticked. It diffs and touches what changed. Right, this is the oldest good idea in software. And somehow, the moment we all started building LLM agents, the industry got collective amnesia and reverted to naive, expensive polling.

2:22 One sentence: it is a thermostat for model calls. That is the core idea, and we can drop React now, because the rest stands on its own. So the claim the project makes, and it is a checkable one, is inference cost that scales with surprise, not with wall-clock time. Okay, that is a strong line. How is it checkable and not just a slogan? Because the runtime keeps a cost rollup you can actually read. It splits tokens into fresh and reused, and it tags every node with a disposition: rendered, skipped, or failed.

2:54 Got it. So fresh tokens are what a real surprise cost, and reused tokens are what you saved by skipping. You read the percentage straight off. Exactly. The bill itemizes what changed, which means you verify the claim instead of believing it. The slogan ships with its own proof. Before we go further, we should name the two layers, because the project has two names and people conflate them. So OpenProse is the language. You write contracts in plain Markdown files, the `.prose.md` files. Reactor is the runtime that reconciles them.

3:26 Yeah, and those contracts are portable, by the way. Imagine pointing a different runtime at the same Markdown files later. Reactor is the recommended fast path, not a cage the language is locked into. So why follow the runtime today instead of the language? Because the runtime is where the memoization, the dependency graph, and the receipts live. So that is the layer that earns our full attention today. From that split, let's open an actual contract, because it demystifies fast. A real one is a Markdown file with frontmatter that says `kind: responsibility`. That is it.

4:02 So not a config language, not YAML soup. Plain Markdown. Plain Markdown. And the contract is the only authored source of intent. That is the first tenet of the whole project: you declare what should be true, in one file. For example, in the surprise-cost demo they ship, the digest contract just declares what the digest keeps true. No imperative steps, no loop. Yeah. You describe the destination, and the runtime figures out the trip. That is the point: intent lives in exactly one place. Now inside that contract, a node declares three things, and once you see them the rest of the system clicks.

4:42 First is Maintains: the truth schema this node keeps current. And Maintains is where it marks which fields actually count as a change, plus the postconditions that have to hold. Right. Second is Requires: the upstream facets it subscribes to. 3rd is Continuity: its wake source, what actually triggers it. Yeah. So a contract says what stays true, what it depends on, and when it wakes. Three declarations and you are done. Exactly. That is the whole shape of a node, which means everything downstream just follows from those three.

5:15 From those sections, there are three kinds of node, and the vocabulary is deliberately small. A gateway is ingress. It turns an external trigger into truth, an entry point into the graph. Then a responsibility is a standing goal. For example, keep the daily digest current. A mounted node in the graph that holds something true over time. Right. And a function is a pure, stateless helper. It declares Parameters and Returns, and it carries no stored truth and no wake source. Yeah. So ingress, standing work, and pure helpers.

5:48 That covers the space without a sprawling taxonomy. A small vocabulary on purpose. Now you can hold the entire node taxonomy in your head, which is the point. Now here is the first half of the question every engineer is already forming, which is, isn't this just caching? Are we just doing Redis caching and giving it a fancy academic name? Yeah, I was about to say it. So the first part of the answer is material versus immaterial fields. Not every field counts as a change. Only material ones move the fingerprint.

6:20 And crucially, the contract author decides what material means. For example, a last-seen timestamp might not count, while a changed dollar amount does. Sure, and plain caching cannot do that. A byte cache keys on the literal bytes; this keys on author-declared meaning. That is the shift the whole runtime is built on. After material fields, we get to the heart of the whole thing. There is exactly one rule, and everything hangs off it. Let's hear it. A node renders if and only if its memo key moved.

6:52 And the memo key is a pair: the contract fingerprint, and the input fingerprints. Nothing else. So no clock at all. The current time never enters the key. Not anywhere. The code comment in the memo module literally says the key is that pair and nothing else. If neither fingerprint changed, the node does not run. And that single rule is the entire thesis. Take a node on any wake: it comes down to one decision, do the fingerprints match the last run. If they do, it skips. Now the second half of the caching question, and honestly this is the part I find most interesting.

7:31 The world-model is written by a model, so it is fuzzy. How does fuzzy text become a stable fingerprint? So that is the canonicalizer's job. It does a canonical serialization over the material projection of that truth, then fingerprints the result. Meaning two model outputs that say the same thing different ways collapse to one fingerprint. Yeah. A re-run with identical meaning skips, even when the raw bytes differ. The real engineering is deciding what counts as the same. So the hard problem was never the cache lookup.

8:05 It was defining sameness over messy model output. And that is the click. Sameness is the whole game here. Let's trace the meter across three wakes on their surprise-cost example. Start cold: the world wakes for the very first time. So nothing is memoized yet. There is no prior truth to compare against. Right. Both nodes fire: the gateway renders, the digest renders, and the fresh meter ticks up by two. Two real model calls. Sure. Every fingerprint is new on a cold start, so every node counts as a surprise by definition.

8:39 And that is the cold-start tax. You pay it exactly once, and now you have a baseline to compare against. From that cold start, now the interesting wake. An identical re-trigger arrives. Same data, nothing actually changed. So walk me through what the gateway does. It recomputes its fingerprint, sees it did not move, and memo-skips. It pushes no facet downstream. So nothing wakes. Wait, so the digest never even runs? Never even woken. Say you fire this thing every 5 minutes on a quiet afternoon: fresh spend on each of those wakes is 0.

9:13 The whole pass is a fingerprint comparison and a shrug. Oh, and that is the payoff of the entire idea, right there. A scheduled wake on a stable world costs you nothing. After that quiet wake, the 3rd case, where you actually force a render. To do it you move the memo key on purpose: you bump the gateway's contract fingerprint over the same ledger. So you are simulating a real change by moving the thing the key watches. Yeah. The fingerprint moves, so the gateway has to render. That pushes the facet the digest subscribes to, so the digest re-renders. The surprise propagates exactly one hop.

9:51 One hop. Oh, wow. So it does not re-run the universe; it follows the one edge that actually changed. And change is the only thing that costs. Now you can see the marquee claim in motion: spend tracked the one real edit, nothing else. Now a question that hits you immediately. Who wires up that dependency graph? Yeah, because wiring a DAG by hand is its own nightmare. Right, so you never do. A component called Forme builds it. It matches each node's Requires facet against another node's Maintains facet and draws the subscription edge.

10:29 So structure is subscription. If I require a facet you maintain, an edge appears between us automatically. Exactly. The topology falls out of the contracts themselves. Nobody draws boxes and arrows. I love that. The graph is a consequence of what each node declares. Now you can trust the topology to never drift out of sync with reality. From that auto-wiring, one more piece worth naming: a single truth can split into independently subscribable parts, called facets. Okay. So a consumer subscribes to exactly the facet it depends on, and wakes only on that lane.

11:07 Yeah. Say a truth has three facets and you only care about one. The other two moving does not wake you. Fine-grained subscription means fine-grained wake. And the default is the atomic whole-truth facet, not a wildcard that grabs everything. Right, so you opt into a narrow lane on purpose. And that is the point: surprise only spreads down lanes someone actually subscribed to. Now step back, because there is a clean split in time here. Forme checks the graph is acyclic, then freezes the topology, the per-node canonicalizer, and the validators at compile time.

11:44 So all the intelligence is frozen ahead of time. Right. And then run time is dumb on purpose. It compares fingerprints, then it skips, renders, or propagates. That is the entire loop. Yep. Compile once, run forever. So the thinking happens up front, which lets the run loop stay fast and predictable. Yeah, and predictable is the goal. That is the shift: at run time there are 0 judgment calls left to make. After that frozen compile, let's look at what happens when a node does need to render, because the unit is smaller than people expect. So it is not a long-lived agent loop chewing on the world.

12:25 No. A render is one bounded model session that computes the next state of its truth. It reads the prior version by reference, so it does not go re-fetch everything. Right. Think of it like a function call that returns the next value: reading the old state is cheap, and the expensive model work is scoped to producing the new one. Right. One bounded unit of model work, then it is done. And that is the trick: each render is a single accountable unit, so the cost stays legible. Now this next part genuinely surprised me when I read it, so let me be careful here. When a render finishes, what checks the work?

13:06 So this is the part people get wrong. There is no judge. No second model grading the first. Wait, really? No critic model in the loop? Nope. The render self-attests its postconditions and signs its own receipt. And where a postcondition can be expressed deterministically, the harness verifies it in plain code at commit time. So the safety check is deterministic, not another stochastic model second-guessing the first. Okay, here is my worry. For example, say the postcondition is subjective, like make sure the email's tone stays polite and professional.

13:40 A deterministic check cannot measure politeness, so doesn't the whole gate fall apart there? Then it cannot live as a postcondition inside that node. The gate has to be deterministically expressible, so for a subjective check you architect a second responsibility node downstream whose whole job is to evaluate the first node's output. Exactly, and people garble the no-judge part constantly, so it is worth saying twice: the guard is plain code, never a second model. You do not let the LLM grade its own homework.

14:12 Now we can see why a failed render has to fall back to safe behavior. From that deterministic gate, what actually happens when a render cannot satisfy its postconditions? So I am guessing it does not just commit anyway and hope. It commits nothing. The prior truth stands untouched, and a failed receipt records exactly why it bailed. So under uncertainty the default is to do nothing, but loudly. The system would rather stall than write a wrong truth. Yeah. Safety outranks cost, and that is an explicit tenet, which means under load you get a clean stall, never a silently wrong answer.

14:50 Now here is the line that makes the whole thing auditable, and it is worth stating crisply. The model authors exactly two things. Which two? The compile step, meaning the topology, the canonicalizer, and the postcondition checks. And the render, meaning the new state. That is the entire surface the model touches. And everything else, the scheduling, the validating, the recording, the executing, is deterministic code the model cannot override. On paper, a stochastic model bolted inside deterministic plumbing looks like oil and water, an impedance mismatch. But that cage is exactly what lets you meter the spend.

15:29 So you can always point to where a stochastic decision entered and where it did not. For anyone running a design review, that is the unlock: a hard line between guesswork and guarantees. After that boundary, let's talk about receipts, because they are a first-class output, not a log line. Every decision the runtime makes leaves a content-addressed receipt. So what is actually in one? The fingerprints involved, the wake cause, the status, and the cost. So imagine a node that re-ran when you expected a skip.

16:02 You can open its receipt and see exactly why it ran. Sure. And content-addressed means each receipt is named by a hash of its own contents, so the trail composes. Right. Trust here is demonstrated rather than asserted. So the next on-call engineer queries the trail instead of grepping logs in a panic. From those receipts, the cost rollup is just the receipts read as an invoice. It itemizes fresh tokens, what each surprise actually cost. And the reused tokens, the spend you avoided on every skip.

16:35 Sure. Plus a surprise cause on every line, which has to echo the wake source that drove the spend. So the invoice does not just say you were charged. It says which trigger caused which spend. Yeah, so every token on the bill is attributable. That is the unlock for anyone who has to defend a model budget. After the invoice, one more property: the ledger is chain-verifiable. You can walk it and confirm it has not been tampered with. So they actually test that, not just assert it. They ship a demo called tamper-forge that attacks a real ledger, and it shows exactly where verification catches the edit.

17:13 So the trail defends itself on purpose. Take an attacker who rewrites one committed world-model: the chain walk flags the exact spot it breaks. Right, the verification is the point. So now you can show the catch instead of promising it. Now, in that same honest spirit, we should be precise about what signed means here, because it is easy to oversell. Yeah, this is the part to get exactly right. So in version one, signed means tamper-evident at the meaning layer. It is not a cryptographic byte hash.

17:48 So editing a published world-model while leaving the chain itself intact is not yet caught. Not yet. And there is no timestamp and no actor in a version-one receipt either. You cannot ask who wrote this and when. It is almost like they built the engine so purely around structural causality that they forgot to put a watch on it. And the project says all of that out loud, in its own README. So we say it too, which means you get the real picture: tamper-evident today, not tamper-proof. From those caveats, here is the part I really like.

18:23 You do not have to trust any slide we just showed. You can run the proof. So how do you run it without burning model credits? One command replays a saved run with no model key and no spend. It prints the dispositions, the surprise causes, the full token accounting, and a chain-verify. So a fully offline, keyless replay. You get the whole accounting without making a single live model call. Right, so the claim ships with its own way to check it. If you are evaluating this, that is the moment it earns trust: run the replay, read the numbers, no faith required.

18:58 Now let's be fair about when this actually pays off, because it is a bet, not magic. A quiet world wins big. Meaning most facets are stable between wakes, so most nodes memo-skip. Right. But flip it. Imagine a trading-signal agent on a price feed that moves every second, or live telemetry off a Formula One car where tire temperature, RPM, and wind speed all change every millisecond: every input churns on every wake, every node is surprised, and the bill never drops at all. Yeah, so this whole thing is a bet on stability.

19:33 If your inputs never settle, it is pure overhead. So the decision that tells you whether it pays off comes down to one question: how quiet is your world, really? After the fit question, the honest status, because the project is young and says so. The benchmarks are openly pending. So they are not claiming a speedup they have not measured. Right. They publish the harness before the numbers, rather than imply a result they have not run. I think that is more trustworthy, not less. Yeah. And the determinism is specifically in the wake-and-commit decision, not in the model output. The render itself is, Still stochastic.

20:12 Yeah. The postcondition gate is what bounds that render. So treat this as an idea and an implementation worth studying, not a drop-in to adopt blind. Now here is what I would actually take home, even if you never run Reactor itself. Several of these ideas graft straight onto your own agent stack. So give me the portable set. Content-addressed world-models. Fingerprint-gated model calls. Receipts used as composition edges. A deterministic, no-judge commit gate. And fail-closed defaults. And each of those stands alone.

20:43 Take fingerprint-gating: bolt it onto your nightly cron agent tomorrow, and the quiet-day bill drops without touching anything else. Right. So the transferable architecture is the real payoff, and the runtime is just one clean packaging of it. Which is what makes it worth studying even if you never deploy it. Yeah. The ideas outlive the implementation, which means that is the lesson worth stealing even if Reactor never ships to your stack. So to close the loop, here is the whole model in three beats.

21:14 Declare the world as it should be. Reconcile only on surprise. Leave a receipt for every decision. And the reason it is cheap to run is that it is honest about what changed. No change, no spend. The real shift is this: the cost of intelligence should not be measured in minutes, it should be measured by how much reality actually moved. Our source pack was the pinned `openprose` slash `prose` repository, plus the language and harness specs. The links are in the description. Next time on Systems in the Open, we take another real system and rebuild its shape from the sources.

21:48 Thanks for listening to Learning Podcasts.