Claude Managed Agents Are Not What They Look Like

Outline

Transcript

0:00 So you already know exactly what it feels like to use Claude Code or, you know, local coding tools like Codex in your terminal. It's this incredibly tight, specific experience, right? Oh, yeah, absolutely. It's super fast. Right. You ask for a change. The agent reads the files, runs some commands, edits the code, and then it just checks the result and keeps going. It loops until it actually has something useful to show you. Yeah, the developer experience there is just so iterative. You're right there with it.

0:27 Exactly. But imagine taking that exact experience, pulling it out of your local terminal and moving it into a backend product. Picture just adding a button to a user interface that simply says, I don't know, investigate this issue or prepare a fix. Right. And the second you try to actually wire up that button, you hit a massive architectural wall. Yeah, because your backend just can't treat that request like a normal chat completion, can it? No, not at all. A synchronous request is, well, it's entirely the wrong shape for the job.

0:58 Because it might time out or, you know, the load balancer just drops the connection entirely. Exactly. Or it just provides absolutely no progress back to the user while it's thinking. And the work triggered by that button, I mean, it could take minutes. It needs to touch files. It requires external connectors. It has to handle credentials safely, stream progress updates, and ultimately it has to produce a reviewable artifact. Which means, you know, we really need to fundamentally shift how we look at this technology from an infrastructure perspective.

1:28 Yeah. So to figure out how to bridge this gap, we're doing a deep dive into Anthropics' new architectural documentation on managed agents. And we've got some field notes from backend engineers who actually tried to build this from scratch. Right. The painful way. Exactly. Our mission is to figure out why your standard API requests are failing your AI features and how to redesign your backend to actually support multi-step workflows. And really, it all boils down to one thesis. Anthropics' managed agents are best understood as a managed worker runtime for multi-step tool-using work, not as a smarter chat endpoint.

2:03 Yeah. And that is such a critical distinction for any backend engineer to internalize. I mean, the mental model here should actually feel really familiar to you. Right. Because you already know the difference between handling a request in line versus, say, handing work off to a background worker. Exactly. Some jobs are small enough for request response. Like if you just need to validate a field or classify a short message or extract a single value from an invoice. For those, a direct model call is perfectly fine.

2:34 Yeah, it's great. But other jobs are just not shaped like that. They take minutes and they branch based on what they actually find in the system. Right. Okay. Let's unpack this because I know exactly what a backend engineer is thinking right now listening to this. Oh, I bet I know too. They're thinking, why can't I just write a while loop? Yep. The famous while loop. Right. You call the model, you parse the JSON tool request, you execute the tool, feed the result back into the context window, and just loop until the model says it's done.

3:00 I mean, it sounds like, what, 20 lines of code? It does. It is incredibly easy to describe. But it is an absolute operational nightmare to actually run in production. Really? Why is it so bad? Well, if you try to build that while loop yourself, you are taking on this massive complex burden. Let's just look at the outer machinery you have to build around that simple little loop. Okay. You need a sandbox to run the code safely. You need an isolated file system so your concurrent requests don't just overwrite each other.

3:30 You need structured logs, a really robust way to handle transient network errors, and crucially, a mechanism for cancellation. Wait, let's stop on cancellation and state management for a second. Let's say I do write that while loop and my user clicks investigate this issue. Okay. Three minutes in, they just close their browser window. Or they click cancel. Or honestly, my own server that's deploying the loop just restarts. What happens to the partial files? Right. So in a naive while loop, that process is now orphaned. Maybe it just wrote half a file, and the next time a job runs, the whole repository is corrupted.

4:02 Oh, wow. That's a disaster. Or worse, the loop gets caught in a trap. I mean, I've seen field notes from a team that actually built their own loop. How did that go? Well, it worked beautifully on a Tuesday afternoon for a local demo. Mm-hmm. But on a Friday, a transient network blip caused the agent to silently retry a file -write operation literally hundreds of times. Wait, hundreds of times? Yeah. It cost them a massive amount in token costs in just 20 minutes, and it completely corrupted the state.

4:33 So, okay, the outer machinery isn't just nice to have. It is the product. It exists. A toy agent demo works perfectly fine in a notebook with, you know, a few hard-coded tool calls, but a production agent needs actual boundaries. What network can it reach? What files can it read? Right. These are back-end questions long before they're AI questions. And that is exactly where managed agents takes over the harness. Instead of you building the while loop, the sandbox, the retries, and the event streaming from scratch, you hand that operational burden over.

5:02 Okay. That makes sense. But to use it effectively, you have to adopt a strict back-end worker job analogy. Yeah. There are basically four structural primitives you need to learn. All right. So if I'm not writing the loop, I assume I need to start by defining the worker's brain, right? Like, what is the agent actually supposed to do, and what is it allowed to use? Yep. And that brings us to the first primitive, the agent. Think of the agent as the worker definition. Okay. The definition. Right. It's a reusable configuration.

5:30 It defines exactly which model to use, what the system instructions are, the narrow typed tools it has access to, and any MCP servers. And just to clarify for everyone, MCP servers are really simply just external tool or resource connectors. Exactly. And you don't rewrite this agent definition every single time a user clicks a button. You version it as something like a support investigator or a release note drafter. Right. But a worker definition just floating in a void can't really do anything.

5:59 We just established that giving an agent access to your entire back-end is a terrible idea. So where does this defined worker actually run? Inside the second primitive, the environment. If the agent is the worker definition, the environment is the container template. Okay. The template. Yeah. It acts as the strict isolation boundary. The environment dictates what packages are installed, what kind of file system is available, and what network access exists. Right. Because an agent is really only as useful as the world it can safely inspect.

6:31 I mean, a coding agent without access to a repository is basically just giving you advice. It can't do the actual work. Exactly. And conversely, a support agent with broad, unrestricted production database credentials is a massive security risk. Oh, absolutely. So the environment turns a vague capability into a strictly bounded runtime. Right. Okay. So we have the worker definition, the agent, and we have the container template, the environment. When a user actually clicks that button on the front-end UI, what gets created?

7:00 That creates the third primitive, the session. The session. Okay. The session is the running job. It is one running instance of a specific agent inside a specific environment. And because it's a managed job, it has strict lifecycle statuses. Like what? Well, it can be idle, running, rescheduling-like after a transient error or terminated. Rescheduling is a really key word there. That solves that orphaned while loop problem we talked about earlier, right? Exactly. The runtime manages the state if the connection drops.

7:27 But wait, if the job is running in the background for, say, 10 minutes, we hit a serious user experience problem. I mean, if you've ever stared at a spinning wheel on a front-end UI for five minutes, just praying your back-end hasn't crashed, you know exactly how frustrating silent background work is. Oh, it's the worst. Which is why the fourth primitive is the observability layer for your application. And that's events, specifically the server-sent events stream or SSE stream. Okay. So this is how your back-end talks to the running session, how the session talks back.

8:01 It streams the input, output, progress, and all those status changes. Right. And it is entirely non-negotiable for a good product experience. The event stream allows your application to show granular progress. So you can actually tell the user, reading repository, running tests, inspecting failure, applying patch, finished. Yes. It turns this terrifying black box into a visible, trustworthy process. You know, I can see why it's really tempting to look at these four primitives agent, environment, session, events, and just call this whole architecture Claude Code in the cloud.

8:34 Right. But what's fascinating here is that using that phrase is actually too narrow, and it severely limits your product design. How so? Well, Claude Code is a product built for developers. Managed agents is a platform primitive for building entirely new products. A coding agent is really just one obvious use case. Oh, I see. You can use the exact same runtime shape to build a finance document processor or an internal operations assistant. Right. Because the nature of the work fundamentally changes when you move from a local terminal to a back-end platform.

9:06 I mean, think about the supervision shift. When you're using local coding tools, the developer is the human supervisor. You're sitting right there. You see the diff in real time. Exactly. If the agent starts doing something reckless, you just interrupt the process. Your own judgment fills in the gaps of the agent's logic. But in a back-end product, your application code must encode that supervision. The developer is not there to watch the terminal output. Right. The goal of a back-end agent is not silent autonomy where it mutates production data and just hopes for the best.

9:38 The goal is bounded work that produces a reviewable artifact. Let's make that concrete. Let's use a GitHub example. Let's say we have an internal developer portal and a junior engineer clicks a button assigning a job to add idempotency key handling to an old payments endpoint. Okay. Great example. If you tried to build that as a direct model call, the model might just explain the concept of idempotency, and it might draft a generic patch based on whatever files you managed to cram into the context window.

10:06 But the real job requires deep inspection. I mean, the agent actually has to go find the relevant route handler, read the existing payment service to understand the company's specific pattern, modify the code, run linters, run the test suite. And that is exactly where a managed agent session shines. Your back-end application starts a session using your coding agent definition, running inside a repository-aware environment. So it actually has access to the code base. Right. The agent reads the files, runs the commands, and streams its progress.

10:37 And if the local tests fail, that failure isn't a dead end. What happens? The failure becomes new evidence for the agent to process. The session continues, tweaking the code until the patch actually passes. And the output is that reviewable artifact, a pull request. Ah, and then the CI pipeline gives an independent binary signal. Exactly. The pull request acts as the review boundary where a human finally steps back in. Here's where it gets really interesting, though. Because your back-end application is now the entity supervising this worker, the most dangerous part of your entire system becomes its action surface.

11:16 Yes. The tools. Tools are the only thing that makes an agent useful, and they are exactly what make it dangerous. Because an agent is only as powerful as the tools you give it. If you give an agent a bash execution tool, that is a massive action surface. File writes are a massive action surface. Web fetch is a massive action surface. We absolutely are. So my instinct as a developer might be to just write a really strict system text. I tell the agent, under no circumstances are you allowed to look at account B's logs.

11:43 Only look at account A. And that is a terrible idea. Yeah. You cannot rely on politely requesting good behavior in the system text. That is a guaranteed data breach. Because prompt injection is still a thing. Exactly. If a user from account A asks the agent a question, and they craft a clever prompt that tricks the agent into using a broad data query tool to read account B's logs, the system text will not save you. Right. Sandboxing the environment is absolutely necessary, but it is not sufficient.

12:11 Authorization must be enforced by your application's code and the tools themselves. But how do we actually build that enforcement if the agent is the one making the decisions? You utilize narrow typed tools. You keep the action surface as strictly constrained as possible. Instead of handing the agent broad credentials like, say, a raw OAuth token or general database access and just hoping it behaves, you expose a strict contextual tool. So, for example, instead of a generic execute SQL tool, you'd provide a get-user-payment-history tool.

12:39 Yes, exactly. And that tool inherently only returns records for the user ID that your backend authenticated when the session started. Exactly. The tool itself enforces the tenant boundary. And when it comes to modifying data, you enforce the principle of read access before write access. Makes sense. If the agent needs to make a change, it should use a tool that proposes a patch or drafts a configuration change instead of a tool that executes an irreversible action directly against the database. Narrow typed tools are really how these products survive contact with real users.

13:14 Okay, so with security locked down and our tools strictly scoped, we really have to talk about how this fits into the operational reality of a backend because we're integrating this into a distributed system. You always have to remember that agents are basically probabilistic control loops, but they are running inside deterministic backend systems. All right. Your backend still needs all its standard safeguards. You still need idempotency keys for your own API endpoints. You still need timeouts and you absolutely need cleanup logic.

13:45 Yeah, that cleanup logic is vital. If a user cancels a job halfway through, your backend has to handle that gracefully, which actually brings up a really common misconception about session state. Oh, let's hear it. Since a managed agent session maintains the history of the conversation in the workspace state, it kind of feels like it could be my new database. Like I can just query the session later to see what happened. No, absolutely not. Session state is just worker execution state. It is highly useful while the job runs and it's completely necessary for the agent to maintain context, but it is not your product database.

14:19 It is not your audit log and it's not your long-term memory architecture. So I shouldn't rely on the session existing forever? No. Your backend application must extract the final outputs, the structured JSON, the generated report, the pull request link, and save them to your own persistent storage. You archive or delete the sessions according to your own company's retention policy. The session is ephemeral. Got it. And what about that event stream we talked about earlier, the SSE stream? The event stream acts as your observability layer.

14:50 You pipe those events into your own logging systems. You use it to see exactly where the agent spent its time, which tools it called, what failed, and why it stopped. That is how you debug an agent in production. Okay. So we have firmly established that this is powerful infrastructure for multi -step jobs. But let's establish some boundaries here. Should I be replacing my normal workflows with managed agents? Firmly, no. Yeah. You have to measure the complexity of the work. If your task is a single, short, transformation like extracting fields from an invoice or classifying an inbound email, a direct model API call is a much better fit.

15:25 Because it's simpler. It is simpler to test, it's faster to execute, and it's way cheaper to operate. Okay. But what about complex, high-volume pipelines? Let's say I have a pipeline that processes thousands of documents a day, and it takes seven distinct steps to validate, extract, and route the data. Even there, if you have a regulated pipeline where every single step is known in advance, fixed workflows beat agents every time. Really? Yeah. If your pipeline always does exactly the same, seven things in the exact same order, keep those steps explicit in your deterministic code.

15:58 In addition to that, agents incur token costs and runtime costs. Right. Because the agent is constantly thinking about what to do next. It has to evaluate the state of the world after every single tool call. Exactly. You only reach for a managed agent runtime when the path to the solution depends entirely on what the system discovers along the way. Oh, that makes total sense. If the path is known, agency is just expensive overhead. If the path requires investigation, reading logs, running commands, and reacting to errors, then a managed agent is the right tool for the job.

16:27 So what does this all mean for the back-end engineer trying to build the next generation of features? Managed agents provide the managed harness. You get the container templates, the running sessions, the event streams, and the tools. But they do not remove your product design responsibilities, and they certainly do not remove your back-end responsibilities. Exactly. Think about what your application code still strictly owns. Your code owns authentication. It owns authorization and tenant isolation.

16:55 It owns the status UI that the user interacts with. It owns cancellation logic and the persistence of the final result. Right. Most importantly, the back-end code still owns what work users may ask for. Yeah. Exposing an agent to your users is not just dropping in an AI feature. You are designing robust infrastructure. You are taking a probabilistic loop and putting it inside a deterministic, safe environment. You are turning bounded work into something users can genuinely trust.