Managed Agents
Managed agents are agent runtimes hosted by a vendor: the provider runs the agent loop, the tool-execution sandbox and the session state on its own infrastructure, and the developer supplies configuration (model, system prompt, tools) and sends events. They are the hosted alternative to a self-built agent harness, where the team writes and operates the loop, the sandbox and the state store itself. The trade is the usual one for managed services: less to build and operate, in exchange for less control over how the loop behaves and where the data sits.
The Brain, Hands and Session Split
The clearest public description of the architecture is Anthropic's engineering post "Scaling Managed Agents: Decoupling the Brain from the Hands" (April 8, 2026). It separates an agent into three parts. The session is an append-only log of everything that happened, stored outside the model's context window. The brain is a stateless harness that calls the model and routes tool calls. The hands are sandboxes where code runs, reached through a tool-shaped interface the post writes as execute(name, input) → string.
Because the harness keeps no state of its own, a crashed one can be replaced: a new instance is started with wake(sessionId) and reads the event log back. Because the sandbox is only a tool, a lost container is a tool error rather than a lost agent. Anthropic reports that moving the harness out of the container cut p50 time-to-first-token by roughly 60% and p95 by over 90%, since inference no longer waits for a container to be provisioned (vendor-reported, for its own service). Credentials stay out of the sandbox: repository tokens are used during sandbox initialization, and tokens for external tools sit in a vault behind a proxy.
What Vendors Offer as of October 2026
Anthropic. Claude Managed Agents is in beta. Its documentation defines four concepts: an agent (model, prompt, tools, MCP servers, skills), an environment (an Anthropic-managed cloud sandbox or a self-hosted one), a session, and events. Features added on top of the base runtime include outcomes, where the developer writes a rubric and a separate grader scores the work in its own context window and sends feedback for another iteration (three iterations by default, twenty at most); memory stores, collections of text files mounted into the sandbox as a directory, with an immutable version written on every change; and multiagent orchestration, in which a coordinator delegates to other agents that share one sandbox and filesystem but each keep a context-isolated thread. A scheduled memory-consolidation process called dreaming is in a more limited research preview. Anthropic says outcomes improved task success by up to 10 points in its own testing, a vendor-reported figure.
OpenAI. OpenAI splits the same pieces differently. The Responses API offers a hosted shell tool that runs commands in OpenAI-managed containers, which have no outbound network access unless an allowlist is configured. The OpenAI Agents SDK keeps the harness in the developer's own process and adds sandbox agents: a manifest describes the workspace, and a sandbox client selects where it runs — local, Docker, or a hosted provider such as E2B, Modal, Daytona, Cloudflare or Vercel. That is closer to a self-hosted harness with rented hands than to a fully managed agent.
Google. Google's documentation describes an Agent Runtime to "deploy, operate, and scale agentic applications", together with Sessions, Memory Bank and a managed Code Execution sandbox, presented as of October 2026 under the name Gemini Enterprise Agent Platform. The runtime accepts agents built with ADK, LangGraph, LangChain, LlamaIndex and other frameworks, so the developer still writes the loop and Google hosts it.
The three are not interchangeable: Anthropic hosts the loop itself, Google hosts a loop the customer wrote, and OpenAI hosts tools and containers.
Trade-offs Against a Self-Hosted Harness
| Dimension | Managed agents | Self-hosted harness |
|---|---|---|
| Control | Loop, compaction and caching are the vendor's and can change between releases | Every prompt, tool and context decision is owned and versioned by the team |
| Data | Session history and outputs are stored server-side by the vendor | State stays in the operator's datastore |
| Cost | Little engineering; pricing set by the vendor | Engineering and operations cost; freedom to mix models |
| Lock-in | Session format, memory and grader are vendor-specific | Portable across models, at the price of maintaining the harness |
The data row is concrete. Anthropic's documentation states that because Managed Agents stores conversation history and sandbox state server-side, it is not currently eligible for Zero Data Retention or HIPAA Business Associate Agreement coverage. Self-hosted sandboxes narrow the gap without closing it: tool execution, files and network traffic stay on the customer's infrastructure, but orchestration stays with Anthropic and tool inputs and outputs still flow to its control plane so the model can read them.
The control row matters more than it first appears. Anthropic's April 23, 2026 postmortem on Claude Code traced a period of quality complaints to three harness-level changes rather than to the model, which shows that a loop someone else maintains can change an agent's behavior without any change on the customer's side. Teams that choose a managed runtime still need their own agent evals to notice. A fuller side-by-side is in managed agents vs self-hosted agent harnesses.
Why the Design Converged
A stateless loop, an external log and disposable execution environments are ordinary distributed-systems practice applied to agents; the Anthropic post frames it with the pets-versus-cattle analogy. The design puts a boundary where review is possible: work happens somewhere disposable, and what leaves that place can be inspected. The same principle shows up outside coding: LightCMS, the CMS serving this site, has agents edit in a copy-on-write fork that a human reviews and merges.
Further Reading
- Scaling Managed Agents: Decoupling the Brain from the Hands — Anthropic Engineering, April 2026
- Claude Managed Agents overview — Claude Platform docs, accessed October 2026
- Define outcomes — Claude Platform docs, accessed October 2026
- Using agent memory — Claude Platform docs, accessed October 2026
- Self-hosted sandboxes — Claude Platform docs, accessed October 2026
- New in Claude Managed Agents: dreaming, outcomes, and multiagent orchestration — Claude blog, May 2026
- Sandbox agents — OpenAI Agents SDK docs, accessed October 2026
- Shell tool guide — OpenAI API docs, accessed October 2026
- Agent Platform overview — Google Cloud docs, accessed October 2026
- An update on recent Claude Code quality reports — Anthropic Engineering, April 2026