Managed Agents vs Self-Hosted Agent Harnesses

Comparison

Managed agents vs self-hosted agent harnesses is the build-or-buy question for the software around a model. With managed agents, a vendor runs the agent loop, the session store and usually the execution sandbox, and the developer configures an agent and exchanges events with it over an API. With a self-hosted agent harness, the developer runs that loop in their own process, using a library such as the Claude Agent SDK, the OpenAI Agents SDK, Google ADK, LangGraph or Microsoft Agent Framework, and owns state, recovery and the agent sandbox.

As of October 2026 the managed option exists at all three large model vendors, in different forms. Anthropic's Claude Managed Agents and OpenAI's Agents API are both in beta and both host the harness itself. Google's Agent Runtime is a managed service for deploying agents written with a framework, so the developer still supplies the harness code. The line between the two columns is also blurring: Anthropic and OpenAI both let the hosted harness drive a sandbox that runs on the customer's infrastructure.

Feature Comparison

DimensionManaged AgentsSelf-Hosted Harness
Who runs the agent loopThe vendor. Anthropic: "pre-built, configurable agent harness that runs in managed infrastructure"; OpenAI: "OpenAI runs the agent harness"The developer, in their own process or cluster
Examples (Oct 2026)Claude Managed Agents (beta); OpenAI Agents API (beta); Google Agent Runtime (managed hosting for framework-built agents)Claude Agent SDK, OpenAI Agents SDK (MIT), Google ADK (Apache 2.0), LangGraph (MIT), Microsoft Agent Framework (MIT)
Session statePersisted server-side by the vendor; event history retrievable through the APIWhatever store the developer chooses: files, SQLite, Postgres, Cosmos DB and so on
Execution sandboxVendor-hosted by default; self-hosted sandbox option at Anthropic and OpenAIDeveloper-provisioned: local process, Docker, or a sandbox provider
Model choiceThe vendor's own modelsDepends on the library; several support multiple providers
Harness customisationConfiguration: model, system prompt, tools, MCP servers, skillsFull: the loop, context management, permissions, hooks and tool routing are code
Built-in extrasCompaction, prompt caching, multi-agent delegation; Anthropic adds rubric-graded outcomes and memoryVaries by library; assembled and tuned by the developer
PricingAnthropic: model tokens plus $0.08 per session-hour while running. OpenAI: model API rates, standard tool rates and container rates. Google Agent Runtime: rates not reproduced hereModel tokens plus the developer's own compute, storage and operations
Data-handling limitsClaude Managed Agents: not eligible for Zero Data Retention or HIPAA BAA. OpenAI Agents API: US data residency only, no ZDRDetermined by the developer's infrastructure and the model API terms selected
Operational burdenLow: no loop, queue, recovery or sandbox fleet to operateHigh: provisioning, crash recovery, upgrades, security patching and observability
MaturityBeta at Anthropic and OpenAI; behaviours "may be refined between releases" (Anthropic)Ranges from pre-1.0 libraries to GA frameworks

Detailed Analysis

What the vendor actually manages

Anthropic's engineering account of its service (April 2026) describes virtualising three components: a session, "the append-only log of everything that happened"; a harness, the loop that calls the model and routes tool calls; and a sandbox, where code runs and files are edited. Keeping them separate means a failed container no longer loses the session, and inference does not wait for a container to start. Anthropic reports that p50 time-to-first-token fell by roughly 60% and p95 by over 90% after the change. Those are vendor-reported figures for its own system.

OpenAI's Agents API documentation uses nearly the same vocabulary of agent, environment, session and events. It says OpenAI "manages sessions, orchestration, context compaction, and recovery" while the application supplies tools and chooses an execution environment: OpenAI-hosted, self-hosted, or none. Google's Agent Runtime is a different kind of product. It is a managed, auto-scaling environment for agents built with ADK, LangChain, LangGraph, LlamaIndex, AG2 or custom code, alongside managed Sessions, Memory Bank, evaluation and code-execution services. The developer still writes and versions the agent loop.

What a self-hosted harness gives back

Running the harness yourself returns control of the parts that most affect agent quality. The Claude Agent SDK exposes the tools, agent loop and context management that power Claude Code as a Python and TypeScript library, with hooks at lifecycle points, a permission system, subagents, sessions that can be resumed or forked, and MCP connections. Open-source frameworks go further by making the loop itself editable and, in several cases, model-independent.

That control matters because the harness, not only the model, shapes results, which is the premise of harness engineering. A managed service makes those decisions for its customers and may revise them between releases, as Anthropic's beta notice says explicitly. A team that needs a frozen, reproducible agent, or a behaviour the vendor does not offer, has to own the harness.

Security and data boundaries

Managed services have a structural security advantage in one respect: credentials can be kept out of the sandbox. Anthropic describes cloning repositories with an access token during sandbox initialisation, and holding OAuth tokens for MCP tools in a vault reached through a proxy, so that code the model writes never sees them. A self-hosted team can build the same pattern but has to build it.

The disadvantage is that state lives with the vendor. Anthropic states that because sessions store conversation history, sandbox state and outputs server-side, Managed Agents is not currently eligible for Zero Data Retention or HIPAA Business Associate Agreement coverage. OpenAI states that its Agents API supports data residency only in the United States and does not support ZDR. Self-hosted sandboxes move tool execution, files and network egress into the customer's environment, but Anthropic's documentation is clear that tool inputs and outputs still flow to its control plane so the model can see results. For workloads where that is unacceptable, only a fully self-hosted harness calling a model API under suitable terms will do.

Cost and features

Managed pricing is simple to state. Claude Managed Agents bills model tokens at standard rates plus $0.08 per session-hour, metered to the millisecond only while a session is running; idle time waiting for a message or a tool confirmation is not charged, and the Batch API discount does not apply. OpenAI bills model usage, tools and hosted containers at their standard rates. In both cases tokens dominate: Anthropic's own worked example of a one-hour session totals $0.705, of which $0.08 is runtime. The saving from self-hosting is therefore mostly not in the runtime fee. It lies in control over caching, model routing and context size, set against engineering and on-call time.

Managed platforms are also where vendors ship features first. Anthropic's May 2026 update added multi-agent orchestration, in which a lead agent delegates to specialists with their own model, prompt and tools, and outcomes, in which a separate grader scores output against a developer-written rubric in its own context window (both public beta), plus a scheduled memory-curation process called dreaming (research preview). Anthropic reports that outcomes improved task success by up to 10 points over a standard prompting loop; this is a vendor-reported result. Each of these can be reproduced in a self-hosted harness, at the cost of building and evaluating it.

Best For

Shipping a first production agent quickly

Managed Agents

No loop, queue, session store or sandbox fleet to build. The vendor handles compaction, caching and recovery.

Zero-data-retention or HIPAA-covered workloads

Self-Hosted Harness

Anthropic and OpenAI both state that their managed agent services do not support ZDR, and Anthropic excludes HIPAA BAA coverage.

Multi-vendor or open-weight models

Self-Hosted Harness

Managed agent services run the vendor's own models. Frameworks such as LangGraph, Google ADK and Microsoft Agent Framework support several providers.

Long-running asynchronous jobs

Managed Agents

Durable server-side sessions, webhooks and scheduled runs are provided, and a lost container does not lose the session.

Custom context management or tool routing

Self-Hosted Harness

Hooks, permissions and the loop itself are code. Managed services expose configuration, not the loop.

Private data with limited operations staff

Hybrid

A managed harness with a self-hosted sandbox keeps files and network egress in-house, though tool inputs and outputs still pass through the vendor.

Reproducible, pinned behaviour for audits or evals

Self-Hosted Harness

A beta managed harness can change between releases. A pinned library version and your own prompts do not.

Small team, variable load

Managed Agents

Runtime is billed only while sessions run, and there is no idle infrastructure to pay for or patch.

The Bottom Line

Managed agents trade control for operations. The vendor runs the loop, keeps the session log and supplies an isolated sandbox with credentials held outside it, and in exchange the customer accepts that vendor's models, its harness decisions, its data-retention posture and, for now, beta status. For teams whose agents fit those constraints, the runtime fee is small next to token cost and the saved engineering is real.

A self-hosted harness is justified when one of the constraints binds: regulated data that cannot sit in a vendor's session store, a requirement for several model providers or open weights, the need to pin behaviour for evaluation, or a harness design the vendor does not offer. It should be chosen knowing that durable sessions, sandbox isolation and credential handling are the hard parts and will have to be built.

The architecture is the same either way: a stateless loop, a durable event log outside the context window, and disposable execution environments. Designing around that shape keeps the decision reversible, and the self-hosted sandbox options at Anthropic and OpenAI mean it is no longer strictly either-or.