Agent Sandbox
An agent sandbox is an isolated, disposable execution environment — typically a container or a microVM — in which an AI agent runs code, shell commands and file operations, so that whatever the agent does is confined to a space that can be thrown away. In current agent architectures the sandbox is reached as a tool: the harness sends a command, the sandbox returns a string, and nothing the agent executes runs in the same process, or with the same privileges, as the system that controls it.
Why Agents Need One
An agent that can run a shell will run code nobody has reviewed: code the model just wrote, dependencies it just installed, and instructions it read from a web page or a file. The last case is the dangerous one. Prompt injection can turn untrusted content into commands, and a model cannot be relied on to refuse every time. A sandbox changes the question from "will the agent ever do something harmful?" to "what can it reach if it does?".
Anthropic's write-up of sandboxing in Claude Code (October 20, 2025) states the requirement plainly: effective sandboxing needs both filesystem isolation and network isolation. Without network isolation a compromised agent can send out files such as SSH keys; without filesystem isolation it can escape the sandbox and then reach the network. The same post reports that sandboxing reduced permission prompts by 84% in Anthropic's internal usage, which points to the second benefit: a bounded environment lets an agent work with fewer interruptions, because fewer individual actions need approval.
Isolation Technologies
| Approach | How it isolates | Documented trade-off |
|---|---|---|
| OS primitives (Linux bubblewrap, macOS Seatbelt) | Restricts a local process's filesystem and network access | Shares the host kernel; used by Claude Code for local runs |
| Containers (Docker, OCI) | Namespaces and cgroups on a shared kernel | Fast and familiar; the kernel is the shared attack surface |
| gVisor | An application kernel written in Go that runs in userspace and intercepts the workload's system calls | Its documentation cites reduced application compatibility and higher per-system-call overhead |
| Firecracker microVMs | KVM-based virtual machines, each with its own guest kernel | Project site states boot in under 125 ms and under 5 MiB of memory overhead per microVM |
Firecracker was built at Amazon Web Services for services such as AWS Lambda and is open source under Apache 2.0. gVisor, developed at Google, describes itself as neither a system-call filter nor a virtual machine in the everyday sense. Hosted sandbox products are mostly packaging around these primitives. E2B's open-source infrastructure repository describes one Firecracker microVM per sandbox, resumed from a snapshot. Modal's documentation says its Sandboxes use gVisor by default, with a virtual-machine option, a default lifetime of 5 minutes and a maximum of 24 hours. Daytona describes sandboxes with a dedicated kernel, filesystem and network stack and advertises start-up in under 90 ms; that number is vendor-reported, as are all the start-up figures here. No independent benchmark comparing these services was found among the sources opened for this page.
The Sandbox Holds No Production Secrets
Isolation limits what code can touch; it does nothing about secrets that are placed inside the boundary. The durable rule is that a sandbox should never contain a credential worth stealing. Vendors implement this in three similar ways.
Anthropic's managed-agents architecture post (April 2026) keeps OAuth tokens in a vault outside the sandbox; tools are called through a proxy that fetches the credential, and repository tokens are used only while the sandbox is initialized. Its vault documentation describes environment-variable credentials stored in the sandbox as an opaque placeholder that is swapped for the real secret at network egress, and only for allowed hosts, so that "the agent never sees the secret value". OpenAI's hosted shell documentation describes the same idea under the name domain secrets: the model sees a placeholder name, and a sidecar applies the real value only for approved destinations. Its hosted containers have no outbound network access by default.
The limits are documented too. Anthropic notes that substitution is outbound only: if a client exchanges the secret for a session token, that token arrives in the sandbox unredacted. A key scoped more broadly than the task needs also widens the damage a misbehaving agent can do, whether or not it ever reads the key. Credential isolation reduces what can be stolen; it does not make an over-privileged key safe.
Disposable by Design
Treating the sandbox as a tool also makes it replaceable. In the brain-hands-session design used by managed agents, the conversation state lives in an external log, so a dead container is an error the harness reports to the model rather than a lost task. Snapshots serve the same goal from the other side: the OpenAI Agents SDK describes a manifest for a fresh workspace and snapshot or session state for reconnecting to earlier work, with interchangeable sandbox clients for local, Docker and hosted providers.
The Sandbox as a Review Boundary
A sandbox is also a place where work can wait for inspection. Coding agents typically produce a branch or a diff that a person or a test suite checks before merging, which makes the sandbox one half of a human-in-the-loop control. The same principle shows up outside coding: LightCMS, the CMS serving this site, has agents edit in a copy-on-write fork of the content that a human reviews and merges before anything goes live. What a sandbox cannot do is judge whether the output is correct; that remains the job of agent evals and review.
Further Reading
- Making Claude Code more secure and autonomous with sandboxing — Anthropic Engineering, October 2025
- Scaling Managed Agents: Decoupling the Brain from the Hands — Anthropic Engineering, April 2026
- Authenticate with vaults — Claude Platform docs, accessed October 2026
- Firecracker: secure and fast microVMs for serverless computing — Firecracker project, accessed October 2026
- What is gVisor? — gVisor project, accessed October 2026
- E2B infrastructure repository — E2B on GitHub, accessed October 2026
- Sandboxes guide — Modal docs, accessed October 2026
- Daytona documentation — Daytona, accessed October 2026
- Shell tool guide — OpenAI API docs, accessed October 2026
- Sandbox clients — OpenAI Agents SDK docs, accessed October 2026