Agent sandboxes compared: container-use, E2B, Daytona, microVMs and Docker
An agent sandbox is a disposable machine that a coding agent such as Claude Code or Codex runs inside, so an unattended run can do anything without reaching your laptop, credentials or production. The options differ in isolation: OS-level wrapping (Claude Code sandbox runtime), containers (container-use, Modal), microVMs (E2B, Docker Sandboxes), and managed or self-hosted platforms (Daytona, Coder).
This page is for the developer who wants an agent to work for an hour without approval prompts, and the tech lead who has to decide where that is allowed. You tried auto mode, then someone on the team ran claude -p --dangerously-skip-permissions on a laptop with a production .env in the repository. The question now is which sandbox makes the next run safe by construction, and how the work gets back into review.
What this sandbox catalogue gives you
Section titled “What this sandbox catalogue gives you”- A decision table and a dated comparison of seven sandboxes on isolation, startup, lifetime, network control and price.
- A worked example: Claude Code headless in E2B, from the E2B cookbook, plus Codex.
- A workflow: an unattended loop in a disposable VM whose branch comes back as a git bundle for review.
- Three copy-paste prompts, traps and failure modes.
The pattern behind all of this (a seeded, per-task stack with its own ports, data and credentials) is in ephemeral environments. The approval modes and OS sandboxes built into each agent are in permissions and sandboxing.
Which agent sandbox should you pick?
Section titled “Which agent sandbox should you pick?”Start from where the run happens and who owns the machine.
| Your situation | Pick | Why |
|---|---|---|
| One developer, several agents, on a laptop | Docker Sandboxes (sbx) or container-use | A microVM (Docker) or container (container-use) per agent, each on its own clone or branch; free for local use |
| Unattended runs of Claude Code only, no Docker allowed | Claude Code sandbox runtime (@anthropic-ai/sandbox-runtime) | Wraps the whole Claude Code process, including hooks and MCP servers, in Seatbelt or bubblewrap |
| Agent runs started by your CI, a queue or your product | E2B, Modal or Daytona | SDK-created sandboxes with per-sandbox egress control and a lifetime you set in code |
| Code and model credentials must stay on your own infrastructure | Coder | Self-hosted workspaces in Terraform; Coder Agents keeps “no API keys in workspaces” |
| A web or phone UI over sandboxed agents | CloudCLI sandbox (experimental) | Wraps Docker Sandboxes behind CloudCLI’s UI |
| Background tasks started from an issue or chat, no infrastructure | The vendors’ own cloud environments | Covered in ephemeral environments, not here |
A worktree is not on this list on purpose. It isolates files only; ports, databases, Docker names and every credential in your shell stay shared (see parallel agents).
How do the sandboxes compare on isolation, startup, lifetime and price?
Section titled “How do the sandboxes compare on isolation, startup, lifetime and price?”Facts below were checked on 2026-09-26 against each project’s README, SDK type definitions or documentation source on GitHub. The vendor websites of E2B, Modal and Daytona were unreachable from our environment on that date, so no price or startup figure from those sites appears here.
| Sandbox | Isolation model | Runs where | What drives startup | Default lifetime | Network control | Pricing |
|---|---|---|---|---|---|---|
Claude Code sandbox runtime (npm @anthropic-ai/sandbox-runtime 0.0.77, beta research preview) | OS primitives around the whole process: sandbox-exec (Seatbelt) on macOS, bubblewrap on Linux, and a Windows alpha (srt-sandbox account plus WFP egress filter) as of 0.0.77 | Your machine | Nothing to boot; it wraps the process | The process | Denies network by default; you allow domains in ~/.srt-settings.json | Free npm package |
| container-use (Dagger, Apache-2.0, “experimental”) | A fresh container per agent, each on its own git branch | Your machine, via Dagger | Image build on first use | Until you merge, apply or discard the environment | Not documented in the README | Free, open source |
Docker Sandboxes (sbx CLI) | A microVM per sandbox with its own Docker daemon, filesystem and network | Your machine, or Docker-managed cloud | First run pulls the agent image; Docker says later runs “start in seconds” | Persists after the agent exits until sbx rm | Global presets Open, Balanced or Locked Down, plus sbx policy allow network <host> | CLI and local compute free, including commercial use; cloud is pay-as-you-go |
E2B (npm e2b 2.51.0) | One Firecracker microVM per sandbox, own cgroup and network namespace | E2B Cloud, or self-hosted from e2b-dev/infra | Boots or resumes from a snapshot of a template you build once | 300 s unless you set timeoutMs; maximum 1 hour on Hobby and 24 hours on Pro (SDK docs) | network.allowOut and denyOut (hostnames, IPs, CIDRs), or allowInternetAccess: false | Tiered (Hobby, Pro); prices not verified here |
Modal Sandboxes (PyPI modal 1.5.5) | A container created on demand with modal.Sandbox.create | Modal’s cloud | Building the modal.Image you define | timeout=300 seconds by default | block_network=True, outbound_domain_allowlist, outbound_cidr_allowlist | Not verified here |
Daytona (PyPI daytona, npm @daytona/sdk 0.218.0) | Hosted sandboxes created through the SDK | Daytona’s hosted service | Creating from a snapshot | Auto-stops after 15 idle minutes; auto-delete is off unless you set it | network_block_all, domain_allow_list, network_allow_list | Not verified here |
| Coder (AGPL-3.0; Premium tier paid) | Workspaces you define in Terraform | Your infrastructure | Your Terraform template | Set by your templates | Yours to configure | Open source core; Premium is paid |
Daytona’s open-source repository says: “This repository is no longer maintained. As of June 2026, Daytona’s core development has moved to a private codebase.” The hosted service and SDKs continue (both packages released 0.218.0 on 2026-09-25). CloudCLI’s sandbox is marked experimental and needs Docker’s sbx CLI; CloudCLI’s managed cloud starts at €7 a month, separate from your model subscription.
Popularity as of 2026-09-26 (GitHub stars, read through the GitHub API on that date): Daytona 71,717 (repository frozen), Coder 16,696, E2B 13,971, CloudCLI 13,811, container-use 4,046 and modal-labs/modal-client 518 (the SDK repository only). Stars measure attention, not fitness for your threat model.
Run Claude Code headless in an E2B sandbox
Section titled “Run Claude Code headless in an E2B sandbox”This is the E2B cookbook’s own example (e2b-dev/e2b-cookbook, examples/anthropic-claude-code-in-sandbox-js), checked on 2026-09-26. It builds a template once, then runs one Claude Code prompt in a sandbox and kills it.
-
Install the SDK and the runner, and export both keys from your secret store (never paste them into the file):
Terminal window npm install e2b dotenvnpm install -D tsxexport E2B_API_KEY ANTHROPIC_API_KEY # values come from your secret manager -
Define the template. The cookbook names it
anthropic-claude-code; that name is yours to choose, and it is not a pre-built E2B template:src/template.ts import { Template } from "e2b";export const templateName = "anthropic-claude-code";export const template = Template().fromNodeImage("24").npmInstall(["@anthropic-ai/claude-code"], { g: true }); -
Build it once. The cookbook uses one vCPU and 1,024 MB of memory:
src/build.prod.ts import "dotenv/config";import { Template, defaultBuildLogger } from "e2b";import { template, templateName } from "./template";await Template.build(template, {alias: templateName,cpuCount: 1,memoryMB: 1024,onBuildLogs: defaultBuildLogger(),});Terminal window npx tsx src/build.prod.ts -
Run a prompt.
timeoutMs: 0on the command removes the per-command timeout, because a Claude Code run can take a long time:src/index.ts import "dotenv/config";import { Sandbox } from "e2b";import { templateName } from "./template";const sbx = await Sandbox.create(templateName, {envs: { ANTHROPIC_API_KEY: process.env.ANTHROPIC_API_KEY! },// Not in the cookbook: deny all egress except the model API.network: { allowOut: ["api.anthropic.com"], denyOut: ["0.0.0.0/0"] },});console.log("Sandbox created", sbx.sandboxId);const result = await sbx.commands.run(`echo 'Create a hello world index.html' | claude -p --dangerously-skip-permissions`,{ timeoutMs: 0 },);console.log(result.stdout);await sbx.kill();Terminal window npx tsx src/index.ts
You should see Sandbox created with an ID, then Claude Code’s final message describing the index.html it wrote. The file exists only inside the microVM, sbx.kill() destroys it, and nothing comes back to you. The next section turns it into a loop whose output you can review.
How do you run —dangerously-skip-permissions safely and still review the result?
Section titled “How do you run —dangerously-skip-permissions safely and still review the result?”The loop has five stages, and the sandbox is only one of them: specify → build in a disposable VM → verify with the harness, not the agent → return a branch → review and merge. The agent never holds a GitHub token and never pushes. Your harness moves code in and out as git bundles, so the sandbox’s only outbound connections are the model API and the package registry.
-
Specify. Write the task as a file with acceptance criteria the harness can check (a test command that must pass). An agent told “make it work” will declare success; an agent told “
npm test -- src/billingmust pass and no file outsidesrc/billing/may change” can be held to it. See how to write acceptance criteria an agent can be held to. -
Build in a disposable VM. Create the sandbox with egress denied except for the hosts the run needs, upload the repository as a bundle, and run the agent unattended.
-
Verify with the harness. After the agent exits, the harness runs the acceptance command itself and records the exit code. The agent’s “all tests pass” is not evidence.
-
Return a branch. The harness bundles the agent’s commits and downloads the bundle. Nothing is pushed from inside the sandbox.
-
Review and merge. You fetch the bundle into a local branch, push it, and open a pull request. CI, a review bot and a human reviewer see it like any other change.
This harness implements stages 2 to 4 for Claude Code on E2B, reusing the cookbook’s template. Run it from the root of your repository:
import { execFileSync } from "node:child_process";import { readFileSync, writeFileSync } from "node:fs";import { Sandbox } from "e2b";
const [taskFile, branch, acceptance] = process.argv.slice(2); // e.g. task.md agent/billing "npm test -- src/billing"const git = (...args: string[]) => execFileSync("git", args, { encoding: "utf8" }).trim();
const base = git("rev-parse", "HEAD");git("bundle", "create", "/tmp/in.bundle", "HEAD");
const sbx = await Sandbox.create("anthropic-claude-code", { timeoutMs: 60 * 60 * 1000, // one hour, the Hobby-tier maximum envs: { ANTHROPIC_API_KEY: process.env.ANTHROPIC_API_KEY! }, network: { allowOut: ["api.anthropic.com", "registry.npmjs.org"], denyOut: ["0.0.0.0/0"], },});
try { const sh = (cmd: string) => sbx.commands.run(cmd, { cwd: "/tmp/work", timeoutMs: 0 }); await sbx.commands.run("git --version && mkdir -p /tmp/work"); await sbx.files.write("/tmp/in.bundle", new Blob([readFileSync("/tmp/in.bundle")])); await sbx.files.write("/tmp/task.md", readFileSync(taskFile, "utf8")); await sh(`git clone -q /tmp/in.bundle . && git checkout -q -b ${branch}`); await sh(`git config user.email agent@sandbox.invalid && git config user.name sandbox-agent`); await sh("npm ci");
// Stage 2: the unattended loop. Safe only because of the egress rules above. const run = await sh("claude -p --dangerously-skip-permissions --max-budget-usd 5 < /tmp/task.md").catch((e) => e); console.log(`agent exit code: ${run.exitCode}`);
// Stage 3: the harness runs the acceptance check, not the agent. const check = await sh(acceptance).catch((e) => e); console.log(`acceptance exit code: ${check.exitCode}`);
// Stage 4: commit anything left uncommitted and bundle only the new commits. await sh(`git add -A && (git diff --cached --quiet || git commit -qm "agent: ${branch}")`); await sh(`git bundle create /tmp/out.bundle ${branch} ^${base}`); writeFileSync("/tmp/out.bundle", await sbx.files.read("/tmp/out.bundle", { format: "bytes" }));} finally { await sbx.kill();}Then, on your machine:
npx tsx scripts/sandbox-task.ts task.md agent/billing "npm test -- src/billing"git fetch /tmp/out.bundle agent/billing:agent/billinggit diff --stat HEAD...agent/billing # scope check before anything elsegit push -u origin agent/billing # CI and review take over from herecommands.run throws when a command exits non-zero, which is why the acceptance step catches the error and reads exitCode from it. A failed acceptance check still returns the branch, so you can see what the agent tried; the exit code tells the reviewer not to merge it. The bundle round trip was tested with git on 2026-09-26, and the E2B calls follow the type definitions of SDK 2.51.0; run the harness once on a throwaway task before you trust it.
Run the same loop with Codex or Cursor
Section titled “Run the same loop with Codex or Cursor”Only the agent install and the unattended command change.
Template: .npmInstall(["@anthropic-ai/claude-code"], { g: true }). Allow api.anthropic.com. Command:
claude -p --dangerously-skip-permissions --max-budget-usd 5 < /tmp/task.md--max-budget-usd caps API spend for the run (Claude Code v2.1.283 --help).
Template: .npmInstall(["@openai/codex"], { g: true }). Pass OPENAI_API_KEY in envs and allow api.openai.com (the default OpenAI provider’s host for API-key login; a custom provider needs the host in its model_providers base_url). Command (flags checked against codex exec --help, codex-cli 0.157.1):
printenv OPENAI_API_KEY | codex login --with-api-keycodex exec --dangerously-bypass-approvals-and-sandbox - < /tmp/task.mdThe help text says the bypass flag is “intended solely for running in environments that are externally sandboxed”, which is what the microVM is. With -, codex exec reads the prompt from stdin.
Docker’s Cursor agent page shows the CLI agent binary as cursor-agent (sbx run cursor runs cursor-agent --yolo). Its install script is hosted on cursor.com, which we could not verify on 2026-09-26, so no E2B template is shown. In the cloud, Cursor’s own Cloud Agents run in a vendor VM, covered in ephemeral environments.
Set up each sandbox with Claude Code, Codex and Cursor
Section titled “Set up each sandbox with Claude Code, Codex and Cursor”container-use: a container and a branch per agent
Section titled “container-use: a container and a branch per agent”container-use is an MCP server, so the agent itself creates and uses environments through its tools. Install the binary first (it is not on npm or PyPI):
brew install dagger/tap/container-use# or, on any platformcurl -fsSL https://raw.githubusercontent.com/dagger/container-use/main/install.sh | bashcd /path/to/repositoryclaude mcp add container-use -- container-use stdiocurl https://raw.githubusercontent.com/dagger/container-use/main/rules/agent.md >> CLAUDE.mdcodex mcp add container-use -- container-use stdiomkdir -p ~/.codexcurl https://raw.githubusercontent.com/dagger/container-use/main/rules/agent.md >> ~/.codex/AGENTS.mdcontainer-use’s own guide shows the equivalent [mcp_servers.container-use] block in ~/.codex/config.toml with command = "container-use" and args = ["stdio"].
Add the server to .cursor/mcp.json, then add the Cursor rules file:
{ "mcpServers": { "container-use": { "command": "container-use", "args": ["stdio"] } }}curl --create-dirs -o .cursor/rules/container-use.mdc https://raw.githubusercontent.com/dagger/container-use/main/rules/cursor.mdcPrompt the agent with a real task, such as “Create a Flask hello-world app in Python.” Your working tree stays unchanged; the agent reports an app URL and an environment ID such as fancy-mallard. Review and decide from the terminal:
container-use list # every environmentcontainer-use diff fancy-mallard # what the agent changedcontainer-use log fancy-mallard # every command it rancontainer-use merge fancy-mallard # keep its commits, or: container-use apply fancy-mallard to stage themContext cost. container-use adds MCP tools to every session; its quickstart’s --allowedTools list names ten of them. Run /context in Claude Code before and after registering the server to see what it costs your window.
Docker Sandboxes: a microVM per agent on your machine
Section titled “Docker Sandboxes: a microVM per agent on your machine”Install sbx (macOS 14+ on Apple silicon, Windows 11 with Hypervisor Platform, or Linux with KVM), then sign in:
brew trust docker/tap && brew install docker/tap/sbx # Windows: winget install -h Docker.sbxsbx loginsbx secret set anthropic # or use /login inside the sandbox with a Claude subscriptionsbx run --clone claude .Docker’s agent page says Claude Code starts with --dangerously-skip-permissions by default.
sbx secret set openai # or: sbx secret set openai --oauthsbx run --clone codex .Docker starts Codex with --dangerously-bypass-approvals-and-sandbox.
sbx run --clone cursor .Docker starts Cursor’s CLI agent with --yolo; it supports an API key stored as a secret or OAuth.
Use --clone. Without it, sbx run mounts your current directory read-write and the agent’s edits land in your working tree as it writes them. In clone mode the agent edits a private clone, your repository is visible read-only at /run/sandbox/source, and sbx rm deletes the clone, so fetch the commits you want first. Clone mode does not work from a linked git worktree, only the main one. On first run, pick the Balanced or Locked Down network preset rather than Open, and inspect it with sbx policy ls. With Locked Down, allow the model host first: sbx policy allow network api.anthropic.com (or your provider’s host). Docker’s get-started page says that under Locked Down “even your model provider API is blocked unless you explicitly allow it”.
Modal, Daytona, Coder and the sandbox runtime
Section titled “Modal, Daytona, Coder and the sandbox runtime”-
Modal. Modal’s example (
modal-labs/modal-examples,13_sandboxes/sandbox_agent.py) installs Claude Code in amodal.Image, clones a repository and runs the call below. Keeppty=True: the example says Claude requires it. It runs as root with a read-only question, so it omits--dangerously-skip-permissions.sandbox.exec("claude", "-p", "What is in this repository?",pty=True,secrets=[modal.Secret.from_name("anthropic-secret", required_keys=["ANTHROPIC_API_KEY"])],workdir="/repo",) -
Daytona.
pip install daytonaornpm install @daytona/sdk. Create withdaytona.create()and run commands throughsandbox.process.exec(...). Setauto_delete_interval(orephemeral),ttl_minutesfor a hard wall-clock cap, andnetwork_block_allordomain_allow_listexplicitly, because a stopped sandbox is kept by default. Daytona’ssecretsparameter mounts an organisation secret as a placeholder and substitutes the real value only on outbound requests to that secret’s allowed hosts. -
Coder.
curl -fsSL https://coder.com/install.sh | sh, thencoder server. Its registry modules run Claude Code, Codex and OpenCode isolated in workspaces, and the AI Gateway centralises authentication, audit and cost. Coder and CloudCLI are AGPL-licensed, so clear them with legal before you embed or host them. -
Claude Code sandbox runtime. Allow writes to your project,
~/.claude,~/.claude.jsonand/tmp, and allowapi.anthropic.comin~/.srt-settings.json. The keys arefilesystem.allowWriteandnetwork.allowedDomains(sandbox-runtime README, 0.0.77):{"filesystem": { "allowWrite": [".", "~/.claude", "~/.claude.json", "/tmp"] },"network": { "allowedDomains": ["api.anthropic.com"] }}Then run
npx @anthropic-ai/sandbox-runtime claude. With no~/.srt-settings.jsonat all,srtstill starts on its built-in defaults (no network, writes confined), so a clean start does not prove your file was found; a file that exists but does not validate makessrtexit with an error (README, 0.0.77).
How do you prove the sandbox boundary holds?
Section titled “How do you prove the sandbox boundary holds?”A sandbox you have not tested is a belief. Run these four checks when you adopt a sandbox and again after every template change. The tech lead owns the sandbox policy and signs off on the checks; the pull request reviewer signs off on the change itself.
| Check | How | Pass condition |
|---|---|---|
| Egress canary | Inside the sandbox, curl -sS -m 5 https://example.com | Fails; curl -sS -m 5 https://api.anthropic.com connects |
| Credential canary | Inside the sandbox, env | grep -i -E 'token|secret|key' | Only the model key (or a Daytona placeholder) appears; no GitHub, cloud or production values |
| Scope check | On your machine, git diff --stat HEAD...agent/<branch> | Only the paths the task allowed changed |
| Teardown check | E2B Sandbox.list(), sbx ls, container-use list | No sandbox left from finished runs |
Quality of the code itself is proven the usual way: the harness’s acceptance exit code, CI on the pushed branch, and a review agent or bot before a human approves. The sandbox makes the run safe; it does not make the output correct.
Copy-paste prompts for agent sandboxes
Section titled “Copy-paste prompts for agent sandboxes”What breaks when agents run in sandboxes, and how to recover
Section titled “What breaks when agents run in sandboxes, and how to recover”- The agent reports success and the acceptance check fails. Recovery: keep the branch, read
BLOCKED.mdif present, tighten the task file and rerun; do not merge on the agent’s summary. - The run dies at the sandbox timeout. Recovery: set the lifetime explicitly (E2B
timeoutMs, Modaltimeout), split the task, and bundle work at checkpoints so a timeout loses minutes, not the run. - Dependency installs fail under a strict allowlist. The registry or a postinstall download host is blocked. Recovery: install dependencies in the template at build time so run time needs only the model API, or add the exact hosts, never
0.0.0.0/0. - Sandboxes pile up and bill you. A crashed harness skips
kill(); Daytona keeps stopped sandboxes by default; Docker Sandboxes persist untilsbx rm. Recovery:try/finallyaround every run, an explicit auto-delete setting plus Daytona’sttl_minuteswall-clock cap, and a scheduled sweep that lists and removes sandboxes older than your maximum run time. - The agent’s edits appear in your working tree. Docker Sandboxes without
--clonemounts the directory read-write. Recovery:git stashor discard, then always create sandboxes with--clone. - Claude Code refuses
--dangerously-skip-permissions. The process runs as root. Recovery: run as a non-root user in the image, or use Docker Sandboxes, which starts the agent in bypass mode for you. git bundle createfails with an empty bundle. The agent made no commits and changed nothing. Recovery: treat it as a failed run, read the agent’s final message and the acceptance output, and fix the task file.- Claude Code hangs or prints nothing in Modal. The process has no PTY. Recovery: pass
pty=Truetosandbox.exec. - A prompt injection exfiltrates data. Any sandbox with open egress can leak what the agent can read. Recovery: deny all egress except named hosts, keep production credentials out of the sandbox, and rotate any key that was present during a suspicious run.
Where to go next with agent sandboxes
Section titled “Where to go next with agent sandboxes”- Permissions and sandboxing: the approval modes and OS sandboxes inside Claude Code, Codex and Cursor, which the tools on this page wrap from the outside.
- Ephemeral environments: give each sandboxed run its own seeded database, stubs and port block.
- Headless agents in CI: run the same loop from a pipeline instead of your terminal.
- Autonomous loops: long-running loops that need exactly this blast-radius limit.
- AI code review bots compared: the review gate the returned branch goes through.
- Agent tools overview: the other shelves, from multiplexers to review bots.
Frequently asked questions
Which sandbox should I use to run Claude Code or Codex unattended?
For local runs, Docker Sandboxes (a microVM per agent) or container-use (a container and git branch per agent). For product or CI runs, a cloud SDK such as E2B, Modal or Daytona. For code that must stay on your own infrastructure, Coder.
Is --dangerously-skip-permissions safe inside a sandbox?
Only inside a disposable, network-restricted sandbox that holds no production credentials, running as a non-root user. Claude Code's own help recommends the flag only for sandboxes with no internet access.
Can I self-host Daytona from GitHub?
Not as a maintained option. The Daytona repository has been unmaintained since June 2026 because core development moved to a private codebase; the hosted service and the daytona and @daytona/sdk packages continue.
Is container-use on npm?
No. container-use is a Go binary installed with Homebrew or an install script and registered as an MCP server with the command container-use stdio. An npx container-use configuration does not work.