Skip to content

Agent sandboxes compared: container-use, E2B, Daytona, microVMs and Docker

An agent sandbox is a disposable machine that a coding agent such as Claude Code or Codex runs inside, so an unattended run can do anything without reaching your laptop, credentials or production. The options differ in isolation: OS-level wrapping (Claude Code sandbox runtime), containers (container-use, Modal), microVMs (E2B, Docker Sandboxes), and managed or self-hosted platforms (Daytona, Coder).

This page is for the developer who wants an agent to work for an hour without approval prompts, and the tech lead who has to decide where that is allowed. You tried auto mode, then someone on the team ran claude -p --dangerously-skip-permissions on a laptop with a production .env in the repository. The question now is which sandbox makes the next run safe by construction, and how the work gets back into review.

  • A decision table and a dated comparison of seven sandboxes on isolation, startup, lifetime, network control and price.
  • A worked example: Claude Code headless in E2B, from the E2B cookbook, plus Codex.
  • A workflow: an unattended loop in a disposable VM whose branch comes back as a git bundle for review.
  • Three copy-paste prompts, traps and failure modes.

The pattern behind all of this (a seeded, per-task stack with its own ports, data and credentials) is in ephemeral environments. The approval modes and OS sandboxes built into each agent are in permissions and sandboxing.

Start from where the run happens and who owns the machine.

Your situationPickWhy
One developer, several agents, on a laptopDocker Sandboxes (sbx) or container-useA microVM (Docker) or container (container-use) per agent, each on its own clone or branch; free for local use
Unattended runs of Claude Code only, no Docker allowedClaude Code sandbox runtime (@anthropic-ai/sandbox-runtime)Wraps the whole Claude Code process, including hooks and MCP servers, in Seatbelt or bubblewrap
Agent runs started by your CI, a queue or your productE2B, Modal or DaytonaSDK-created sandboxes with per-sandbox egress control and a lifetime you set in code
Code and model credentials must stay on your own infrastructureCoderSelf-hosted workspaces in Terraform; Coder Agents keeps “no API keys in workspaces”
A web or phone UI over sandboxed agentsCloudCLI sandbox (experimental)Wraps Docker Sandboxes behind CloudCLI’s UI
Background tasks started from an issue or chat, no infrastructureThe vendors’ own cloud environmentsCovered in ephemeral environments, not here

A worktree is not on this list on purpose. It isolates files only; ports, databases, Docker names and every credential in your shell stay shared (see parallel agents).

How do the sandboxes compare on isolation, startup, lifetime and price?

Section titled “How do the sandboxes compare on isolation, startup, lifetime and price?”

Facts below were checked on 2026-09-26 against each project’s README, SDK type definitions or documentation source on GitHub. The vendor websites of E2B, Modal and Daytona were unreachable from our environment on that date, so no price or startup figure from those sites appears here.

SandboxIsolation modelRuns whereWhat drives startupDefault lifetimeNetwork controlPricing
Claude Code sandbox runtime (npm @anthropic-ai/sandbox-runtime 0.0.77, beta research preview)OS primitives around the whole process: sandbox-exec (Seatbelt) on macOS, bubblewrap on Linux, and a Windows alpha (srt-sandbox account plus WFP egress filter) as of 0.0.77Your machineNothing to boot; it wraps the processThe processDenies network by default; you allow domains in ~/.srt-settings.jsonFree npm package
container-use (Dagger, Apache-2.0, “experimental”)A fresh container per agent, each on its own git branchYour machine, via DaggerImage build on first useUntil you merge, apply or discard the environmentNot documented in the READMEFree, open source
Docker Sandboxes (sbx CLI)A microVM per sandbox with its own Docker daemon, filesystem and networkYour machine, or Docker-managed cloudFirst run pulls the agent image; Docker says later runs “start in seconds”Persists after the agent exits until sbx rmGlobal presets Open, Balanced or Locked Down, plus sbx policy allow network <host>CLI and local compute free, including commercial use; cloud is pay-as-you-go
E2B (npm e2b 2.51.0)One Firecracker microVM per sandbox, own cgroup and network namespaceE2B Cloud, or self-hosted from e2b-dev/infraBoots or resumes from a snapshot of a template you build once300 s unless you set timeoutMs; maximum 1 hour on Hobby and 24 hours on Pro (SDK docs)network.allowOut and denyOut (hostnames, IPs, CIDRs), or allowInternetAccess: falseTiered (Hobby, Pro); prices not verified here
Modal Sandboxes (PyPI modal 1.5.5)A container created on demand with modal.Sandbox.createModal’s cloudBuilding the modal.Image you definetimeout=300 seconds by defaultblock_network=True, outbound_domain_allowlist, outbound_cidr_allowlistNot verified here
Daytona (PyPI daytona, npm @daytona/sdk 0.218.0)Hosted sandboxes created through the SDKDaytona’s hosted serviceCreating from a snapshotAuto-stops after 15 idle minutes; auto-delete is off unless you set itnetwork_block_all, domain_allow_list, network_allow_listNot verified here
Coder (AGPL-3.0; Premium tier paid)Workspaces you define in TerraformYour infrastructureYour Terraform templateSet by your templatesYours to configureOpen source core; Premium is paid

Daytona’s open-source repository says: “This repository is no longer maintained. As of June 2026, Daytona’s core development has moved to a private codebase.” The hosted service and SDKs continue (both packages released 0.218.0 on 2026-09-25). CloudCLI’s sandbox is marked experimental and needs Docker’s sbx CLI; CloudCLI’s managed cloud starts at €7 a month, separate from your model subscription.

Popularity as of 2026-09-26 (GitHub stars, read through the GitHub API on that date): Daytona 71,717 (repository frozen), Coder 16,696, E2B 13,971, CloudCLI 13,811, container-use 4,046 and modal-labs/modal-client 518 (the SDK repository only). Stars measure attention, not fitness for your threat model.

Run Claude Code headless in an E2B sandbox

Section titled “Run Claude Code headless in an E2B sandbox”

This is the E2B cookbook’s own example (e2b-dev/e2b-cookbook, examples/anthropic-claude-code-in-sandbox-js), checked on 2026-09-26. It builds a template once, then runs one Claude Code prompt in a sandbox and kills it.

  1. Install the SDK and the runner, and export both keys from your secret store (never paste them into the file):

    Terminal window
    npm install e2b dotenv
    npm install -D tsx
    export E2B_API_KEY ANTHROPIC_API_KEY # values come from your secret manager
  2. Define the template. The cookbook names it anthropic-claude-code; that name is yours to choose, and it is not a pre-built E2B template:

    src/template.ts
    import { Template } from "e2b";
    export const templateName = "anthropic-claude-code";
    export const template = Template()
    .fromNodeImage("24")
    .npmInstall(["@anthropic-ai/claude-code"], { g: true });
  3. Build it once. The cookbook uses one vCPU and 1,024 MB of memory:

    src/build.prod.ts
    import "dotenv/config";
    import { Template, defaultBuildLogger } from "e2b";
    import { template, templateName } from "./template";
    await Template.build(template, {
    alias: templateName,
    cpuCount: 1,
    memoryMB: 1024,
    onBuildLogs: defaultBuildLogger(),
    });
    Terminal window
    npx tsx src/build.prod.ts
  4. Run a prompt. timeoutMs: 0 on the command removes the per-command timeout, because a Claude Code run can take a long time:

    src/index.ts
    import "dotenv/config";
    import { Sandbox } from "e2b";
    import { templateName } from "./template";
    const sbx = await Sandbox.create(templateName, {
    envs: { ANTHROPIC_API_KEY: process.env.ANTHROPIC_API_KEY! },
    // Not in the cookbook: deny all egress except the model API.
    network: { allowOut: ["api.anthropic.com"], denyOut: ["0.0.0.0/0"] },
    });
    console.log("Sandbox created", sbx.sandboxId);
    const result = await sbx.commands.run(
    `echo 'Create a hello world index.html' | claude -p --dangerously-skip-permissions`,
    { timeoutMs: 0 },
    );
    console.log(result.stdout);
    await sbx.kill();
    Terminal window
    npx tsx src/index.ts

You should see Sandbox created with an ID, then Claude Code’s final message describing the index.html it wrote. The file exists only inside the microVM, sbx.kill() destroys it, and nothing comes back to you. The next section turns it into a loop whose output you can review.

How do you run —dangerously-skip-permissions safely and still review the result?

Section titled “How do you run —dangerously-skip-permissions safely and still review the result?”

The loop has five stages, and the sandbox is only one of them: specify → build in a disposable VM → verify with the harness, not the agent → return a branch → review and merge. The agent never holds a GitHub token and never pushes. Your harness moves code in and out as git bundles, so the sandbox’s only outbound connections are the model API and the package registry.

  1. Specify. Write the task as a file with acceptance criteria the harness can check (a test command that must pass). An agent told “make it work” will declare success; an agent told “npm test -- src/billing must pass and no file outside src/billing/ may change” can be held to it. See how to write acceptance criteria an agent can be held to.

  2. Build in a disposable VM. Create the sandbox with egress denied except for the hosts the run needs, upload the repository as a bundle, and run the agent unattended.

  3. Verify with the harness. After the agent exits, the harness runs the acceptance command itself and records the exit code. The agent’s “all tests pass” is not evidence.

  4. Return a branch. The harness bundles the agent’s commits and downloads the bundle. Nothing is pushed from inside the sandbox.

  5. Review and merge. You fetch the bundle into a local branch, push it, and open a pull request. CI, a review bot and a human reviewer see it like any other change.

This harness implements stages 2 to 4 for Claude Code on E2B, reusing the cookbook’s template. Run it from the root of your repository:

scripts/sandbox-task.ts
import { execFileSync } from "node:child_process";
import { readFileSync, writeFileSync } from "node:fs";
import { Sandbox } from "e2b";
const [taskFile, branch, acceptance] = process.argv.slice(2); // e.g. task.md agent/billing "npm test -- src/billing"
const git = (...args: string[]) => execFileSync("git", args, { encoding: "utf8" }).trim();
const base = git("rev-parse", "HEAD");
git("bundle", "create", "/tmp/in.bundle", "HEAD");
const sbx = await Sandbox.create("anthropic-claude-code", {
timeoutMs: 60 * 60 * 1000, // one hour, the Hobby-tier maximum
envs: { ANTHROPIC_API_KEY: process.env.ANTHROPIC_API_KEY! },
network: {
allowOut: ["api.anthropic.com", "registry.npmjs.org"],
denyOut: ["0.0.0.0/0"],
},
});
try {
const sh = (cmd: string) => sbx.commands.run(cmd, { cwd: "/tmp/work", timeoutMs: 0 });
await sbx.commands.run("git --version && mkdir -p /tmp/work");
await sbx.files.write("/tmp/in.bundle", new Blob([readFileSync("/tmp/in.bundle")]));
await sbx.files.write("/tmp/task.md", readFileSync(taskFile, "utf8"));
await sh(`git clone -q /tmp/in.bundle . && git checkout -q -b ${branch}`);
await sh(`git config user.email agent@sandbox.invalid && git config user.name sandbox-agent`);
await sh("npm ci");
// Stage 2: the unattended loop. Safe only because of the egress rules above.
const run = await sh("claude -p --dangerously-skip-permissions --max-budget-usd 5 < /tmp/task.md").catch((e) => e);
console.log(`agent exit code: ${run.exitCode}`);
// Stage 3: the harness runs the acceptance check, not the agent.
const check = await sh(acceptance).catch((e) => e);
console.log(`acceptance exit code: ${check.exitCode}`);
// Stage 4: commit anything left uncommitted and bundle only the new commits.
await sh(`git add -A && (git diff --cached --quiet || git commit -qm "agent: ${branch}")`);
await sh(`git bundle create /tmp/out.bundle ${branch} ^${base}`);
writeFileSync("/tmp/out.bundle", await sbx.files.read("/tmp/out.bundle", { format: "bytes" }));
} finally {
await sbx.kill();
}

Then, on your machine:

Terminal window
npx tsx scripts/sandbox-task.ts task.md agent/billing "npm test -- src/billing"
git fetch /tmp/out.bundle agent/billing:agent/billing
git diff --stat HEAD...agent/billing # scope check before anything else
git push -u origin agent/billing # CI and review take over from here

commands.run throws when a command exits non-zero, which is why the acceptance step catches the error and reads exitCode from it. A failed acceptance check still returns the branch, so you can see what the agent tried; the exit code tells the reviewer not to merge it. The bundle round trip was tested with git on 2026-09-26, and the E2B calls follow the type definitions of SDK 2.51.0; run the harness once on a throwaway task before you trust it.

Only the agent install and the unattended command change.

Template: .npmInstall(["@anthropic-ai/claude-code"], { g: true }). Allow api.anthropic.com. Command:

Terminal window
claude -p --dangerously-skip-permissions --max-budget-usd 5 < /tmp/task.md

--max-budget-usd caps API spend for the run (Claude Code v2.1.283 --help).

Set up each sandbox with Claude Code, Codex and Cursor

Section titled “Set up each sandbox with Claude Code, Codex and Cursor”

container-use: a container and a branch per agent

Section titled “container-use: a container and a branch per agent”

container-use is an MCP server, so the agent itself creates and uses environments through its tools. Install the binary first (it is not on npm or PyPI):

Terminal window
brew install dagger/tap/container-use
# or, on any platform
curl -fsSL https://raw.githubusercontent.com/dagger/container-use/main/install.sh | bash
Terminal window
cd /path/to/repository
claude mcp add container-use -- container-use stdio
curl https://raw.githubusercontent.com/dagger/container-use/main/rules/agent.md >> CLAUDE.md

Prompt the agent with a real task, such as “Create a Flask hello-world app in Python.” Your working tree stays unchanged; the agent reports an app URL and an environment ID such as fancy-mallard. Review and decide from the terminal:

Terminal window
container-use list # every environment
container-use diff fancy-mallard # what the agent changed
container-use log fancy-mallard # every command it ran
container-use merge fancy-mallard # keep its commits, or: container-use apply fancy-mallard to stage them

Context cost. container-use adds MCP tools to every session; its quickstart’s --allowedTools list names ten of them. Run /context in Claude Code before and after registering the server to see what it costs your window.

Docker Sandboxes: a microVM per agent on your machine

Section titled “Docker Sandboxes: a microVM per agent on your machine”

Install sbx (macOS 14+ on Apple silicon, Windows 11 with Hypervisor Platform, or Linux with KVM), then sign in:

Terminal window
brew trust docker/tap && brew install docker/tap/sbx # Windows: winget install -h Docker.sbx
sbx login
Terminal window
sbx secret set anthropic # or use /login inside the sandbox with a Claude subscription
sbx run --clone claude .

Docker’s agent page says Claude Code starts with --dangerously-skip-permissions by default.

Use --clone. Without it, sbx run mounts your current directory read-write and the agent’s edits land in your working tree as it writes them. In clone mode the agent edits a private clone, your repository is visible read-only at /run/sandbox/source, and sbx rm deletes the clone, so fetch the commits you want first. Clone mode does not work from a linked git worktree, only the main one. On first run, pick the Balanced or Locked Down network preset rather than Open, and inspect it with sbx policy ls. With Locked Down, allow the model host first: sbx policy allow network api.anthropic.com (or your provider’s host). Docker’s get-started page says that under Locked Down “even your model provider API is blocked unless you explicitly allow it”.

Section titled “Modal, Daytona, Coder and the sandbox runtime”
  • Modal. Modal’s example (modal-labs/modal-examples, 13_sandboxes/sandbox_agent.py) installs Claude Code in a modal.Image, clones a repository and runs the call below. Keep pty=True: the example says Claude requires it. It runs as root with a read-only question, so it omits --dangerously-skip-permissions.

    sandbox.exec(
    "claude", "-p", "What is in this repository?",
    pty=True,
    secrets=[modal.Secret.from_name("anthropic-secret", required_keys=["ANTHROPIC_API_KEY"])],
    workdir="/repo",
    )
  • Daytona. pip install daytona or npm install @daytona/sdk. Create with daytona.create() and run commands through sandbox.process.exec(...). Set auto_delete_interval (or ephemeral), ttl_minutes for a hard wall-clock cap, and network_block_all or domain_allow_list explicitly, because a stopped sandbox is kept by default. Daytona’s secrets parameter mounts an organisation secret as a placeholder and substitutes the real value only on outbound requests to that secret’s allowed hosts.

  • Coder. curl -fsSL https://coder.com/install.sh | sh, then coder server. Its registry modules run Claude Code, Codex and OpenCode isolated in workspaces, and the AI Gateway centralises authentication, audit and cost. Coder and CloudCLI are AGPL-licensed, so clear them with legal before you embed or host them.

  • Claude Code sandbox runtime. Allow writes to your project, ~/.claude, ~/.claude.json and /tmp, and allow api.anthropic.com in ~/.srt-settings.json. The keys are filesystem.allowWrite and network.allowedDomains (sandbox-runtime README, 0.0.77):

    {
    "filesystem": { "allowWrite": [".", "~/.claude", "~/.claude.json", "/tmp"] },
    "network": { "allowedDomains": ["api.anthropic.com"] }
    }

    Then run npx @anthropic-ai/sandbox-runtime claude. With no ~/.srt-settings.json at all, srt still starts on its built-in defaults (no network, writes confined), so a clean start does not prove your file was found; a file that exists but does not validate makes srt exit with an error (README, 0.0.77).

How do you prove the sandbox boundary holds?

Section titled “How do you prove the sandbox boundary holds?”

A sandbox you have not tested is a belief. Run these four checks when you adopt a sandbox and again after every template change. The tech lead owns the sandbox policy and signs off on the checks; the pull request reviewer signs off on the change itself.

CheckHowPass condition
Egress canaryInside the sandbox, curl -sS -m 5 https://example.comFails; curl -sS -m 5 https://api.anthropic.com connects
Credential canaryInside the sandbox, env | grep -i -E 'token|secret|key'Only the model key (or a Daytona placeholder) appears; no GitHub, cloud or production values
Scope checkOn your machine, git diff --stat HEAD...agent/<branch>Only the paths the task allowed changed
Teardown checkE2B Sandbox.list(), sbx ls, container-use listNo sandbox left from finished runs

Quality of the code itself is proven the usual way: the harness’s acceptance exit code, CI on the pushed branch, and a review agent or bot before a human approves. The sandbox makes the run safe; it does not make the output correct.

What breaks when agents run in sandboxes, and how to recover

Section titled “What breaks when agents run in sandboxes, and how to recover”
  • The agent reports success and the acceptance check fails. Recovery: keep the branch, read BLOCKED.md if present, tighten the task file and rerun; do not merge on the agent’s summary.
  • The run dies at the sandbox timeout. Recovery: set the lifetime explicitly (E2B timeoutMs, Modal timeout), split the task, and bundle work at checkpoints so a timeout loses minutes, not the run.
  • Dependency installs fail under a strict allowlist. The registry or a postinstall download host is blocked. Recovery: install dependencies in the template at build time so run time needs only the model API, or add the exact hosts, never 0.0.0.0/0.
  • Sandboxes pile up and bill you. A crashed harness skips kill(); Daytona keeps stopped sandboxes by default; Docker Sandboxes persist until sbx rm. Recovery: try/finally around every run, an explicit auto-delete setting plus Daytona’s ttl_minutes wall-clock cap, and a scheduled sweep that lists and removes sandboxes older than your maximum run time.
  • The agent’s edits appear in your working tree. Docker Sandboxes without --clone mounts the directory read-write. Recovery: git stash or discard, then always create sandboxes with --clone.
  • Claude Code refuses --dangerously-skip-permissions. The process runs as root. Recovery: run as a non-root user in the image, or use Docker Sandboxes, which starts the agent in bypass mode for you.
  • git bundle create fails with an empty bundle. The agent made no commits and changed nothing. Recovery: treat it as a failed run, read the agent’s final message and the acceptance output, and fix the task file.
  • Claude Code hangs or prints nothing in Modal. The process has no PTY. Recovery: pass pty=True to sandbox.exec.
  • A prompt injection exfiltrates data. Any sandbox with open egress can leak what the agent can read. Recovery: deny all egress except named hosts, keep production credentials out of the sandbox, and rotate any key that was present during a suspicious run.

Frequently asked questions

Which sandbox should I use to run Claude Code or Codex unattended?

For local runs, Docker Sandboxes (a microVM per agent) or container-use (a container and git branch per agent). For product or CI runs, a cloud SDK such as E2B, Modal or Daytona. For code that must stay on your own infrastructure, Coder.

Is --dangerously-skip-permissions safe inside a sandbox?

Only inside a disposable, network-restricted sandbox that holds no production credentials, running as a non-root user. Claude Code's own help recommends the flag only for sandboxes with no internet access.

Can I self-host Daytona from GitHub?

Not as a maintained option. The Daytona repository has been unmaintained since June 2026 because core development moved to a private codebase; the hosted service and the daytona and @daytona/sdk packages continue.

Is container-use on npm?

No. container-use is a Go binary installed with Homebrew or an install script and registered as an MCP server with the command container-use stdio. An npx container-use configuration does not work.