Ephemeral environments: a full stack per agent task
An ephemeral environment for an agent task is a disposable, seeded copy of the whole stack (code checkout, ports, database, stubbed third-party services and test-only credentials) that exists for one run and is destroyed after it. Each agent then proves behaviour end to end without colliding with parallel runs or touching shared data. A worktree isolates only the files.
This page is for the developer who runs several agents at once and the tech lead who wants each of those runs to leave evidence a reviewer can trust. Three agents each got a worktree. The first started the dev server on port 3000, the second’s server silently moved to 3001, and the third’s end-to-end suite attached to the first agent’s server and reported green against code it never tested. Meanwhile two of them ran migrations against the same local Postgres. Every run “passed”, and none of them proved anything.
What you’ll walk away with from a per-task environment
Section titled “What you’ll walk away with from a per-task environment”- A six-part environment contract (checkout, port block, data, stubs, credentials, lifetime) that says what each run gets to itself.
- Three scripts you can adapt:
bin/env-upallocates a port block and starts a seeded stack under its own Compose project,bin/env-checkproves it is live and isolated, andbin/env-downremoves it. - The wiring for each tool: Claude Code worktrees and cloud environments, Codex worktrees and cloud tasks, and Cursor worktrees and Cloud Agents.
- Three copy-paste prompts: build the contract for your repository, run a task inside it, and audit isolation with two concurrent runs.
- An isolation test that proves the setup works before you trust it with unattended runs, and the failure modes that make a green run meaningless.
What does a worktree not isolate?
Section titled “What does a worktree not isolate?”A git worktree gives each agent its own files and branch. Everything the code talks to at runtime is still shared: network ports, the database, caches, message queues, Docker container names, third-party sandboxes and any credential in your shell. An agent that can only run unit tests does not care. An agent that must prove a feature works end to end, which is the point of verifying behaviour instead of reading diffs, needs all of it, and needs it to itself.
The environment is the seventh layer of the harness. The permissions and sandbox layer limits what a run may do; the environment decides what the run can see and break. An environment with no production credentials in it is also the strongest permission rule there is: the agent cannot misuse a key that does not exist on the machine.
Which six parts make up the environment contract?
Section titled “Which six parts make up the environment contract?”| Part | Isolated by | What goes wrong without it |
|---|---|---|
| Checkout | A worktree locally, a fresh clone in a cloud VM | Two agents edit the same files; a checkout carries another task’s uncommitted changes |
| Port block | A fixed block of ports per task index, written to a file every tool reads | Dev servers fall back to the next free port and tests hit a neighbour’s server |
| Data | A per-task database created from migrations and a deterministic seed | Runs corrupt each other’s rows; a test passes only because another run left data behind |
| Service stubs | A stub server per task for each third-party API, with committed mappings | Tests call a real payment, email or ads API, or fail when its sandbox is down |
| Credentials | A generated env file with test-mode values only | A prompt injection or a wrong command reaches production, real customers or real money |
| Lifetime | A teardown command and a time-to-live sweep | Stacks leak until the machine runs out of disk, memory or ports |
A run proves behaviour only when all six hold.
How much isolation does each run need?
Section titled “How much isolation does each run need?”Pick the lightest level that isolates everything the task touches. The tool catalogue, with each product’s isolation model and a worked example, is agent sandboxes compared; this page covers the pattern.
| Level | What it isolates | Use it when | Main cost |
|---|---|---|---|
| Worktree + port block + Compose project on your machine | Files, ports, containers, database | You run two to five agents locally and review their results the same day | Your machine’s CPU and memory; you own teardown |
| Container per task (for example container-use, which gives each agent a container on its own git branch) | The above plus the process tree and installed tools | Tasks install packages or need different toolchains | Image build time; Docker on every machine |
| Vendor cloud environment (Claude Code cloud sessions, Codex cloud, Cursor Cloud Agents) | A whole VM per task, on the vendor’s infrastructure | Background runs, runs started from an issue or chat, and runs that must survive your laptop closing | Setup-script discipline; vendor resource limits; network allowlists |
| Self-hosted cloud runners | A VM or container on your infrastructure | The code or data may not leave your network | You operate the fleet |
To try the container-per-task level with Claude Code, install container-use (brew install dagger/tap/container-use) and register it as an MCP server with claude mcp add container-use -- container-use stdio; it is not published on npm, so an npx container-use config does not work.
Whichever level you choose, keep the environment definition in the repository as scripts, not in a vendor settings screen alone. The same bin/env-up then runs on your laptop, in a vendor cloud VM and in CI, so a result from one is comparable with the others.
Build the environment contract for your repository
Section titled “Build the environment contract for your repository”The example is a Node.js service with Postgres and a third-party payments API. Swap the image names and commands for your stack; keep the structure.
-
Allocate a port block per task. Each task gets an index from 1 to 49, and every port is a base plus ten times the index. The stride of ten keeps a server that falls back to the next free port inside its own block. Allocation uses
mkdir, which is atomic, so two agents starting at the same moment cannot take the same index.#!/usr/bin/env bash# bin/env-up — start an isolated, seeded stack for one agent taskset -euo pipefailslug=$(printf '%s' "${1:?usage: bin/env-up <task-slug>}" | tr '[:upper:]' '[:lower:]' | tr -cs 'a-z0-9' '-' | sed 's/^-*//;s/-*$//')state="${AGENT_ENV_HOME:-$HOME/.agent-envs}"mkdir -p "$state"idx=""for dir in "$state"/*/; do # rerun of the same task: reuse its block[ -f "$dir/slug" ] && [ "$(cat "$dir/slug")" = "$slug" ] && idx=$(basename "$dir")doneif [ -z "$idx" ]; thenfor i in $(seq 1 49); doif mkdir "$state/$i" 2>/dev/null; then idx=$i; echo "$slug" > "$state/$i/slug"; break; fidonefi[ -n "$idx" ] || { echo "no free port block in $state" >&2; exit 1; }cat > .env.task <<EOFTASK_SLUG=$slugTASK_INDEX=$idxCOMPOSE_PROJECT_NAME=task-$slugAPP_PORT=$((3000 + idx * 10))DB_PORT=$((5432 + idx * 10))STUB_PORT=$((8080 + idx * 10))DATABASE_URL=postgres://app:app@127.0.0.1:$((5432 + idx * 10))/appPAYMENTS_API_URL=http://127.0.0.1:$((8080 + idx * 10))PAYMENTS_API_KEY=test_only_not_a_secretEOFset -a; . ./.env.task; set +adocker compose -p "$COMPOSE_PROJECT_NAME" --env-file .env.task -f compose.agent.yml up -d --waitnpm run db:migratenpm run db:seedecho "task $slug: app :$APP_PORT, db :$DB_PORT, stubs :$STUB_PORT"Add
.env.taskto.gitignore. The file is the single source of ports: the dev server, the test runner and the agent all read it, so no tool has to remember a number. -
Define the per-task services under a Compose project. The project name (
-p task-<slug>) namespaces every container, network and volume, so two tasks never share one. Never setcontainer_name: a fixed name defeats the project namespace and the second task fails to start.# compose.agent.yml — services one agent task needs, nothing sharedservices:db:image: postgres:16environment:POSTGRES_USER: appPOSTGRES_PASSWORD: appPOSTGRES_DB: appports: ["127.0.0.1:${DB_PORT}:5432"]tmpfs: /var/lib/postgresql/data # data dies with the containerhealthcheck:test: ["CMD-SHELL", "pg_isready -U app"]interval: 2sretries: 30payments-stub:image: wiremock/wiremockports: ["127.0.0.1:${STUB_PORT}:8080"]volumes: ["./stubs/payments:/home/wiremock"]Binding to
127.0.0.1keeps the stack off your LAN.--waitmakesenv-upreturn only when the database health check passes, so the agent’s first migration does not race the container. -
Seed deterministic data.
db:seedmust produce the same rows on every run: a fixed random seed, fixed timestamps and no calls to live services. Seed the cases your acceptance criteria name (an expired subscription, a user with two organisations, a refund in flight), not a generic sample. If seeding takes minutes, seed one template database once and create each task’s database from it with Postgres’sCREATE DATABASE task_db TEMPLATE app_template, which copies the files instead of replaying the seed. Test data management covers seed design in depth. -
Stub every third-party API. WireMock’s image reads stub mappings from
/home/wiremock/mappingsand serves them on port 8080 inside the container. Commit the mappings understubs/payments/mappings/, including the failure responses (declined card, timeout, 429) your code must handle. Point the application at the stub through the env file (PAYMENTS_API_URL), never through code the agent can edit. -
Generate test-only credentials. The env file carries sandbox or dummy values only. Real tokens stay out of the environment entirely; agent identity and secrets covers how to issue scoped credentials when a task genuinely needs one.
-
Prove the stack is live and isolated. The agent runs this before its first test and pastes the output into its report.
#!/usr/bin/env bash# bin/env-check — prove this task's stack is up, seeded and its ownset -euo pipefailset -a; . ./.env.task; set +adocker compose -p "$COMPOSE_PROJECT_NAME" -f compose.agent.yml --env-file .env.task ps --format '{{.Service}} {{.State}}'curl -fsS "http://127.0.0.1:$STUB_PORT/__admin/health" > /dev/null && echo "stubs: healthy"docker compose -p "$COMPOSE_PROJECT_NAME" -f compose.agent.yml --env-file .env.task exec -T db \psql -U app -d app -tAc "select 'seeded users: ' || count(*) from users"echo "project=$COMPOSE_PROJECT_NAME app=$APP_PORT db=$DB_PORT stubs=$STUB_PORT" -
Tear it down, and sweep what leaks.
env-downremoves the containers and volumes and frees the index. Runbin/env-sweepon a schedule for stacks whose task ended without teardown, such as a headless run that crashed.#!/usr/bin/env bash# bin/env-down — destroy this task's stack and release its port blockset -euo pipefailset -a; . ./.env.task; set +adocker compose -p "$COMPOSE_PROJECT_NAME" -f compose.agent.yml --env-file .env.task down -v --remove-orphansrm -rf "${AGENT_ENV_HOME:-$HOME/.agent-envs}/$TASK_INDEX" .env.taskThe sweep assumes each worktree directory is named after its task slug, as
claude --worktree task-142andbin/env-up task-142produce. It tears down every stack whose worktree is gone, so remove abandoned worktrees withgit worktree removefirst.#!/usr/bin/env bash# bin/env-sweep — remove stacks whose task has no live worktree (run from cron or CI)set -euo pipefailstate="${AGENT_ENV_HOME:-$HOME/.agent-envs}"live=$(git worktree list --porcelain | sed -n 's|^worktree ||p')for dir in "$state"/*/; do[ -f "$dir/slug" ] || continueslug=$(cat "$dir/slug")grep -q "/$slug\$" <<<"$live" && continue # worktree named after the slug still existsecho "sweeping task-$slug"docker compose -p "task-$slug" down -v --remove-orphansrm -rf "$dir"done
Finally, make the dev server and the end-to-end runner read the env file and fail if the port is taken: set Vite’s server.strictPort to true, and derive the dev server port, webServer.url and use.baseURL from APP_PORT, so the suite can only test this task’s server.
Wire the environment into Claude Code, Codex and Cursor
Section titled “Wire the environment into Claude Code, Codex and Cursor”The scripts are the same everywhere; what differs is who starts them and where the stack runs.
Local. Start each task in its own worktree with claude --worktree task-142 (or -w). Claude Code creates it under .claude/worktrees/. A worktree is a fresh checkout, so ask Claude to run bin/env-up task-142 first, or run it yourself. To copy gitignored files such as .env into every new worktree, list them in a .worktreeinclude file at the project root. Non-interactive -p runs do not clean up their worktrees, so pair them with your sweep.
Cloud sessions (claude.ai/code, claude --cloud "<task>", routines) run in an isolated VM configured by a cloud environment: network access level, environment variables and a setup script. Checked on 2026-09-26 against Anthropic’s cloud environments guide:
- The VM already has Docker with
docker compose, PostgreSQL 16 and Redis 7.0, on Ubuntu 24.04. Anthropic-hosted sessions have approximate ceilings of 4 vCPUs, 16 GB of RAM and 30 GB of disk. - The setup script runs as root before Claude Code launches, must exit zero, and is cached as a filesystem snapshot when it finishes in roughly five minutes; the cache rebuilds when you change the script or allowed hosts, or after roughly seven days. Running processes are not cached, so pull images in the setup script and start the stack per session.
- The default network level, Trusted, allows package registries, GitHub and Docker Hub. None makes installs and image pulls fail.
Put image pulls in the environment’s Setup script field:
#!/bin/bash# cloud environment setup script: cached, so pull here, start laterdocker pull postgres:16 || truedocker pull wiremock/wiremock || trueThen start the stack on every cloud session with a SessionStart hook in the repository’s .claude/settings.json. The hook runs locally too, so the script exits unless CLAUDE_CODE_REMOTE is true:
{ "hooks": { "SessionStart": [ { "matcher": "startup|resume", "hooks": [ { "type": "command", "command": "bash \"$CLAUDE_PROJECT_DIR\"/scripts/cloud-env-up.sh", "timeout": 300 } ] } ] }}#!/bin/bash# scripts/cloud-env-up.sh — one task per VM, so one fixed slug is enough[ "$CLAUDE_CODE_REMOTE" = "true" ] || exit 0cd "$CLAUDE_PROJECT_DIR" && bin/env-up cloudA session with several repositories does not load hooks from any repository’s .claude/settings.json, so in that case put “run bin/env-up cloud first” at the top of the task prompt instead. To run cloud sessions on your own compute, Team and Enterprise can use self-hosted environments (public beta) and start a session on one with --environment <id>. Bring a finished cloud session to your terminal with claude --teleport.
Local. codex --worktree runs the session in a new managed git worktree (Codex CLI 0.157.1). Run bin/env-up inside it before the task, and name the command in AGENTS.md so Codex runs it unprompted.
Cloud. Codex cloud runs tasks “in isolated cloud environments”, and each environment is where you “customize dependencies and tools for Codex” (OpenAI docs, checked 2026-08-28; the docs host was unreachable from the writing environment on 2026-09-26). From the terminal, the experimental codex cloud command submits and collects tasks:
# terminal, repository root — ENV_ID comes from browsing `codex cloud`codex cloud exec --env ENV_ID --attempts 2 "Run bin/env-up cloud, then implement issue 142 and prove it with the e2e suite"codex cloud listcodex cloud diff TASK_ID # review the unified diffcodex cloud apply TASK_ID --attempt N # apply the attempt you accepted, then run the checks yourself--attempts asks for best-of-N attempts. Review each one separately with codex cloud diff TASK_ID --attempt N, and accept only an attempt whose report includes the bin/env-check output and a green end-to-end run.
To debug the environment before you rely on it, OpenAI publishes a reference of the cloud base image, ghcr.io/openai/codex-universal, that you can run locally with the same runtime variables (CODEX_ENV_NODE_VERSION, CODEX_ENV_PYTHON_VERSION and others). OpenAI says it “is not an identical environment”, so treat a local pass as necessary, not sufficient. Whether a Codex cloud environment can run Docker was not verified for this page; if yours cannot, run Postgres and the stubs as processes started by bin/env-up instead of containers, and keep the same env file.
Local. Cursor’s Worktrees “let Agent work in isolated Git checkouts”. Run bin/env-up in each one and reference the command from your project rules so the agent starts there.
Cloud. Cursor Cloud Agents “run in isolated VMs in the cloud with full development environments”, and Builds “prepare your Cloud Agent environment in the background”. Since 2026-08-19, subagents can run on their own VMs, each with “an isolated copy of the project”. All three were verified on cursor.com on 2026-08-28; cursor.com was unreachable on 2026-09-26, so the environment configuration file and its keys are left out here. Check Cursor’s cloud agent documentation for where the install and start commands go, and make them call bin/env-up so the definition stays in your repository.
How do you prove an ephemeral environment isolates anything?
Section titled “How do you prove an ephemeral environment isolates anything?”An environment is harness code, so it gets a test like any other code. Run these checks when you introduce it, after changing the scripts, and after upgrading the agent tool.
- The concurrency test. Start two tasks at once and run the full suite in both. Both must pass, each
env-checkmust show different ports and a different Compose project, and a row inserted in one database must not appear in the other. This is the test that fails when anything in the contract is shared. - The wrong-server test. Stop task A’s dev server while task B’s is still up, then run task A’s end-to-end suite. It must fail to connect, not pass against task B.
- The seed-reproducibility test. Run
env-up, dump the database, tear down, repeat, and diff the two dumps. Any difference is non-determinism that will surface later as a flaky test. - The teardown test. After
env-down,docker ps -a --filter label=com.docker.compose.project=task-<slug>returns nothing and the port block directory is gone. - The canary. Break a stub mapping on purpose (return 500 for a call the feature needs) and check that the agent’s end-to-end run goes red. A suite that stays green is not exercising the stub.
Then make every agent run leave the evidence behind: the env-check output, the test summary and the commit it ran against. A Claude Code cloud session can also link its own transcript, because the VM exposes CLAUDE_CODE_REMOTE_SESSION_ID; put that link in the pull request body so a reviewer can open the run that produced the change. A reviewer who sees two concurrent green runs on separate port blocks against a seeded database does not need to re-run the feature by hand. The evidence bundle page defines what each run should attach to its pull request.
Ownership matters as much as the scripts. The tech lead, or the platform team where one exists, owns bin/env-*, compose.agent.yml, the seeds and the stubs. Protect those paths with CODEOWNERS so an agent’s pull request cannot quietly change the environment its own tests run in.
What breaks in ephemeral environments, and how do you recover?
Section titled “What breaks in ephemeral environments, and how do you recover?”- A server moves to the next free port and nobody notices. Vite, for example, falls back silently unless
server.strictPortis set, so the agent’s app runs outside its block. Recovery: strict port options on every server, and base URLs derived from.env.taskin one place. - The end-to-end runner attaches to another task’s server. Playwright’s
webServer.reuseExistingServerattaches to whatever answers on the URL, so the run goes green against code that was never under test. Recovery: derive the runner’s server URL from the task’s port, and run the wrong-server test above. - A fixed
container_nameor a named external volume. The second task fails to start, or both share state. Recovery: let the Compose project name every resource, and grep forcontainer_namein review. - Stacks leak. Crashed and headless runs skip teardown; Claude Code’s
-pruns also leave their worktrees behind. Recovery: runbin/env-sweepfrom step 7 on a schedule. - The cloud session starts without the stack. A setup script that exits non-zero fails the session; one that runs longer than about five minutes is not cached, so every session is slow; a stack started in the setup script is gone because snapshots keep files, not processes. Recovery: pull and install in the setup script, start services in a
SessionStarthook, and move one-off long downloads out of setup. - The agent makes the environment agree with its code. It edits a stub mapping or a seed so a failing test passes. Recovery: CODEOWNERS on
stubs/, seeds andcompose.agent.yml, a prompt rule forbidding those edits, and the practices in protect the oracle. - Stubs drift from the real API. Every run passes against a stub that no longer matches the provider. Recovery: a scheduled CI job that runs the same contract tests against the provider’s own sandbox and fails when the stub and the sandbox disagree.
- A real credential rides along. A token set as a plain environment variable in a cloud environment is readable by anyone who uses that environment, according to Anthropic’s cloud environments guide (checked 2026-09-26). Recovery: test-mode values in
.env.task, and for anything real either a scoped agent identity or Anthropic’s API credentials, which the agent proxy attaches outside the VM (Pro and Max plans only; Team and Enterprise do not have them yet, per the same guide). - The VM runs out of resources. A full stack plus a browser test run can exceed a hosted VM’s ceilings, and the VM may stop the task. Recovery: trim
compose.agent.ymlto what the task needs, or move heavy suites to self-hosted runners or CI.
Where to go next with agent environments
Section titled “Where to go next with agent environments”- Prerequisite: make the codebase agent-ready, because a per-task stack only helps when one command can check the result and the seed is deterministic.
- Permissions and sandboxes sets the approval layer and network egress that sit around this environment.
- Agent sandboxes compared catalogues container-use, E2B, Daytona and microVMs when you outgrow Compose on one machine.
- Parallel agents covers the tools that manage many worktrees and sessions at once.
- Background and cloud agents compared compares the vendors’ hosted runs on triggers, concurrency and cost.
- Next: end-to-end verification by the agent, which turns the stack into proof a reviewer can accept.