Skip to content

Ephemeral environments: a full stack per agent task

An ephemeral environment for an agent task is a disposable, seeded copy of the whole stack (code checkout, ports, database, stubbed third-party services and test-only credentials) that exists for one run and is destroyed after it. Each agent then proves behaviour end to end without colliding with parallel runs or touching shared data. A worktree isolates only the files.

This page is for the developer who runs several agents at once and the tech lead who wants each of those runs to leave evidence a reviewer can trust. Three agents each got a worktree. The first started the dev server on port 3000, the second’s server silently moved to 3001, and the third’s end-to-end suite attached to the first agent’s server and reported green against code it never tested. Meanwhile two of them ran migrations against the same local Postgres. Every run “passed”, and none of them proved anything.

What you’ll walk away with from a per-task environment

Section titled “What you’ll walk away with from a per-task environment”
  • A six-part environment contract (checkout, port block, data, stubs, credentials, lifetime) that says what each run gets to itself.
  • Three scripts you can adapt: bin/env-up allocates a port block and starts a seeded stack under its own Compose project, bin/env-check proves it is live and isolated, and bin/env-down removes it.
  • The wiring for each tool: Claude Code worktrees and cloud environments, Codex worktrees and cloud tasks, and Cursor worktrees and Cloud Agents.
  • Three copy-paste prompts: build the contract for your repository, run a task inside it, and audit isolation with two concurrent runs.
  • An isolation test that proves the setup works before you trust it with unattended runs, and the failure modes that make a green run meaningless.

A git worktree gives each agent its own files and branch. Everything the code talks to at runtime is still shared: network ports, the database, caches, message queues, Docker container names, third-party sandboxes and any credential in your shell. An agent that can only run unit tests does not care. An agent that must prove a feature works end to end, which is the point of verifying behaviour instead of reading diffs, needs all of it, and needs it to itself.

The environment is the seventh layer of the harness. The permissions and sandbox layer limits what a run may do; the environment decides what the run can see and break. An environment with no production credentials in it is also the strongest permission rule there is: the agent cannot misuse a key that does not exist on the machine.

Which six parts make up the environment contract?

Section titled “Which six parts make up the environment contract?”
PartIsolated byWhat goes wrong without it
CheckoutA worktree locally, a fresh clone in a cloud VMTwo agents edit the same files; a checkout carries another task’s uncommitted changes
Port blockA fixed block of ports per task index, written to a file every tool readsDev servers fall back to the next free port and tests hit a neighbour’s server
DataA per-task database created from migrations and a deterministic seedRuns corrupt each other’s rows; a test passes only because another run left data behind
Service stubsA stub server per task for each third-party API, with committed mappingsTests call a real payment, email or ads API, or fail when its sandbox is down
CredentialsA generated env file with test-mode values onlyA prompt injection or a wrong command reaches production, real customers or real money
LifetimeA teardown command and a time-to-live sweepStacks leak until the machine runs out of disk, memory or ports

A run proves behaviour only when all six hold.

Pick the lightest level that isolates everything the task touches. The tool catalogue, with each product’s isolation model and a worked example, is agent sandboxes compared; this page covers the pattern.

LevelWhat it isolatesUse it whenMain cost
Worktree + port block + Compose project on your machineFiles, ports, containers, databaseYou run two to five agents locally and review their results the same dayYour machine’s CPU and memory; you own teardown
Container per task (for example container-use, which gives each agent a container on its own git branch)The above plus the process tree and installed toolsTasks install packages or need different toolchainsImage build time; Docker on every machine
Vendor cloud environment (Claude Code cloud sessions, Codex cloud, Cursor Cloud Agents)A whole VM per task, on the vendor’s infrastructureBackground runs, runs started from an issue or chat, and runs that must survive your laptop closingSetup-script discipline; vendor resource limits; network allowlists
Self-hosted cloud runnersA VM or container on your infrastructureThe code or data may not leave your networkYou operate the fleet

To try the container-per-task level with Claude Code, install container-use (brew install dagger/tap/container-use) and register it as an MCP server with claude mcp add container-use -- container-use stdio; it is not published on npm, so an npx container-use config does not work.

Whichever level you choose, keep the environment definition in the repository as scripts, not in a vendor settings screen alone. The same bin/env-up then runs on your laptop, in a vendor cloud VM and in CI, so a result from one is comparable with the others.

Build the environment contract for your repository

Section titled “Build the environment contract for your repository”

The example is a Node.js service with Postgres and a third-party payments API. Swap the image names and commands for your stack; keep the structure.

  1. Allocate a port block per task. Each task gets an index from 1 to 49, and every port is a base plus ten times the index. The stride of ten keeps a server that falls back to the next free port inside its own block. Allocation uses mkdir, which is atomic, so two agents starting at the same moment cannot take the same index.

    #!/usr/bin/env bash
    # bin/env-up — start an isolated, seeded stack for one agent task
    set -euo pipefail
    slug=$(printf '%s' "${1:?usage: bin/env-up <task-slug>}" | tr '[:upper:]' '[:lower:]' | tr -cs 'a-z0-9' '-' | sed 's/^-*//;s/-*$//')
    state="${AGENT_ENV_HOME:-$HOME/.agent-envs}"
    mkdir -p "$state"
    idx=""
    for dir in "$state"/*/; do # rerun of the same task: reuse its block
    [ -f "$dir/slug" ] && [ "$(cat "$dir/slug")" = "$slug" ] && idx=$(basename "$dir")
    done
    if [ -z "$idx" ]; then
    for i in $(seq 1 49); do
    if mkdir "$state/$i" 2>/dev/null; then idx=$i; echo "$slug" > "$state/$i/slug"; break; fi
    done
    fi
    [ -n "$idx" ] || { echo "no free port block in $state" >&2; exit 1; }
    cat > .env.task <<EOF
    TASK_SLUG=$slug
    TASK_INDEX=$idx
    COMPOSE_PROJECT_NAME=task-$slug
    APP_PORT=$((3000 + idx * 10))
    DB_PORT=$((5432 + idx * 10))
    STUB_PORT=$((8080 + idx * 10))
    DATABASE_URL=postgres://app:app@127.0.0.1:$((5432 + idx * 10))/app
    PAYMENTS_API_URL=http://127.0.0.1:$((8080 + idx * 10))
    PAYMENTS_API_KEY=test_only_not_a_secret
    EOF
    set -a; . ./.env.task; set +a
    docker compose -p "$COMPOSE_PROJECT_NAME" --env-file .env.task -f compose.agent.yml up -d --wait
    npm run db:migrate
    npm run db:seed
    echo "task $slug: app :$APP_PORT, db :$DB_PORT, stubs :$STUB_PORT"

    Add .env.task to .gitignore. The file is the single source of ports: the dev server, the test runner and the agent all read it, so no tool has to remember a number.

  2. Define the per-task services under a Compose project. The project name (-p task-<slug>) namespaces every container, network and volume, so two tasks never share one. Never set container_name: a fixed name defeats the project namespace and the second task fails to start.

    # compose.agent.yml — services one agent task needs, nothing shared
    services:
    db:
    image: postgres:16
    environment:
    POSTGRES_USER: app
    POSTGRES_PASSWORD: app
    POSTGRES_DB: app
    ports: ["127.0.0.1:${DB_PORT}:5432"]
    tmpfs: /var/lib/postgresql/data # data dies with the container
    healthcheck:
    test: ["CMD-SHELL", "pg_isready -U app"]
    interval: 2s
    retries: 30
    payments-stub:
    image: wiremock/wiremock
    ports: ["127.0.0.1:${STUB_PORT}:8080"]
    volumes: ["./stubs/payments:/home/wiremock"]

    Binding to 127.0.0.1 keeps the stack off your LAN. --wait makes env-up return only when the database health check passes, so the agent’s first migration does not race the container.

  3. Seed deterministic data. db:seed must produce the same rows on every run: a fixed random seed, fixed timestamps and no calls to live services. Seed the cases your acceptance criteria name (an expired subscription, a user with two organisations, a refund in flight), not a generic sample. If seeding takes minutes, seed one template database once and create each task’s database from it with Postgres’s CREATE DATABASE task_db TEMPLATE app_template, which copies the files instead of replaying the seed. Test data management covers seed design in depth.

  4. Stub every third-party API. WireMock’s image reads stub mappings from /home/wiremock/mappings and serves them on port 8080 inside the container. Commit the mappings under stubs/payments/mappings/, including the failure responses (declined card, timeout, 429) your code must handle. Point the application at the stub through the env file (PAYMENTS_API_URL), never through code the agent can edit.

  5. Generate test-only credentials. The env file carries sandbox or dummy values only. Real tokens stay out of the environment entirely; agent identity and secrets covers how to issue scoped credentials when a task genuinely needs one.

  6. Prove the stack is live and isolated. The agent runs this before its first test and pastes the output into its report.

    #!/usr/bin/env bash
    # bin/env-check — prove this task's stack is up, seeded and its own
    set -euo pipefail
    set -a; . ./.env.task; set +a
    docker compose -p "$COMPOSE_PROJECT_NAME" -f compose.agent.yml --env-file .env.task ps --format '{{.Service}} {{.State}}'
    curl -fsS "http://127.0.0.1:$STUB_PORT/__admin/health" > /dev/null && echo "stubs: healthy"
    docker compose -p "$COMPOSE_PROJECT_NAME" -f compose.agent.yml --env-file .env.task exec -T db \
    psql -U app -d app -tAc "select 'seeded users: ' || count(*) from users"
    echo "project=$COMPOSE_PROJECT_NAME app=$APP_PORT db=$DB_PORT stubs=$STUB_PORT"
  7. Tear it down, and sweep what leaks. env-down removes the containers and volumes and frees the index. Run bin/env-sweep on a schedule for stacks whose task ended without teardown, such as a headless run that crashed.

    #!/usr/bin/env bash
    # bin/env-down — destroy this task's stack and release its port block
    set -euo pipefail
    set -a; . ./.env.task; set +a
    docker compose -p "$COMPOSE_PROJECT_NAME" -f compose.agent.yml --env-file .env.task down -v --remove-orphans
    rm -rf "${AGENT_ENV_HOME:-$HOME/.agent-envs}/$TASK_INDEX" .env.task

    The sweep assumes each worktree directory is named after its task slug, as claude --worktree task-142 and bin/env-up task-142 produce. It tears down every stack whose worktree is gone, so remove abandoned worktrees with git worktree remove first.

    #!/usr/bin/env bash
    # bin/env-sweep — remove stacks whose task has no live worktree (run from cron or CI)
    set -euo pipefail
    state="${AGENT_ENV_HOME:-$HOME/.agent-envs}"
    live=$(git worktree list --porcelain | sed -n 's|^worktree ||p')
    for dir in "$state"/*/; do
    [ -f "$dir/slug" ] || continue
    slug=$(cat "$dir/slug")
    grep -q "/$slug\$" <<<"$live" && continue # worktree named after the slug still exists
    echo "sweeping task-$slug"
    docker compose -p "task-$slug" down -v --remove-orphans
    rm -rf "$dir"
    done

Finally, make the dev server and the end-to-end runner read the env file and fail if the port is taken: set Vite’s server.strictPort to true, and derive the dev server port, webServer.url and use.baseURL from APP_PORT, so the suite can only test this task’s server.

Wire the environment into Claude Code, Codex and Cursor

Section titled “Wire the environment into Claude Code, Codex and Cursor”

The scripts are the same everywhere; what differs is who starts them and where the stack runs.

Local. Start each task in its own worktree with claude --worktree task-142 (or -w). Claude Code creates it under .claude/worktrees/. A worktree is a fresh checkout, so ask Claude to run bin/env-up task-142 first, or run it yourself. To copy gitignored files such as .env into every new worktree, list them in a .worktreeinclude file at the project root. Non-interactive -p runs do not clean up their worktrees, so pair them with your sweep.

Cloud sessions (claude.ai/code, claude --cloud "<task>", routines) run in an isolated VM configured by a cloud environment: network access level, environment variables and a setup script. Checked on 2026-09-26 against Anthropic’s cloud environments guide:

  • The VM already has Docker with docker compose, PostgreSQL 16 and Redis 7.0, on Ubuntu 24.04. Anthropic-hosted sessions have approximate ceilings of 4 vCPUs, 16 GB of RAM and 30 GB of disk.
  • The setup script runs as root before Claude Code launches, must exit zero, and is cached as a filesystem snapshot when it finishes in roughly five minutes; the cache rebuilds when you change the script or allowed hosts, or after roughly seven days. Running processes are not cached, so pull images in the setup script and start the stack per session.
  • The default network level, Trusted, allows package registries, GitHub and Docker Hub. None makes installs and image pulls fail.

Put image pulls in the environment’s Setup script field:

#!/bin/bash
# cloud environment setup script: cached, so pull here, start later
docker pull postgres:16 || true
docker pull wiremock/wiremock || true

Then start the stack on every cloud session with a SessionStart hook in the repository’s .claude/settings.json. The hook runs locally too, so the script exits unless CLAUDE_CODE_REMOTE is true:

{
"hooks": {
"SessionStart": [
{
"matcher": "startup|resume",
"hooks": [
{ "type": "command", "command": "bash \"$CLAUDE_PROJECT_DIR\"/scripts/cloud-env-up.sh", "timeout": 300 }
]
}
]
}
}
#!/bin/bash
# scripts/cloud-env-up.sh — one task per VM, so one fixed slug is enough
[ "$CLAUDE_CODE_REMOTE" = "true" ] || exit 0
cd "$CLAUDE_PROJECT_DIR" && bin/env-up cloud

A session with several repositories does not load hooks from any repository’s .claude/settings.json, so in that case put “run bin/env-up cloud first” at the top of the task prompt instead. To run cloud sessions on your own compute, Team and Enterprise can use self-hosted environments (public beta) and start a session on one with --environment <id>. Bring a finished cloud session to your terminal with claude --teleport.

How do you prove an ephemeral environment isolates anything?

Section titled “How do you prove an ephemeral environment isolates anything?”

An environment is harness code, so it gets a test like any other code. Run these checks when you introduce it, after changing the scripts, and after upgrading the agent tool.

  • The concurrency test. Start two tasks at once and run the full suite in both. Both must pass, each env-check must show different ports and a different Compose project, and a row inserted in one database must not appear in the other. This is the test that fails when anything in the contract is shared.
  • The wrong-server test. Stop task A’s dev server while task B’s is still up, then run task A’s end-to-end suite. It must fail to connect, not pass against task B.
  • The seed-reproducibility test. Run env-up, dump the database, tear down, repeat, and diff the two dumps. Any difference is non-determinism that will surface later as a flaky test.
  • The teardown test. After env-down, docker ps -a --filter label=com.docker.compose.project=task-<slug> returns nothing and the port block directory is gone.
  • The canary. Break a stub mapping on purpose (return 500 for a call the feature needs) and check that the agent’s end-to-end run goes red. A suite that stays green is not exercising the stub.

Then make every agent run leave the evidence behind: the env-check output, the test summary and the commit it ran against. A Claude Code cloud session can also link its own transcript, because the VM exposes CLAUDE_CODE_REMOTE_SESSION_ID; put that link in the pull request body so a reviewer can open the run that produced the change. A reviewer who sees two concurrent green runs on separate port blocks against a seeded database does not need to re-run the feature by hand. The evidence bundle page defines what each run should attach to its pull request.

Ownership matters as much as the scripts. The tech lead, or the platform team where one exists, owns bin/env-*, compose.agent.yml, the seeds and the stubs. Protect those paths with CODEOWNERS so an agent’s pull request cannot quietly change the environment its own tests run in.

What breaks in ephemeral environments, and how do you recover?

Section titled “What breaks in ephemeral environments, and how do you recover?”
  • A server moves to the next free port and nobody notices. Vite, for example, falls back silently unless server.strictPort is set, so the agent’s app runs outside its block. Recovery: strict port options on every server, and base URLs derived from .env.task in one place.
  • The end-to-end runner attaches to another task’s server. Playwright’s webServer.reuseExistingServer attaches to whatever answers on the URL, so the run goes green against code that was never under test. Recovery: derive the runner’s server URL from the task’s port, and run the wrong-server test above.
  • A fixed container_name or a named external volume. The second task fails to start, or both share state. Recovery: let the Compose project name every resource, and grep for container_name in review.
  • Stacks leak. Crashed and headless runs skip teardown; Claude Code’s -p runs also leave their worktrees behind. Recovery: run bin/env-sweep from step 7 on a schedule.
  • The cloud session starts without the stack. A setup script that exits non-zero fails the session; one that runs longer than about five minutes is not cached, so every session is slow; a stack started in the setup script is gone because snapshots keep files, not processes. Recovery: pull and install in the setup script, start services in a SessionStart hook, and move one-off long downloads out of setup.
  • The agent makes the environment agree with its code. It edits a stub mapping or a seed so a failing test passes. Recovery: CODEOWNERS on stubs/, seeds and compose.agent.yml, a prompt rule forbidding those edits, and the practices in protect the oracle.
  • Stubs drift from the real API. Every run passes against a stub that no longer matches the provider. Recovery: a scheduled CI job that runs the same contract tests against the provider’s own sandbox and fails when the stub and the sandbox disagree.
  • A real credential rides along. A token set as a plain environment variable in a cloud environment is readable by anyone who uses that environment, according to Anthropic’s cloud environments guide (checked 2026-09-26). Recovery: test-mode values in .env.task, and for anything real either a scoped agent identity or Anthropic’s API credentials, which the agent proxy attaches outside the VM (Pro and Max plans only; Team and Enterprise do not have them yet, per the same guide).
  • The VM runs out of resources. A full stack plus a browser test run can exceed a hosted VM’s ceilings, and the VM may stop the task. Recovery: trim compose.agent.yml to what the task needs, or move heavy suites to self-hosted runners or CI.