Background and cloud agents compared
Background and cloud agents are coding agents that run on remote machines and report back through a branch, pull request or session link. Claude Code, Codex, Cursor, GitHub Copilot and Jules all offer one, and they differ most in what starts a run, what the run can reach, how long it may last and whose bill it lands on.
Your team uses three agents, and each vendor now sells “an agent that works while you sleep”. One developer fires tasks at Claude Code on the web, another schedules Cursor Automations, and a third assigns issues to Copilot. Nobody can say which of them may push to which branch, which ones see production secrets, or why last month’s bill doubled. The vendor pages each describe one product; this page puts them in one frame so you can choose deliberately and verify what comes back.
It is written for developers who delegate tasks, tech leads who standardise how a team does it, and CTOs who approve the tools. It assumes you already know how to give an agent a bounded task; if not, start with Level 4 of the autonomy ladder.
What this comparison gives you
Section titled “What this comparison gives you”- One table per decision axis: environment, triggers, concurrency and limits, permissions, cost.
- The exact command or API call that starts a background run in each tool, all checked on 2026-09-26.
- A job-to-tool decision table for the six background jobs teams actually run.
- Three copy-paste task briefs that work in every one of these agents.
- A ten-question adoption checklist for a tech lead or CTO.
- The failure modes specific to remote runs and how to recover from each.
Which background agent fits which job?
Section titled “Which background agent fits which job?”Start from the job, not the vendor. Most teams need two of these patterns, rarely more.
| Job | Reach for | Why this one |
|---|---|---|
| Hand off a spec’d task and keep working | Claude Code claude --cloud, Codex codex cloud exec, Copilot (assign the issue) | One command or one click, and the result comes back as a branch you review |
| Get several independent attempts at a hard fix | Codex codex cloud exec --attempts 4 | Best-of-N is a first-class flag (1 to 4 attempts), and you apply the attempt you prefer |
| Recurring chores with no laptop involved | Claude Code routines, Cursor Automations, Copilot automations | All three run on a schedule in the vendor’s cloud |
| Turn an alert into a draft fix | Claude Code routine with an API trigger, Cursor Automations (Sentry, PagerDuty and webhook triggers) | The alert system calls the agent directly |
| Babysit a pull request’s CI and review comments | Claude Code auto-fix (/autofix-pr), @copilot on the pull request | The agent reacts to failed checks and comments until the PR is green |
| Keep code and tool execution inside your network | Claude Code self-hosted environments, Copilot on self-hosted Actions runners, Cursor private workers | The only three with a documented self-hosted path |
Codex automations are the exception to “background means cloud”: the desktop app runs them locally, against a project folder or a worktree, so they need the app open and the machine awake. They suit personal checks that need local state; Scheduled Codex automations covers them.
Where does each background agent run?
Section titled “Where does each background agent run?”The environment decides what the agent can build, test and reach. Everything you configured only on your laptop is missing unless you commit it or put it in the environment definition.
| Claude Code cloud session | Codex cloud task | Cursor Cloud Agent | Copilot cloud agent | Jules | |
|---|---|---|---|---|---|
| Machine | Isolated Anthropic-managed VM: Ubuntu 24.04, x86_64, about 4 vCPUs, 16 GB RAM, 30 GB disk | Container from OpenAI’s universal image or a cached container (2026-08-28) | “Isolated VMs in the cloud with full development environments” (2026-08-28) | Ephemeral GitHub Actions environment, upgradeable to larger runners | A cloud VM |
| Defined by | A cloud environment: network level, variables, setup script, cached between sessions | An environment per repository: setup script, optional maintenance script, variables, setup-only secrets | Environment plus Builds, which “prepare your Cloud Agent environment in the background” | .github/workflows/copilot-setup-steps.yml with one copilot-setup-steps job | The GitHub repository and starting branch |
| Repository config it reads | CLAUDE.md, .claude/ skills, agents, rules; project .mcp.json and settings in single-repo sessions | AGENTS.md | Not verifiable on 2026-09-26 | Custom instructions, skills, repository MCP settings (GitHub and Playwright MCP on by default) | The prompt, plus optional last commit and log in the action |
| Self-hosted option | Self-hosted environments, Team and Enterprise, public beta (--environment ccpool_…) | None found in the sources read | SDK env.type of pool or machine; Cloud Agents API /v0/private-workers | Self-hosted Actions runners via runs-on in the setup file | None found |
Devin, Cognition’s agent, is not in these tables: its product pages were unreachable on 2026-09-26, and only the existence of the Devin CLI repository could be checked. Put it through the adoption checklist below before you adopt it.
Two differences matter most in practice. As documented on 2026-09-26, Claude Code does not carry your user-level setup into the cloud: ~/.claude/CLAUDE.md, user skills, and MCP servers added at local or user scope stay on your machine, and plugins enabled in the repository’s settings are not installed. Copilot’s setup file runs only when it is on the default branch, so a branch that adds it does not change the agent’s environment until it merges.
What can start a background run?
Section titled “What can start a background run?”Triggers decide who, or what, can put an agent to work, and so they are also your prompt-injection surface.
| Claude Code | Codex | Cursor | Copilot | Jules | |
|---|---|---|---|---|---|
| By hand | claude.ai/code, the mobile app, the Desktop app (Cloud), claude --cloud | Codex web, codex cloud exec | Agents window, Cloud Agents API, @cursor/sdk | Agents panel on GitHub, assigning an issue, VS Code, @copilot on a pull request | Web app, REST API (v1alpha) |
| Schedule | Routines: hourly, daily, weekdays, weekly, cron via /schedule update, one-offs; minimum interval one hour | Desktop-app automations (local) | Automations: Scheduled | Automations: hourly, daily or weekly | Your CI scheduler calling the action |
| Events | Routines: GitHub pull request and release events with filters; API /fire endpoint | GitHub, GitLab, Linear and Slack start points (2026-08-28) | Automations: source control, Slack, webhook, Linear, Sentry, PagerDuty (2026-08-28) | Automations: issue created, pull request opened or synchronized, with search and path filters | Any GitHub Actions event via google-labs-code/jules-invoke |
| Chat tools | Claude Tag in Slack (Team and Enterprise) | Slack @Codex, Linear (2026-08-28) | Slack, Linear | Slack, Microsoft Teams, Jira, Linear, Azure Boards | — |
How many can run at once, and for how long?
Section titled “How many can run at once, and for how long?”None of the sources read on 2026-09-26 publishes a hard concurrency number for interactive cloud sessions. The limits you meet are usage limits and per-run caps.
| Parallelism | Per-run limits | Other caps | |
|---|---|---|---|
| Claude Code | Each claude --cloud call is an independent session; Projects (public beta, Pro and Max) coordinate many from one conversation | Idle sessions expire and the VM is reclaimed; background subagents and shell commands are not restored on reopen | Routines: per-account daily run cap; GitHub events have per-routine and per-account hourly caps, and extra events are dropped |
| Codex | Parallel tasks; --attempts 1..4 per task | Not published in the sources read | Plan allowance on 5-hour and weekly windows |
| Cursor | Not documented in the sources read; subagents can run on their own VMs (changelog, 2026-08-19) | Not verifiable on 2026-09-26 | Not verifiable on 2026-09-26 |
| Copilot | One branch and one pull request per task; one repository per run | 59 minutes per session, hard limit | Automations are scoped to one repository |
| Jules | One session per API call | Not published in the sources read | — |
Copilot’s 59-minute ceiling is the one that changes how you write tasks: anything that needs a long build plus a long test run must be split. The same discipline helps elsewhere, because a short task is also a short diff to verify.
What can a background agent touch?
Section titled “What can a background agent touch?”This is the axis a security review asks about first. Every tool isolates the machine; they differ in network egress, credentials and whether anyone approves actions mid-run.
| Network by default | Credentials | Approvals during the run | Where results land | |
|---|---|---|---|---|
| Claude Code cloud session | Trusted: allowlisted registries, GitHub, cloud SDKs; also None, Full or Custom | GitHub credentials stay outside the VM behind a proxy; API credentials on Pro and Max stay outside the sandbox too | You pick a permission mode per session | A branch; you create the pull request |
| Claude Code routine | Same environment levels | Acts as you: commits, pull requests and connector actions carry your identity, and the routine belongs to your personal account, not the team | None: routines have no permission-mode picker and run without stopping for approval | claude/-prefixed branches; pushes to protected branches are rejected |
| Codex cloud task | Setup script has internet; agent phase is off unless you allow it (2026-08-28) | Secrets exist only during the setup script (2026-08-28) | None | A diff you apply locally (codex cloud apply) or a pull request |
| Cursor Cloud Agent | Not verifiable on 2026-09-26 | Per-session envVars, encrypted at rest and deleted with the agent (SDK 1.0.32) | None | A branch, or a pull request with autoCreatePR; opened as the Cursor GitHub App by default for service-account keys |
| Copilot cloud agent | Firewall on, with a recommended allowlist for dependencies | Repository secrets and variables for Copilot | Tools are chosen per automation | A branch and one pull request; its workflows wait until a user with write access approves them |
| Jules | Not published in the sources read | Jules API key | The action sends "requirePlanApproval": false | A pull request ("automationMode": "AUTO_CREATE_PR") |
Two defaults deserve a deliberate decision rather than acceptance. Claude Code routines include all of your connected MCP connectors unless you remove them, and Claude may use every tool of an included connector, writes included, without asking. Cursor’s automation docs, as excerpted on 2026-09-26, describe a memory tool that carries notes between runs. That is useful for triage and also a place where injected text persists, so if your automations have it, treat what it stores as untrusted input.
The broader comparison of sandboxes and approval modes, including local runs, is in permissions, sandboxes and approval modes.
What does a background run cost?
Section titled “What does a background run cost?”Each vendor bills a background run through a different meter, which is why the bill surprises people.
| What you pay with | What to watch | |
|---|---|---|
| Claude Code | Subscription usage shared with all your other Claude use; “no separate compute charge for the cloud VM” | Parallel sessions consume limits proportionately; routines add a daily run cap, then metered overage if usage credits are on |
| Codex | Your ChatGPT plan’s Codex allowance on 5-hour and weekly windows | --attempts 4 is up to four runs of work for one task |
| Cursor | Cloud Agent usage on the account that owns the API key | Tag runs with SDK metadata so you can join them to usage exports |
| Copilot | GitHub Actions minutes plus GitHub AI Credits ($0.01 each), billed to the user who created the automation | An automation on “pull request synchronized” runs on every push |
| Jules | Your Jules plan (pricing not verifiable on 2026-09-26) | Scheduled workflows run even when nothing changed |
Current model prices are on the models hub. The number worth tracking is not cost per run but cost per merged pull request, because a cheap run that produces a discarded branch is pure waste. Cost visibility for agent work shows how to build that view.
How do you start the same task in each tool?
Section titled “How do you start the same task in each tool?”The task below is the same in every tab: fix one flaky test and prove the fix. Only the launcher changes.
A cloud session clones your GitHub remote at your current branch, not your working tree, so push first. Run these in your terminal:
git push origin HEADclaude --cloud "Fix the flaky test in tests/checkout.spec.ts. Done means: npx vitest run tests/checkout.spec.ts passes 5 times in a row, and npm run typecheck passes. Report the command output as evidence."The command prints a session link. To steer the running session from any logged-in machine, queue a message; to continue locally, teleport it back:
claude -p "Also run npm run lint and include its output" --cloud SESSION_IDclaude --teleport SESSION_IDSESSION_ID is the session_… ID or the claude.ai/code URL. --remote is the older, deprecated spelling of --cloud (checked on CLI 2.1.283), and --environment ccpool_… sends the session to your organization’s self-hosted environment instead of Anthropic’s VMs. When the pull request exists, /autofix-pr on its branch starts a cloud session that answers CI failures and review comments. --cloud needs a claude.ai login on a Pro, Max or Team plan, or an Enterprise premium or Chat + Claude Code seat: it does not work with Amazon Bedrock, Google Cloud’s Agent Platform or other third-party providers, and organizations with Zero Data Retention cannot use cloud sessions at all.
codex cloud exec needs an environment ID; run codex cloud to browse your environments and tasks. Then submit two attempts and pick the better one (CLI 0.157.1, marked experimental):
codex cloud exec --env ENV_ID --attempts 2 \ "Fix the flaky test in tests/checkout.spec.ts. Done means: npx vitest run tests/checkout.spec.ts passes 5 times in a row, and npm run typecheck passes. Report the command output as evidence."codex cloud list --env ENV_IDcodex cloud diff TASK_ID --attempt 2codex cloud apply TASK_ID --attempt 2ENV_ID comes from codex cloud, and TASK_ID from codex cloud list. apply writes the chosen attempt’s diff into your local checkout, where your own gates run before you push. --branch runs the task on a branch other than your current one.
Cursor’s cloud agents are started from the Agents window, from Automations, or from code. This Node script uses @cursor/sdk 1.0.32 and opens the pull request itself:
import { Agent } from '@cursor/sdk';
const agent = await Agent.create({ apiKey: process.env.CURSOR_API_KEY, cloud: { repos: [{ url: 'https://github.com/acme/webapp', startingRef: 'main' }], autoCreatePR: true, metadata: { task: 'flaky-checkout-test' }, },});const run = await agent.send( 'Fix the flaky test in tests/checkout.spec.ts. Done means: npx vitest run tests/checkout.spec.ts passes 5 times in a row, and npm run typecheck passes. Report the command output as evidence.',);const result = await run.wait();console.log(result.status, result.git?.branches?.[0]?.prUrl ?? 'no PR');Run it with CURSOR_API_KEY set in the environment. The SDK exposes no tool allowlist for cloud agents, so containment is Cursor’s VM plus your repository’s protections. Driving agents from code covers the SDK in depth.
Put the brief in an issue and assign it to Copilot, or mention @copilot in a pull request comment. To dispatch Copilot from Claude Code, Codex or another agent, add the GitHub MCP server’s copilot toolset, which provides assign_copilot_to_issue and create_pull_request_with_copilot:
# Claude Code (terminal). Single quotes store the ${GITHUB_PAT} reference,# not the token; Claude Code expands it from your environment at load time.claude mcp add --transport http github-copilot https://api.githubcopilot.com/mcp/x/copilot \ -H 'Authorization: Bearer ${GITHUB_PAT}'
# Codex (terminal)codex mcp add github-copilot --url https://api.githubcopilot.com/mcp/x/copilot \ --bearer-token-env-var GITHUB_PATThen ask your local agent: “Assign Copilot to issue 482 in acme/webapp, with custom instructions that the fix must pass npx vitest run tests/checkout.spec.ts five times in a row.” Copilot works in its Actions environment, and its pull request’s workflows run after a user with write access approves them.
The official google-labs-code/jules-invoke action sends this request; you can send it from any script. Set JULES_API_KEY in the environment first:
curl 'https://jules.googleapis.com/v1alpha/sessions' \ -X POST \ -H "Content-Type: application/json" \ -H "X-Goog-Api-Key: $JULES_API_KEY" \ -d '{ "prompt": "Fix the flaky test in tests/checkout.spec.ts. Done means: npx vitest run tests/checkout.spec.ts passes 5 times in a row, and npm run typecheck passes.", "sourceContext": { "source": "sources/github/acme/webapp", "githubRepoContext": { "startingBranch": "main" } }, "requirePlanApproval": false, "automationMode": "AUTO_CREATE_PR" }'The API is v1alpha, so pin the payload in one script and expect changes.
How do you prove a background agent’s output is good?
Section titled “How do you prove a background agent’s output is good?”A background run removes you from the loop while it works, so the proof has to come from checks the agent did not write and cannot change. The agent’s own report (“all tests pass”) is a claim, not evidence.
-
Require a pull request for every background result. A ruleset on the default branch requires one approving review and your CI checks. Routines already push only to
claude/branches; for the other tools the ruleset does the same job. -
Re-run the gates in CI, on the pull request. Tests, type check, lint and any fitness functions run in your pipeline, not in the agent’s VM. The agent may run them too, but only CI’s result counts.
-
Protect the oracle. Route test files, gate scripts, lockfiles and
.github/throughCODEOWNERS, so an agent that “fixes” a test by weakening it needs a named human’s approval. Protecting the oracle lists the paths. -
Attach an evidence bundle. CI posts what ran, what passed and what changed in risky paths. The reviewer reads the bundle and the diff of protected paths, not every line. See the evidence bundle.
-
Name the person who signs off. The owner of the task, or of the automation, approves the merge. On Copilot, pull requests from an automation are attributed to its creator, who cannot approve them, so a second person must. Adopt the same two-person rule for the other tools.
Read the run, not its status light. A green status in the Claude Code routine run list means only that the session started and exited without an infrastructure error, not that the task succeeded; blocked network requests and task failures show up in the transcript. The same holds for every tool here: the pull request’s CI result is the signal, the run badge is not.
Copilot adds one manual checkpoint: its pull requests’ workflows wait for approval from someone with write access. Before you click it, read the pull request’s changes to .github/ and to test files, because approving runs whatever the agent put there.
Cursor, Copilot and Claude Code also sell AI review on the resulting pull request. Treat that as a second opinion that finds issues, not as the approval; reviewing an agent’s pull request explains how to combine it with human sign-off.
Copy-paste task briefs for background agents
Section titled “Copy-paste task briefs for background agents”A remote agent cannot ask you a clarifying question at 2 a.m. and get an answer, so the brief carries the stop conditions. These three work unchanged in claude --cloud, codex cloud exec, a Cursor cloud agent, a Copilot issue body and a Jules session.
The second prompt names the routine-fire-payload block on purpose. Claude Code wraps the text of a /fire call in that block and labels it untrusted, so a routine that does not reference it treats the alert as inert context and does nothing useful.
Adoption checklist for a tech lead or CTO
Section titled “Adoption checklist for a tech lead or CTO”Answer these ten questions before a background agent gets write access to a shared repository. Keep the answers in the repository’s AGENTS.md or your agent policy, so the next person does not have to rediscover them.
| # | Question | A good answer looks like |
|---|---|---|
| 1 | Which jobs may run in the background, and which may not? | A list of change classes: yes for tests, lint fixes, dependency bumps; no for auth, billing and migrations |
| 2 | What starts a run, and can an outsider trigger it? | Maintainer-applied labels or schedules; event triggers filtered to authors with write access |
| 3 | Whose identity does the run act as? | A named owner per automation; service identities where the tool supports them |
| 4 | What network egress does the environment allow? | The narrowest level that builds: Trusted or Custom, never Full by default |
| 5 | Which secrets can the agent read at run time? | None beyond test credentials; production secrets never |
| 6 | Which connectors or MCP servers does each automation carry? | Only those its prompt needs, reviewed when the prompt changes |
| 7 | What gates must pass before merge, and who can change them? | CI on the pull request; gate files owned in CODEOWNERS |
| 8 | Who signs off, and is it someone other than the automation’s owner? | Two-person rule for every background pull request |
| 9 | What is the spend ceiling, and who sees it first? | Monthly budget per team, cost per merged pull request reviewed monthly |
| 10 | How is a runaway automation stopped? | Owner and admin both know the off switch: the Routines toggle, the Copilot organization policy, in_app_local_automation in Codex requirements.toml, and whoever administers your Cursor team |
Question 10 has a concrete answer in each tool. Team and Enterprise owners can turn off routines for everyone at Admin settings > Claude Code. Copilot automations need both the cloud agent and automations allowed by the organization. Codex’s desktop automations are gated by the in_app_local_automation requirement, which admins set in requirements.toml.
To switch them off, put this in the admin-managed requirements.toml:
[features]in_app_local_automation = falseThe Codex source marks this key a requirements-only gate, meant to be set from requirements rather than from a user’s config.toml (checked on Codex 0.157.1).
What breaks when agents run in the background?
Section titled “What breaks when agents run in the background?”The cloud session cannot find a tool or config you use every day. Your user-level CLAUDE.md, user skills and locally scoped MCP servers do not travel to the VM. Recovery: commit what the agent needs to the repository (.claude/, .mcp.json via claude mcp add --scope project), and install toolchains in the environment’s setup script, which the environment cache keeps.
claude --cloud refuses to start. The message names the cause: Cloud sessions are disabled by your organization's policy means an Owner has not turned on allow_remote_sessions; a message naming Amazon Bedrock or another provider means Claude Code is configured for a third-party provider. Recovery: ask an Owner to enable cloud sessions at Admin settings > Claude Code, or unset the provider variables and run claude auth login with a claude.ai account.
Every Anthropic-hosted session fails with an authentication error. Your organization has IP allowlisting on, and cloud sessions, routines and Code Review call the API from Anthropic’s network. Recovery: ask Anthropic support to exempt Anthropic-hosted services, or route those runs to a self-hosted environment.
A Copilot session stops mid-task. It hit the 59-minute hard limit. Recovery: split the task so build, test and fix fit in one session, and warm dependencies in copilot-setup-steps.yml rather than in the task.
Copilot cannot open or update its pull request. A ruleset or branch protection rule that it cannot satisfy, such as a commit-author restriction, blocks it. Recovery: add Copilot as a bypass actor on that ruleset, scoped to the branch pattern it uses.
A routine ignores the alert text you sent it. Fire text arrives wrapped as untrusted data. Recovery: reference the routine-fire-payload block explicitly in the saved prompt, as the alert-triage prompt above does.
GitHub-triggered routines silently skip events. Events beyond the per-routine or per-account hourly cap are dropped during the research preview. Recovery: filter on labels or draft status so only the events you need fire, and use a schedule for sweeps.
Auto-fix replies trigger your deployment bot. Claude posts review replies under your GitHub account, which can fire issue_comment automation such as Atlantis. Recovery: turn auto-fix off on repositories where a comment can deploy or run privileged operations.
A reopened session lost its background work. The VM was reclaimed after inactivity, and background subagents and shell commands are not restored. Recovery: have long jobs write progress to a committed file or pushed branch, and restart the job from there.
The bill doubled. Parallel sessions, best-of-N attempts and event-triggered automations multiply quietly. Recovery: set a per-team ceiling, tag runs (Cursor metadata, routine names, automation owners), and review cost per merged pull request monthly in agent observability.
Where to go next with background agents
Section titled “Where to go next with background agents”For many agents working on one body of work, continue with multi-agent orchestration patterns. For who should own the identities these runs act as, see agent identity and secrets.
Frequently asked questions
Which background coding agents run when my laptop is closed?
Claude Code cloud sessions and routines, Codex cloud tasks, Cursor Cloud Agents and Automations, the GitHub Copilot cloud agent and its automations, and Jules all run on vendor infrastructure. Codex automations in the desktop app run locally and need the app running and the machine awake.
Which cloud agent has a hard time limit?
The GitHub Copilot cloud agent has a documented hard limit of 59 minutes per session that cannot be extended. Claude Code routines have a minimum schedule interval of one hour and a per-account daily run cap. Checked on 26 September 2026.
How is a Claude Code cloud session billed?
Cloud sessions share rate limits with all other Claude and Claude Code usage on the account, and Anthropic states there is no separate compute charge for the cloud VM. Routines also count against a daily run cap, with metered overage when usage credits are on.
Can I run a background agent on my own infrastructure?
Yes, in three of them: Claude Code self-hosted environments (Team and Enterprise, public beta), the Copilot cloud agent on self-hosted GitHub Actions runners, and Cursor cloud agents on private workers or pools. For the others, code and prompts run on the vendor's machines.