The coding-agent landscape beyond the big three
Beyond Claude Code, Codex and Cursor, coding agents in September 2026 fall into three groups: cloud agents that turn an issue into a pull request (GitHub Copilot, Jules), scriptable terminal agents (Gemini CLI, OpenCode, Amp, Factory Droid), and editor or desktop environments (Kiro, Devin Desktop, Cline, Kilo Code, Junie, Warp). Only a bake-off on real closed issues ranks them.
Your CTO asks “should we also buy Copilot’s agent, or move to Kiro?” Vendor pages, star counts and leaderboards cannot answer, because none of them measures your codebase, your gates or your reviewers’ time.
What this landscape map gives you
Section titled “What this landscape map gives you”- A dated map of 14 agents, and a table from your need to the agents worth testing.
- The site’s one bake-off protocol (also used by the agent shells survey and the frameworks hub): scorecard, runner commands and a decision rule.
Which coding agents run where?
Section titled “Which coding agents run where?”Checked on 2026-09-26 against vendor repositories, github/docs and npm. Plan prices are on the pricing page, model prices on the models hub.
| Agent | Vendor, licence | Where it runs | What it automates | Install or entry point |
|---|---|---|---|---|
| GitHub Copilot | GitHub, proprietary | Cloud agent on GitHub Actions; desktop app; CLI | Tasks that end in a pull request, @copilot on pull requests, automations; hosts Claude and Codex (public preview) | npm install -g @github/copilot (CLI 1.0.88, command copilot) |
| Gemini CLI | Google, Apache-2.0 | Terminal, headless, GitHub Action | Scripted tasks with JSON output, @gemini-cli on issues | npm install -g @google/gemini-cli (0.61.0, command gemini) |
| Jules and Antigravity | Jules: cloud VM; Antigravity: quota in Google AI plans (secondary: CloudZero) | Jules opens pull requests (jules-action); Antigravity supports Agent Skills | Vendor accounts | |
| Kiro | AWS | IDE, CLI, Web, Mobile (preview) | Specs, property-based tests, steering, Agent Skills | Vendor download |
| Devin Desktop and Devin CLI | Cognition | Editor (formerly Windsurf), CLI | Editor agent; SWE-2 model since 2026-09-10 | Vendor download |
| OpenCode | anomalyco, open source | Terminal, desktop app (beta) | Any model provider; read-only plan agent | npm i -g opencode-ai@latest (1.18.32) |
| Amp | Amp, proprietary | Terminal | Execute mode amp -x, subagents, “Oracle” second opinion | npm install -g @ampcode/cli (command amp) |
| Factory Droid | Factory, proprietary | Terminal, headless | Unattended CLI tasks | npm install -g @factory/cli (0.228.0, command droid) |
| Cline | Cline, Apache-2.0 | VS Code, JetBrains, headless CLI | Your own model keys | npm install -g cline (3.0.65) |
| Kilo Code | Kilo, MIT | VS Code, JetBrains, CLI (an OpenCode fork) | “500+ models” through one agent | npm install -g @kilocode/cli (7.8.1, command kilo) |
| Junie | JetBrains | JetBrains IDEs, terminal, CI/CD | “An LLM-agnostic coding agent” | npm install -g @jetbrains/junie (3110.7.0) |
| Warp | Warp, AGPL-3.0 client | Desktop terminal | Hosts its own agent, Claude Code, Codex, Gemini CLI | Vendor download |
The Devin Desktop rename (June 2026) and SWE-2 date are SECONDARY (creeta, aipricing.guru). More open-source terminal agents (Goose, Crush, Qwen Code, Aider) are in other coding agents.
Popularity, dated. GitHub stars on 2026-09-26: OpenCode 210,085, Gemini CLI 107,166, Cline 69,340, Warp 65,162, Kilo Code 27,419, Copilot CLI 11,211 (a repository without source code, so stars understate use). Stars measure attention, not fitness.
Which agent should you test for which need?
Section titled “Which agent should you test for which need?”| If you need… | Test these first | Why |
|---|---|---|
| Issue to pull request with nobody at a laptop | Copilot cloud agent, Jules | Vendor-managed runs that open pull requests |
| A second agent under an existing Copilot contract | Copilot app or CLI, then Claude or Codex inside Copilot | Seats you already pay for; AI credits since 2026-06-01 |
| An open-source agent for any model, including self-hosted | OpenCode, Cline, Kilo Code | Open licences, any provider; see open-weight models |
| A JetBrains-first team | Junie, then Cline, Kilo Code | Stays in the team’s IDE |
| Spec-driven work with property-based tests | Kiro | Built into the product |
| A free terminal agent for scripts | Gemini CLI | “60 requests/min and 1,000 requests/day with personal Google account” (README, 2026-09-26) |
| One GUI terminal for the agents you already run | Warp | Hosts your agents side by side; clear the AGPL-3.0 client licence with legal before you modify or redistribute the client |
Keep the harness portable: Copilot, Kiro and Antigravity support Agent Skills, Copilot and Gemini CLI speak MCP, and one AGENTS.md can hold your rules for all of them. A candidate with its own instruction format is a lock-in cost; see lock-in and portability.
How do you run a coding-agent bake-off?
Section titled “How do you run a coding-agent bake-off?”A bake-off replays closed work through the incumbent and one or two candidates under identical conditions and scores the results with gates, not opinions. The pilot design then tests whether the winner changes delivery.
-
Write the decision rule first. Name the owner, the incumbent, at most two candidates and the thresholds (table below).
-
Build the task set. Pick 10 to 20 issues closed in the last 90 days with a merged fix: about 40% bugs, 30% small features, 20% refactors, 10% tests.
-
Hold out the oracle. Turn each merged fix into acceptance tests outside the agent’s worktree; see protect the oracle.
-
Freeze the harness. Same commit,
AGENTS.md, MCP servers and budget cap, and a fresh worktree or container per run. Use each tool’s default model, plus one shared-model run where possible, to separate harness from model. -
Run each task three times per agent. One run measures luck.
-
Score automatically, then review only what passed. Gates, held-out tests, a review agent; a human reviews only what cleared all three and records the minutes.
-
Decide and schedule the re-run. Apply the rule, then re-run when a vendor ships a new default model; see the new-model playbook.
Run the same task headless in each agent
Section titled “Run the same task headless in each agent”Commands checked against the installed CLI or vendor docs on 2026-09-26. Every command below runs its agent fully open, so run it only inside a disposable container.
# Terminal, inside the task's disposable containerclaude -p "$(cat task.md)" --permission-mode bypassPermissions \ --output-format json --max-budget-usd 5 > run.json# Only in an externally sandboxed container (codex 0.157.1)codex exec --dangerously-bypass-approvals-and-sandbox --json -o last-message.md "$(cat task.md)" > run.jsonlStart the same task.md as a Cloud Agent or editor agent on a fresh branch and fill the scorecard by hand. The Cursor CLI command name was not verified on 2026-09-26, so it is not scripted here.
# Candidates, same task file, same fully open scope# (gemini, amp: vendor docs; copilot: CLI 1.0.88 --help; 2026-09-26)gemini -p "$(cat task.md)" --approval-mode yolo --output-format json > run.jsoncopilot -p "$(cat task.md)" --allow-all-tools --max-ai-credits 500 -s > run.txtamp -x "$(cat task.md)"Keep one permission scope across agents, or you measure the permissions, not the agent. For a narrower scope, limit every agent to edits plus the test command: Claude Code --permission-mode acceptEdits --allowedTools "Bash(npm test *)", Codex --sandbox workspace-write, Copilot patterns from copilot help permissions such as --allow-tool='shell(npm:*)' (try that once first: Copilot CLI 1.0.88 calls --allow-all-tools “required for non-interactive mode”).
Claude Code (--max-budget-usd) and Copilot CLI (--max-ai-credits, 1 AI credit = $0.01; CLI 1.0.88) cap spend from the CLI. For Codex, Gemini CLI and Amp (no spend flag; checked codex 0.157.1), enforce the step 4 cap with a timeout (for example timeout 20m) and read cost_usd from the provider’s billing export. Other headless modes: see headless agents in CI.
Scorecard and decision rule to adopt
Section titled “Scorecard and decision rule to adopt”One row per run; accepted is 1 only when gates, held-out tests and human review all pass.
task_id,agent,model,run,gates_pass,heldout_pass,blocking_review_findings,interventions,wall_minutes,cost_usd,reviewer_minutes,acceptedBUG-4312,incumbent,default,1,1,1,0,0,14,1.80,6,1| Metric | Definition | Example threshold for “adopt the candidate” |
|---|---|---|
| Accepted rate | accepted runs ÷ all runs | Within 5 points of the incumbent, or better |
| Cost per accepted change | (usage cost + reviewer minutes × loaded rate) ÷ accepted runs; see economics | At most 80% of the incumbent’s |
| Interventions | human messages or approvals | No higher than the incumbent’s median |
| Held-out escapes | runs that pass gates but fail held-out tests | Not higher than the incumbent’s |
Two copy-paste prompts: build the task set and write the held-out tests
PR_NUMBER is the merged pull request that closed the issue.
What breaks in a coding-agent bake-off?
Section titled “What breaks in a coding-agent bake-off?”- The candidate saw the answer. A task file quoting the fix, or tests in the worktree, inflates scores. Recover: regenerate
task.mdfrom the issue only. - The result measures the model, not the agent. Recover: add the shared-model run from step 4.
- The instruction file never loaded. Claude Code, for example, reads
AGENTS.mdonly when there is noCLAUDE.md(v2.1.277 and later,latestchannel). Recover: start each run by asking the agent to quote your first project rule. - The tool changed under you. Roo Code shut down on 15 May 2026 per its README, and Continue’s README calls it “no longer actively maintained and is read-only”. Recover: prefer portable config (MCP, Agent Skills,
AGENTS.md) and re-run before renewals. - A cloud agent got more access than the pilot needed. Recover: connect a test repository or fork; see permissions and sandboxing.
Where to go next from the agent landscape
Section titled “Where to go next from the agent landscape”Frequently asked questions
Which coding agents exist beyond Claude Code, Codex and Cursor?
As of 2026-09-26: GitHub Copilot (cloud agent, desktop app, CLI), Google's Gemini CLI, Jules and Antigravity, AWS Kiro, Cognition's Devin Desktop (formerly Windsurf) and Devin CLI, and the independents OpenCode, Amp, Factory Droid, Cline, Kilo Code, JetBrains Junie and Warp.
Which one is best?
This page ranks none of them, because no public benchmark measures your repository, your gates and your review cost. Shortlist by where the agent runs and what it automates, then run the bake-off protocol on 10 to 20 of your own closed issues.
How do I compare a new coding agent with the one we already use?
Replay the same closed issues through both agents from the same commit, with the same instruction file and budget, three runs each. Score them with your existing gates and held-out acceptance tests, and compare accepted changes, interventions and cost per accepted change against a decision rule written before the first run.