Skip to content

The coding-agent landscape beyond the big three

Beyond Claude Code, Codex and Cursor, coding agents in September 2026 fall into three groups: cloud agents that turn an issue into a pull request (GitHub Copilot, Jules), scriptable terminal agents (Gemini CLI, OpenCode, Amp, Factory Droid), and editor or desktop environments (Kiro, Devin Desktop, Cline, Kilo Code, Junie, Warp). Only a bake-off on real closed issues ranks them.

Your CTO asks “should we also buy Copilot’s agent, or move to Kiro?” Vendor pages, star counts and leaderboards cannot answer, because none of them measures your codebase, your gates or your reviewers’ time.

  • A dated map of 14 agents, and a table from your need to the agents worth testing.
  • The site’s one bake-off protocol (also used by the agent shells survey and the frameworks hub): scorecard, runner commands and a decision rule.

Checked on 2026-09-26 against vendor repositories, github/docs and npm. Plan prices are on the pricing page, model prices on the models hub.

AgentVendor, licenceWhere it runsWhat it automatesInstall or entry point
GitHub CopilotGitHub, proprietaryCloud agent on GitHub Actions; desktop app; CLITasks that end in a pull request, @copilot on pull requests, automations; hosts Claude and Codex (public preview)npm install -g @github/copilot (CLI 1.0.88, command copilot)
Gemini CLIGoogle, Apache-2.0Terminal, headless, GitHub ActionScripted tasks with JSON output, @gemini-cli on issuesnpm install -g @google/gemini-cli (0.61.0, command gemini)
Jules and AntigravityGoogleJules: cloud VM; Antigravity: quota in Google AI plans (secondary: CloudZero)Jules opens pull requests (jules-action); Antigravity supports Agent SkillsVendor accounts
KiroAWSIDE, CLI, Web, Mobile (preview)Specs, property-based tests, steering, Agent SkillsVendor download
Devin Desktop and Devin CLICognitionEditor (formerly Windsurf), CLIEditor agent; SWE-2 model since 2026-09-10Vendor download
OpenCodeanomalyco, open sourceTerminal, desktop app (beta)Any model provider; read-only plan agentnpm i -g opencode-ai@latest (1.18.32)
AmpAmp, proprietaryTerminalExecute mode amp -x, subagents, “Oracle” second opinionnpm install -g @ampcode/cli (command amp)
Factory DroidFactory, proprietaryTerminal, headlessUnattended CLI tasksnpm install -g @factory/cli (0.228.0, command droid)
ClineCline, Apache-2.0VS Code, JetBrains, headless CLIYour own model keysnpm install -g cline (3.0.65)
Kilo CodeKilo, MITVS Code, JetBrains, CLI (an OpenCode fork)“500+ models” through one agentnpm install -g @kilocode/cli (7.8.1, command kilo)
JunieJetBrainsJetBrains IDEs, terminal, CI/CD“An LLM-agnostic coding agent”npm install -g @jetbrains/junie (3110.7.0)
WarpWarp, AGPL-3.0 clientDesktop terminalHosts its own agent, Claude Code, Codex, Gemini CLIVendor download

The Devin Desktop rename (June 2026) and SWE-2 date are SECONDARY (creeta, aipricing.guru). More open-source terminal agents (Goose, Crush, Qwen Code, Aider) are in other coding agents.

Popularity, dated. GitHub stars on 2026-09-26: OpenCode 210,085, Gemini CLI 107,166, Cline 69,340, Warp 65,162, Kilo Code 27,419, Copilot CLI 11,211 (a repository without source code, so stars understate use). Stars measure attention, not fitness.

Which agent should you test for which need?

Section titled “Which agent should you test for which need?”
If you need…Test these firstWhy
Issue to pull request with nobody at a laptopCopilot cloud agent, JulesVendor-managed runs that open pull requests
A second agent under an existing Copilot contractCopilot app or CLI, then Claude or Codex inside CopilotSeats you already pay for; AI credits since 2026-06-01
An open-source agent for any model, including self-hostedOpenCode, Cline, Kilo CodeOpen licences, any provider; see open-weight models
A JetBrains-first teamJunie, then Cline, Kilo CodeStays in the team’s IDE
Spec-driven work with property-based testsKiroBuilt into the product
A free terminal agent for scriptsGemini CLI“60 requests/min and 1,000 requests/day with personal Google account” (README, 2026-09-26)
One GUI terminal for the agents you already runWarpHosts your agents side by side; clear the AGPL-3.0 client licence with legal before you modify or redistribute the client

Keep the harness portable: Copilot, Kiro and Antigravity support Agent Skills, Copilot and Gemini CLI speak MCP, and one AGENTS.md can hold your rules for all of them. A candidate with its own instruction format is a lock-in cost; see lock-in and portability.

A bake-off replays closed work through the incumbent and one or two candidates under identical conditions and scores the results with gates, not opinions. The pilot design then tests whether the winner changes delivery.

  1. Write the decision rule first. Name the owner, the incumbent, at most two candidates and the thresholds (table below).

  2. Build the task set. Pick 10 to 20 issues closed in the last 90 days with a merged fix: about 40% bugs, 30% small features, 20% refactors, 10% tests.

  3. Hold out the oracle. Turn each merged fix into acceptance tests outside the agent’s worktree; see protect the oracle.

  4. Freeze the harness. Same commit, AGENTS.md, MCP servers and budget cap, and a fresh worktree or container per run. Use each tool’s default model, plus one shared-model run where possible, to separate harness from model.

  5. Run each task three times per agent. One run measures luck.

  6. Score automatically, then review only what passed. Gates, held-out tests, a review agent; a human reviews only what cleared all three and records the minutes.

  7. Decide and schedule the re-run. Apply the rule, then re-run when a vendor ships a new default model; see the new-model playbook.

Commands checked against the installed CLI or vendor docs on 2026-09-26. Every command below runs its agent fully open, so run it only inside a disposable container.

Terminal window
# Terminal, inside the task's disposable container
claude -p "$(cat task.md)" --permission-mode bypassPermissions \
--output-format json --max-budget-usd 5 > run.json
Terminal window
# Candidates, same task file, same fully open scope
# (gemini, amp: vendor docs; copilot: CLI 1.0.88 --help; 2026-09-26)
gemini -p "$(cat task.md)" --approval-mode yolo --output-format json > run.json
copilot -p "$(cat task.md)" --allow-all-tools --max-ai-credits 500 -s > run.txt
amp -x "$(cat task.md)"

Keep one permission scope across agents, or you measure the permissions, not the agent. For a narrower scope, limit every agent to edits plus the test command: Claude Code --permission-mode acceptEdits --allowedTools "Bash(npm test *)", Codex --sandbox workspace-write, Copilot patterns from copilot help permissions such as --allow-tool='shell(npm:*)' (try that once first: Copilot CLI 1.0.88 calls --allow-all-tools “required for non-interactive mode”).

Claude Code (--max-budget-usd) and Copilot CLI (--max-ai-credits, 1 AI credit = $0.01; CLI 1.0.88) cap spend from the CLI. For Codex, Gemini CLI and Amp (no spend flag; checked codex 0.157.1), enforce the step 4 cap with a timeout (for example timeout 20m) and read cost_usd from the provider’s billing export. Other headless modes: see headless agents in CI.

One row per run; accepted is 1 only when gates, held-out tests and human review all pass.

task_id,agent,model,run,gates_pass,heldout_pass,blocking_review_findings,interventions,wall_minutes,cost_usd,reviewer_minutes,accepted
BUG-4312,incumbent,default,1,1,1,0,0,14,1.80,6,1
MetricDefinitionExample threshold for “adopt the candidate”
Accepted rateaccepted runs ÷ all runsWithin 5 points of the incumbent, or better
Cost per accepted change(usage cost + reviewer minutes × loaded rate) ÷ accepted runs; see economicsAt most 80% of the incumbent’s
Interventionshuman messages or approvalsNo higher than the incumbent’s median
Held-out escapesruns that pass gates but fail held-out testsNot higher than the incumbent’s
Two copy-paste prompts: build the task set and write the held-out tests

PR_NUMBER is the merged pull request that closed the issue.

  • The candidate saw the answer. A task file quoting the fix, or tests in the worktree, inflates scores. Recover: regenerate task.md from the issue only.
  • The result measures the model, not the agent. Recover: add the shared-model run from step 4.
  • The instruction file never loaded. Claude Code, for example, reads AGENTS.md only when there is no CLAUDE.md (v2.1.277 and later, latest channel). Recover: start each run by asking the agent to quote your first project rule.
  • The tool changed under you. Roo Code shut down on 15 May 2026 per its README, and Continue’s README calls it “no longer actively maintained and is read-only”. Recover: prefer portable config (MCP, Agent Skills, AGENTS.md) and re-run before renewals.
  • A cloud agent got more access than the pilot needed. Recover: connect a test repository or fork; see permissions and sandboxing.

Frequently asked questions

Which coding agents exist beyond Claude Code, Codex and Cursor?

As of 2026-09-26: GitHub Copilot (cloud agent, desktop app, CLI), Google's Gemini CLI, Jules and Antigravity, AWS Kiro, Cognition's Devin Desktop (formerly Windsurf) and Devin CLI, and the independents OpenCode, Amp, Factory Droid, Cline, Kilo Code, JetBrains Junie and Warp.

Which one is best?

This page ranks none of them, because no public benchmark measures your repository, your gates and your review cost. Shortlist by where the agent runs and what it automates, then run the bake-off protocol on 10 to 20 of your own closed issues.

How do I compare a new coding agent with the one we already use?

Replay the same closed issues through both agents from the same commit, with the same instruction file and budget, three runs each. Score them with your existing gates and held-out acceptance tests, and compare accepted changes, interventions and cost per accepted change against a decision rule written before the first run.