AI model comparison for coding agents
On 2026-09-26, Claude Opus 5.5 is the Claude Code default from v2.1.280 (the latest channel) and GPT-6 Astra is the Codex default. The rule: start on the default, raise effort before switching models, and switch only when evals show a gain: Claude Fable 5.1 for the longest-horizon work; Sonnet 5, GPT-6 Sol, GPT-6 Luna or Haiku 4.5 when cost dominates.
In September 2026 Anthropic shipped Claude Fable 5.1 and Opus 5.5, OpenAI shipped GPT-6 Astra, Sol and Luna, Google shipped Gemini 3.8 Flash and, according to MarkTechPost, SpaceXAI shipped Grok 4.7. Your model picker changed under you, a launch chart says your current model is behind, and someone on the team wants to switch today. This page is for the developer choosing a model for a session and the tech lead or CTO who sets the default and budget for a team. It is also the one page on this site that carries model versions and prices; every other page links here.
What this page gives you
Section titled “What this page gives you”- A job-by-job model recommendation for Claude Code, Codex and Cursor.
- Current prices, model IDs and context limits, each checked on 2026-09-26.
- The exact commands and settings to change model and effort, and to pin them.
- Two copy-paste prompts: one finds stale model IDs in a repository, one runs an effort sweep.
- The failure modes that show up after a model change, with the fix for each.
Which model should you use for each job?
Section titled “Which model should you use for each job?”Anthropic’s own advice is: “If you’re unsure which model to use, start with Claude Opus 5.5 for most workloads.” OpenAI’s Codex catalog makes GPT-6 Astra priority one. The table follows both vendors’ defaults.
| Job | Claude Code | Codex | Why |
|---|---|---|---|
| Hard agentic coding (the default case) | Claude Opus 5.5 at default medium effort, raised to high or xhigh when work does not land | GPT-6 Astra (default effort low) | Each tool’s default; Anthropic’s “start with” model; priority 1 in the Codex catalog |
| Everyday work where cost matters | Claude Sonnet 5 | GPT-6 Sol (“Workhorse model for coding and everyday work”) | Half or less of the flagship price |
| Cheap, high volume: classification, extraction, subagent fan-out | Claude Haiku 4.5 | GPT-6 Luna (“Fast and affordable model for easier tasks”) | Lowest per-token price in each lineup |
| Longest horizon, hardest reasoning | Claude Fable 5.1, chosen by hand (/model fable) | GPT-6 Astra at max or ultra | Anthropic: use Fable 5.1 “when your evals on Claude Opus 5.5 at higher effort still fall short” |
| Self-hosted or air-gapped | See open-weight models for coding agents | Same, via codex --oss | Licence and harness support differ per model |
Cursor. SpaceXAI’s Grok 4.7 was available in Cursor from its release day, according to MarkTechPost (2026-09-21). No Cursor default model, model pool or Cursor-side price could be verified on 2026-09-26, so check the model picker and Cursor’s pricing page before you rely on one.
What do the current models cost?
Section titled “What do the current models cost?”Prices are USD per million tokens (MTok), input / output, standard tier, checked 2026-09-26 on each vendor’s price page unless the row names another source.
| Model | API ID | Price in / out | Context / max output | Status |
|---|---|---|---|---|
| Claude Opus 5.5 | claude-opus-5-5 | $4 / $20 | 1M / 128K | Current default, released 2026-09-22 |
| Claude Fable 5.1 | claude-fable-5-1 | $10 / $50 | 1M / 128K | Current; never a default |
| Claude Sonnet 5 | claude-sonnet-5 | $2 / $10, permanent | 1M / 128K | Current; the sonnet alias |
| Claude Haiku 4.5 | claude-haiku-4-5 | $1 / $5 | 200K / 64K | Current Haiku; no effort parameter |
| Claude Opus 5 | claude-opus-5 | $5 / $25 | 1M / 128K | Legacy, still served |
| Claude Opus 4.8 | claude-opus-4-8 | $5 / $25 | — | Legacy, still served |
Batch processing halves every rate. Cache reads cost 0.1x input, except 0.05x on Opus 5.5 ($0.20) and 0.025x on Fable 5.1 ($0.25). Opus 5.5 fast mode is $8 / $40, a research preview on the Claude API only. Claude 4.7 and later models use a tokenizer that produces about 30% more tokens for the same text, so compare cost per task, not price per token. Bedrock IDs add the anthropic. prefix, for example anthropic.claude-opus-5-5.
| Model | API ID | Price in / out | Context in Codex | Status |
|---|---|---|---|---|
| GPT-6 Astra | gpt-6-astra | $10 / $50 | 272K default, up to 872K | Codex bundled default since 2026-09-04 |
| GPT-6 Sol | gpt-6-sol | $2 / $10 | 272K default, up to 872K | Current, released 2026-09-22 |
| GPT-6 Luna | gpt-6-luna | $0.10 / $0.50 | 272K default, up to 872K | Current, released 2026-09-22 |
| GPT-5.6 Sol, Terra, Luna | gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-luna | not verified | 272K default, up to 872K | Older, still selectable in Codex |
OpenAI’s price page could not be read on 2026-09-26. The GPT-6 list prices come from press reports (CloudZero and Evolink for Astra; The New Stack and VentureBeat for Sol and Luna), and GitHub Copilot’s published billing uses the same rates. The context column comes from the Codex 0.157.1 model catalog (codex debug models --bundled). There is no GPT-6 Terra, and GPT-6 Pro is a ChatGPT product name, not an API model ID.
| Model | Price in / out | Context | Status |
|---|---|---|---|
| Gemini 3.8 Flash | $0.75 / $3.75 through 2026-12-31, then $1.50 / $7.50 | not verified | Current Flash, released 2026-09-02 (secondary: The Register) |
| Gemini 3.1 Pro Preview | $2 / $12 up to 200K input, $4 / $18 above | not verified | Google still lists it as a preview |
| Grok 4.7 | $2 / $6 up to 200K, $4 / $12 above (GitHub Copilot’s billed rate) | 500K (secondary: MarkTechPost) | Newest Grok, released 2026-09-21 (secondary: MarkTechPost) |
Google prices are global-endpoint prices from Vertex AI pricing; regional endpoints cost 10% more. Grok list prices were not verifiable on SpaceXAI’s site, so the row shows the rate GitHub Copilot bills.
Tool subscription prices (Claude Pro and Max, ChatGPT plans, Cursor plans) live on the tool pricing comparison, not here.
How much does one task cost on each model?
Section titled “How much does one task cost on each model?”A task that reads 200K uncached input tokens and writes 20K output tokens costs, at list price and before caching:
| Model | Cost per task |
|---|---|
| Claude Fable 5.1 · GPT-6 Astra | $3.00 |
| Claude Opus 5.5 | $1.20 |
| Claude Sonnet 5 · GPT-6 Sol | $0.60 |
| Claude Haiku 4.5 | $0.30 |
| GPT-6 Luna | $0.03 |
Token counts for the same text differ between vendors, so equal rows do not mean equal cost for the same task.
This is arithmetic, not a measurement. Agents re-read context on every turn, so caching changes real cost more than the list price does, and a cheaper model that needs three attempts costs more than an expensive one that needs one. Anthropic’s Claude Code cost documentation puts typical enterprise spend at “around $13 per developer per active day and $150-250 per developer per month” (checked 2026-09-26). Measure cost per accepted task on your own work before you budget.
How do you change model and effort in each tool?
Section titled “How do you change model and effort in each tool?”From v2.1.280 on the latest release channel, Opus 5.5 is the default on every paid plan, the Anthropic API, Claude Platform on AWS, Bedrock and Google Cloud’s Agent Platform. The stable channel (v2.1.274 on 2026-09-26) had not reached v2.1.280: there, Pro and Team Standard default to Sonnet 5 and the other plans to Opus 5. Microsoft Foundry defaults to Sonnet 4.5 (Sonnet 4.5’s retirement commitment is “not sooner than” 2026-09-29; check the deprecations page before relying on this default). Opus 5.5 runs at medium effort by default; Fable 5.1 and Sonnet 5 run at high, and Haiku 4.5 has no effort setting.
# Terminal: start a session on a chosen model and effortclaude --model opus --effort high
# Try fallbacks in order when the primary is overloadedclaude --model opus --fallback-model sonnetInside a session, /model switches model and /effort sets low, medium, high, xhigh or max, and /effort ultracode runs xhigh with automatic workflow orchestration. The aliases are opus (Opus 5.5), sonnet (Sonnet 5), fable (Fable 5.1) and opusplan (Opus in plan mode, Sonnet for execution). To pin models for a team, use the availableModels, enforceAvailableModels and maxEffortLevel settings in managed settings.
GPT-6 Astra has been the bundled default since CLI 0.153.4 (2026-09-04), with default effort low. The server can override the default for signed-in accounts, so pin the model if you need a fixed one. GPT-6 Sol and Luna default to medium. Codex adds an ultra level (“Maximum reasoning with automatic task delegation”) on Astra and Sol; Luna has no ultra.
# Terminal: one session on GPT-6 Sol at high effortcodex -m gpt-6-sol -c model_reasoning_effort="high"
# List the models this Codex build knows aboutcodex debug models --bundled# ~/.codex/config.toml: pin the team defaultmodel = "gpt-6-astra"model_reasoning_effort = "medium"model_context_window = 400000 # raise from the 272K default, up to 872KAbove 272K input, requests are billed at a higher tier (GitHub Copilot bills GPT-6 Astra at $20 / $75 there, against $10 / $50 below it). Raise the window only for sessions that need it, for example with -c model_context_window=400000 on one run.
Choose the model in the model picker for each chat or agent. Cursor’s model list, defaults and per-model rates were not verifiable on 2026-09-26, so read them in Cursor’s settings and pricing page rather than from this site. The effort and eval rules below apply unchanged.
How do you prove a model change was worth it?
Section titled “How do you prove a model change was worth it?”A model switch is a change to your production system, so it gets the same evidence as a code change. You do not judge it by reading the agent’s diffs; you judge it by what passes your gates.
-
Pin today’s setup. Record the exact model ID (
claude-opus-5-5,gpt-6-astra), the effort level and the tool version, so an alias that moves later does not change behaviour silently. -
Sweep effort before models. Run the same tasks at the next effort level up. Anthropic’s Claude Code docs say “Opus 5.5 at
mediummatches or exceeds Opus 5 athigh”, which is the pattern to expect: effort is the cheaper lever. -
Run your own eval set. Use 20–40 tasks taken from your merged pull requests, each with a test command that decides pass or fail. Building that eval set is covered on the benchmarks page.
-
Compare three numbers. Accepted-task rate (tests, type check and lint pass, and the change survives review), cost per accepted task, and wall-clock time per task.
-
Roll out one workflow (loop) at a time, with a rollback. The tech lead who owns the managed settings signs off, changes the pin for one agent workflow at a time, and reverts the pin if the accepted-task rate drops. The new-model playbook has the full 48-hour procedure.
Replace LEVEL with medium, high or xhigh, MERGED_SHA with the merge commit of a real pull request, ISSUE_TEXT with its original issue text, and the two commands with your own test and type-check commands. Run the prompt in three fresh sessions, one per effort level: /effort medium, high and xhigh in Claude Code, or -c model_reasoning_effort="medium" and so on in Codex. Start each session in its own worktree (claude --worktree or codex --worktree) so the runs cannot see each other’s changes. Read the cost of each run in /usage in Claude Code; in Codex, /status shows estimated credits or cost for eligible workspaces (0.148.0+) and /usage shows account usage (0.156.0+); otherwise compare wall-clock time and pass or fail. Then compare the three reports.
What do the benchmarks say about these models?
Section titled “What do the benchmarks say about these models?”On the official Terminal-Bench 4.0 leaderboard (read 2026-09-26), Claude Fable 5.1 (max effort) running in Claude Code scores 57.9% ± 3.8, Claude Opus 5 (max) in Claude Code 51.8% ± 3.4, and GPT-5.6 Sol (max) in Codex 37.3% ± 3.8. Every one of those entries is a max-effort run, so none of them shows how a model performs at its default effort. Neither Claude Opus 5.5 nor GPT-6 Astra had an entry on the board that day. Anthropic reports 66.4% for Opus 5.5 at xhigh in its 2026-09-22 launch table, and itself warns: “at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.”
Every score measures a model inside one agent harness on someone else’s tasks. Read how to read coding-agent benchmarks before you quote one, and let your own eval set make the decision.
What breaks after a model change?
Section titled “What breaks after a model change?”Where to go next with model choice
Section titled “Where to go next with model choice”- Reading coding-agent benchmarks, and building your own evals: the eval set that makes the decision.
- When a new model ships: a 48-hour evaluation playbook: the rollout and rollback procedure.
- Open-weight and self-hosted models for coding agents: GLM, Qwen, Kimi, DeepSeek and Mistral.
- Model routing by evidence, not brand: routing different jobs to different models.
- Where the model runs: gateways, cloud routes, residency and zero retention: for CTOs choosing a hosting route.
- Claude Code cost control and Codex cost management: keeping spend inside budget.
Sources checked on 2026-09-26
Section titled “Sources checked on 2026-09-26”- Anthropic: models overview, pricing, model deprecations, Claude Opus 5.5 announcement, Claude Code model configuration, and Claude Code costs.
- OpenAI: Codex release 0.153.4, Codex release 0.156.1, the local
codex debug models --bundledcatalog (0.157.1), and GitHub Copilot’s model pricing table. - Google: Vertex AI generative AI pricing.
- SpaceXAI: MarkTechPost on Grok 4.7 (secondary).
- Benchmarks: Terminal-Bench 4.0 leaderboard data.
Frequently asked questions
Which model is the default in Claude Code and in Codex?
Claude Opus 5.5 is the Claude Code default on every paid plan and the Anthropic API from v2.1.280 on the latest release channel; the stable channel still gives older defaults. GPT-6 Astra is the Codex bundled default since CLI 0.153.4 (2026-09-04).
How much do the current Claude models cost per million tokens?
Claude Opus 5.5 is $4 / $20, Claude Fable 5.1 is $10 / $50, Claude Sonnet 5 is $2 / $10 (the planned rise to $3 / $15 was cancelled), and Claude Haiku 4.5 is $1 / $5, input / output, checked 2026-09-26.
Should I raise effort or switch to a bigger model?
Raise effort first. Opus 5.5 runs at medium effort by default in Claude Code and GPT-6 Astra at low in Codex; move to high or xhigh with /effort or model_reasoning_effort, and switch model only when your own evals show the gain.
When is Claude Fable 5.1 worth its price?
For the longest-horizon and hardest reasoning work, chosen by hand with /model fable. It is never a default, costs $10 / $50 per million tokens, and may bill to usage credits depending on plan.