Skip to content

AI model comparison for coding agents

On 2026-09-26, Claude Opus 5.5 is the Claude Code default from v2.1.280 (the latest channel) and GPT-6 Astra is the Codex default. The rule: start on the default, raise effort before switching models, and switch only when evals show a gain: Claude Fable 5.1 for the longest-horizon work; Sonnet 5, GPT-6 Sol, GPT-6 Luna or Haiku 4.5 when cost dominates.

In September 2026 Anthropic shipped Claude Fable 5.1 and Opus 5.5, OpenAI shipped GPT-6 Astra, Sol and Luna, Google shipped Gemini 3.8 Flash and, according to MarkTechPost, SpaceXAI shipped Grok 4.7. Your model picker changed under you, a launch chart says your current model is behind, and someone on the team wants to switch today. This page is for the developer choosing a model for a session and the tech lead or CTO who sets the default and budget for a team. It is also the one page on this site that carries model versions and prices; every other page links here.

  • A job-by-job model recommendation for Claude Code, Codex and Cursor.
  • Current prices, model IDs and context limits, each checked on 2026-09-26.
  • The exact commands and settings to change model and effort, and to pin them.
  • Two copy-paste prompts: one finds stale model IDs in a repository, one runs an effort sweep.
  • The failure modes that show up after a model change, with the fix for each.

Anthropic’s own advice is: “If you’re unsure which model to use, start with Claude Opus 5.5 for most workloads.” OpenAI’s Codex catalog makes GPT-6 Astra priority one. The table follows both vendors’ defaults.

JobClaude CodeCodexWhy
Hard agentic coding (the default case)Claude Opus 5.5 at default medium effort, raised to high or xhigh when work does not landGPT-6 Astra (default effort low)Each tool’s default; Anthropic’s “start with” model; priority 1 in the Codex catalog
Everyday work where cost mattersClaude Sonnet 5GPT-6 Sol (“Workhorse model for coding and everyday work”)Half or less of the flagship price
Cheap, high volume: classification, extraction, subagent fan-outClaude Haiku 4.5GPT-6 Luna (“Fast and affordable model for easier tasks”)Lowest per-token price in each lineup
Longest horizon, hardest reasoningClaude Fable 5.1, chosen by hand (/model fable)GPT-6 Astra at max or ultraAnthropic: use Fable 5.1 “when your evals on Claude Opus 5.5 at higher effort still fall short”
Self-hosted or air-gappedSee open-weight models for coding agentsSame, via codex --ossLicence and harness support differ per model

Cursor. SpaceXAI’s Grok 4.7 was available in Cursor from its release day, according to MarkTechPost (2026-09-21). No Cursor default model, model pool or Cursor-side price could be verified on 2026-09-26, so check the model picker and Cursor’s pricing page before you rely on one.

Prices are USD per million tokens (MTok), input / output, standard tier, checked 2026-09-26 on each vendor’s price page unless the row names another source.

ModelAPI IDPrice in / outContext / max outputStatus
Claude Opus 5.5claude-opus-5-5$4 / $201M / 128KCurrent default, released 2026-09-22
Claude Fable 5.1claude-fable-5-1$10 / $501M / 128KCurrent; never a default
Claude Sonnet 5claude-sonnet-5$2 / $10, permanent1M / 128KCurrent; the sonnet alias
Claude Haiku 4.5claude-haiku-4-5$1 / $5200K / 64KCurrent Haiku; no effort parameter
Claude Opus 5claude-opus-5$5 / $251M / 128KLegacy, still served
Claude Opus 4.8claude-opus-4-8$5 / $25—Legacy, still served

Batch processing halves every rate. Cache reads cost 0.1x input, except 0.05x on Opus 5.5 ($0.20) and 0.025x on Fable 5.1 ($0.25). Opus 5.5 fast mode is $8 / $40, a research preview on the Claude API only. Claude 4.7 and later models use a tokenizer that produces about 30% more tokens for the same text, so compare cost per task, not price per token. Bedrock IDs add the anthropic. prefix, for example anthropic.claude-opus-5-5.

Tool subscription prices (Claude Pro and Max, ChatGPT plans, Cursor plans) live on the tool pricing comparison, not here.

How much does one task cost on each model?

Section titled “How much does one task cost on each model?”

A task that reads 200K uncached input tokens and writes 20K output tokens costs, at list price and before caching:

ModelCost per task
Claude Fable 5.1 · GPT-6 Astra$3.00
Claude Opus 5.5$1.20
Claude Sonnet 5 · GPT-6 Sol$0.60
Claude Haiku 4.5$0.30
GPT-6 Luna$0.03

Token counts for the same text differ between vendors, so equal rows do not mean equal cost for the same task.

This is arithmetic, not a measurement. Agents re-read context on every turn, so caching changes real cost more than the list price does, and a cheaper model that needs three attempts costs more than an expensive one that needs one. Anthropic’s Claude Code cost documentation puts typical enterprise spend at “around $13 per developer per active day and $150-250 per developer per month” (checked 2026-09-26). Measure cost per accepted task on your own work before you budget.

How do you change model and effort in each tool?

Section titled “How do you change model and effort in each tool?”

From v2.1.280 on the latest release channel, Opus 5.5 is the default on every paid plan, the Anthropic API, Claude Platform on AWS, Bedrock and Google Cloud’s Agent Platform. The stable channel (v2.1.274 on 2026-09-26) had not reached v2.1.280: there, Pro and Team Standard default to Sonnet 5 and the other plans to Opus 5. Microsoft Foundry defaults to Sonnet 4.5 (Sonnet 4.5’s retirement commitment is “not sooner than” 2026-09-29; check the deprecations page before relying on this default). Opus 5.5 runs at medium effort by default; Fable 5.1 and Sonnet 5 run at high, and Haiku 4.5 has no effort setting.

Terminal window
# Terminal: start a session on a chosen model and effort
claude --model opus --effort high
# Try fallbacks in order when the primary is overloaded
claude --model opus --fallback-model sonnet

Inside a session, /model switches model and /effort sets low, medium, high, xhigh or max, and /effort ultracode runs xhigh with automatic workflow orchestration. The aliases are opus (Opus 5.5), sonnet (Sonnet 5), fable (Fable 5.1) and opusplan (Opus in plan mode, Sonnet for execution). To pin models for a team, use the availableModels, enforceAvailableModels and maxEffortLevel settings in managed settings.

How do you prove a model change was worth it?

Section titled “How do you prove a model change was worth it?”

A model switch is a change to your production system, so it gets the same evidence as a code change. You do not judge it by reading the agent’s diffs; you judge it by what passes your gates.

  1. Pin today’s setup. Record the exact model ID (claude-opus-5-5, gpt-6-astra), the effort level and the tool version, so an alias that moves later does not change behaviour silently.

  2. Sweep effort before models. Run the same tasks at the next effort level up. Anthropic’s Claude Code docs say “Opus 5.5 at medium matches or exceeds Opus 5 at high”, which is the pattern to expect: effort is the cheaper lever.

  3. Run your own eval set. Use 20–40 tasks taken from your merged pull requests, each with a test command that decides pass or fail. Building that eval set is covered on the benchmarks page.

  4. Compare three numbers. Accepted-task rate (tests, type check and lint pass, and the change survives review), cost per accepted task, and wall-clock time per task.

  5. Roll out one workflow (loop) at a time, with a rollback. The tech lead who owns the managed settings signs off, changes the pin for one agent workflow at a time, and reverts the pin if the accepted-task rate drops. The new-model playbook has the full 48-hour procedure.

Replace LEVEL with medium, high or xhigh, MERGED_SHA with the merge commit of a real pull request, ISSUE_TEXT with its original issue text, and the two commands with your own test and type-check commands. Run the prompt in three fresh sessions, one per effort level: /effort medium, high and xhigh in Claude Code, or -c model_reasoning_effort="medium" and so on in Codex. Start each session in its own worktree (claude --worktree or codex --worktree) so the runs cannot see each other’s changes. Read the cost of each run in /usage in Claude Code; in Codex, /status shows estimated credits or cost for eligible workspaces (0.148.0+) and /usage shows account usage (0.156.0+); otherwise compare wall-clock time and pass or fail. Then compare the three reports.

What do the benchmarks say about these models?

Section titled “What do the benchmarks say about these models?”

On the official Terminal-Bench 4.0 leaderboard (read 2026-09-26), Claude Fable 5.1 (max effort) running in Claude Code scores 57.9% ± 3.8, Claude Opus 5 (max) in Claude Code 51.8% ± 3.4, and GPT-5.6 Sol (max) in Codex 37.3% ± 3.8. Every one of those entries is a max-effort run, so none of them shows how a model performs at its default effort. Neither Claude Opus 5.5 nor GPT-6 Astra had an entry on the board that day. Anthropic reports 66.4% for Opus 5.5 at xhigh in its 2026-09-22 launch table, and itself warns: “at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.”

Every score measures a model inside one agent harness on someone else’s tasks. Read how to read coding-agent benchmarks before you quote one, and let your own eval set make the decision.

Frequently asked questions

Which model is the default in Claude Code and in Codex?

Claude Opus 5.5 is the Claude Code default on every paid plan and the Anthropic API from v2.1.280 on the latest release channel; the stable channel still gives older defaults. GPT-6 Astra is the Codex bundled default since CLI 0.153.4 (2026-09-04).

How much do the current Claude models cost per million tokens?

Claude Opus 5.5 is $4 / $20, Claude Fable 5.1 is $10 / $50, Claude Sonnet 5 is $2 / $10 (the planned rise to $3 / $15 was cancelled), and Claude Haiku 4.5 is $1 / $5, input / output, checked 2026-09-26.

Should I raise effort or switch to a bigger model?

Raise effort first. Opus 5.5 runs at medium effort by default in Claude Code and GPT-6 Astra at low in Codex; move to high or xhigh with /effort or model_reasoning_effort, and switch model only when your own evals show the gain.

When is Claude Fable 5.1 worth its price?

For the longest-horizon and hardest reasoning work, chosen by hand with /model fable. It is never a default, costs $10 / $50 per million tokens, and may bill to usage credits depending on plan.