Skip to content

AI Model Comparison Guide

You open the model picker and see several options. Each has different strengths, context windows, and price points. This guide tells you which model to use for which task, when to switch, and how much it costs.

  • A clear default model recommendation for each tool
  • Decision criteria for when to switch models
  • Pricing breakdowns per request type
  • A model routing strategy you can use immediately
TaskRecommended ModelWhy
Complex agentic coding, multi-file refactorsClaude Opus 5Anthropic’s default Opus; leads Fable 5 on most measured benchmarks at half the price
Tasks larger than a single sittingClaude Fable 5Highest-capability tier; its lead grows with task length and complexity
Everyday Claude Code workClaude Sonnet 5Recommended cost/performance starting point; account default varies
Cheap parallel / bulk workClaude Haiku 4.5Half Sonnet’s introductory token rate; one-third after Sonnet’s promotion
Hardest Codex tasksGPT-5.6 SolFlagship tier for complex coding and professional work
Everyday Codex workGPT-5.6 TerraBalanced capability, latency, and cost
High-volume Codex workGPT-5.6 LunaLowest-cost GPT-5.6 tier
Long-running Cursor tasksGrok 4.5First-party frontier model for software and broader computer work
Fast iteration (Cursor)Cursor Composer 2.5In-house frontier-speed coding model
Large codebase (>200K tokens)Opus 5, Sonnet 5, GPT-5.6 family, or Gemini 3.1 Pro1M-class context windows
Multimodal (images, video)Gemini 3.1 ProStrong multimodal support with a 1M context window
Architecture and designClaude Opus 5Deep reasoning at high or xhigh effort
Computer use and browser agentsClaude Opus 5Beats Fable 5 on OSWorld 2.0 at roughly a third of the cost
ModelProviderContextOutput LimitRoleInput $/1MOutput $/1MSpeed
Claude Fable 5Anthropic1M128KPeak Claude$10$50Slower
Claude Opus 5Anthropic1M128KDefault Claude$5$25Standard
Claude Opus 4.8Anthropic1M128KPrevious Opus$5$25Standard
Claude Sonnet 5Anthropic1M128KEveryday Claude$2*$10*Standard
Claude Haiku 4.5Anthropic200K64KLow-cost Claude$1$5Fast
GPT-5.6 SolOpenAI1.05M128KFlagship$5$30Standard
GPT-5.6 TerraOpenAI1.05M128KBalanced$2.50$15Faster
GPT-5.6 LunaOpenAI1.05M128KHigh-volume$1$6Fastest
Grok 4.5Cursor / SpaceXAI500K API**Broad frontier$2$6Standard
Cursor Composer 2.5Cursor200KCoding specialist$0.50$2.50Fast
Gemini 3.1 ProGoogle1MMultimodal$2***$12***Standard

* Sonnet 5 introductory pricing through August 31, 2026; standard pricing is $3 / $15 per MTok afterward.

** SpaceXAI documents a 500K API context for Grok 4.5. The context exposed by a host such as Cursor can be lower, so verify the current picker before a large-context run.

*** Gemini 3.1 Pro prices per tier: $2 / $12 per MTok for prompts up to 200K input tokens, rising to $4 / $18 above that threshold.

The highest-capability tier, for work that outlasts a single sitting.

  • Released: June 9, 2026
  • Context window: 1M tokens with a 128K output limit
  • Key strength: Anthropic states Fable 5’s lead over its other models grows with task length and complexity — it is the tier for long autonomous runs, not for routine throughput. Benchmark rankings depend on effort, harness, and task mix, and Opus 5 now leads it on most published measurements.
  • Available in: Claude Code v2.1.170+ (/model fable), Cursor (model picker), Claude API (claude-fable-5)

When to use: When the task horizon itself is the differentiator — overnight runs, migrations that span many sessions — select Fable 5 explicitly for the main session. Fable is never the automatic account default, and subagents can inherit the parent model, so pin cheaper subagents to Sonnet or Haiku when cost matters. Before paying 2x, try raising Opus 5’s effort level: low and medium effort on the current generation often beat xhigh on prior models.

Pricing: $10 / $50 per 1M tokens (input/output) — exactly 2x Opus 5. Effort levels run low, medium, high, xhigh, and max. Thinking is always on and cannot be disabled; sending thinking: {"type": "disabled"} returns a 400. Fable 5 requires 30-day data retention and is unavailable to organizations under zero data retention.

Fable 5 is the generally-available, safety-tuned member of the Mythos class — in Anthropic’s words, “a Mythos-class model that we’ve made safe for general use.” Its sibling, Claude Mythos 5, is the same underlying model with safeguards lifted in some areas; initial access is restricted to Project Glasswing cyber defenders and critical-infrastructure providers.

The default Claude tier for complex agentic coding and enterprise work.

  • Released: July 24, 2026
  • Context window: 1M tokens — both the default and the maximum — with a 128K output limit
  • Key strength: Anthropic reports state-of-the-art results on Frontier-Bench v0.1 (more than doubling Opus 4.8 at a lower cost per task) and GDPval-AA v2, roughly 3x the next-best model on ARC-AGI-3, and a win over Fable 5 on OSWorld 2.0 at about a third of the cost. At max effort it lands within 0.5 points of Fable 5’s peak CursorBench 3.2 score for half the cost per task. Its May 2026 knowledge cutoff is the most current of any Claude model
  • Available in: Claude Code v2.1.219+, Cursor (model picker), Claude API (claude-opus-5), Bedrock (anthropic.claude-opus-5), Google Cloud, Microsoft Foundry

When to use: Make it the default. Architecture decisions, multi-file refactors, complex debugging, long-running agents, computer use, and knowledge work at scale. Opus 5 is the account default on Max, Team Premium, Enterprise pay-as-you-go, Anthropic API, Claude Platform on AWS, Bedrock, and Google Cloud sessions; Pro, Team Standard, and Enterprise seats default to Sonnet 5.

Pricing: $5 / $25 per 1M tokens (input/output) — identical to Opus 4.8 and half of Fable 5. Fast mode runs at $10 / $50 for roughly 2.5x the output speed, on the Claude API only. Batch processing halves both rates. The minimum cacheable prompt drops to 512 tokens, down from 1024 on Opus 4.8, so shorter prompts now cache with no code change.

The previous Opus, still available — and the fallback target for flagged requests.

  • Released: May 28, 2026
  • Context window: 1M tokens with a 128K output limit
  • Key strength: Around four times less likely than Opus 4.7 to leave flaws in its own code unflagged, with strong long-horizon agentic performance
  • Available in: Claude Code, Cursor (model picker), Anthropic API, Bedrock, Google Cloud

When to use: Opus 5 supersedes it at the same price, so new work should start on Opus 5. Opus 4.8 remains worth knowing about for two reasons: it is not deprecated and still serves pinned production traffic, and it is the model that cybersecurity-flagged requests on both Opus 5 and Fable 5 automatically re-run on. It also still supports Priority Tier, which Opus 5 does not.

Pricing: $5 / $25 per 1M tokens (input/output) — unchanged from Opus 4.7. Fast mode runs at 2x the standard rate for 2.5x the speed. Effort defaults to high.

The cost-effective everyday Claude tier and default on standard subscription seats.

  • Released: June 30, 2026
  • Context window: Native 1M tokens
  • Key strength: A strict improvement over Sonnet 4.6 with a wider cost/performance range; high effort can match an Opus tier on some tasks
  • Available in: Claude Code, Cursor, Anthropic API

When to use: Use it for everyday coding, large-codebase analysis, and most cost-sensitive agentic work; increase effort before moving to Opus or Fable. Sonnet 5 is the account default on Pro, Team Standard, and Enterprise subscription seats. Max, Team Premium, Enterprise pay-as-you-go, and Anthropic API sessions default to Opus 5; organization policy can override either mapping.

Pricing: $2 / $10 per 1M tokens through August 31, 2026; $3 / $15 afterward.

The cheap, fast tier that powers parallel work.

  • Released: October 2025
  • Context window: 200K tokens
  • Key strength: Fast and inexpensive enough to drive subagents, codemods, and bulk file edits; its $1/$5 rate is half Sonnet 5’s introductory $2/$10 rate and one-third of the later $3/$15 rate
  • Available in: Claude Code (subagents and /model), Anthropic API

When to use: Read-only exploration, bulk scans, fan-out subagents, and simple formatting where you don’t need frontier reasoning. The Tier 1 model in a model-routing strategy.

Pricing: $1 / $5 per 1M tokens (input/output).

Three durable capability tiers for Codex, ChatGPT Work, and the OpenAI API.

  • Released: General availability July 9, 2026 (preview June 26)
  • Context window: 1.05M tokens with a 128K output limit for Sol, Terra, and Luna
  • Tiers: gpt-5.6-sol for frontier work, gpt-5.6-terra for balanced everyday work, and gpt-5.6-luna for efficient high-volume workloads; the gpt-5.6 alias routes to Sol
  • Available in: Codex, ChatGPT Work, and API for all three tiers; standard ChatGPT conversations expose Sol on eligible paid plans

When to use: Sol for the hardest coding, research, design, computer-use, and professional tasks; Terra for the everyday default; Luna for extraction, routing, bulk operations, and latency/cost-sensitive work. In Codex, Free and Go use Terra; Plus, Pro, Business, and Enterprise can choose all three.

Pricing: Sol $5 / $30, Terra $2.50 / $15, Luna $1 / $6 per 1M input/output tokens. Cache reads cost 10% of input and cache writes 1.25× uncached input. Requests above 272K input tokens are billed at 2× input / 1.5× output for the full request.

GPT-5.6 supports none, low, medium, high, xhigh, and max reasoning. Pro is a reasoning mode on the selected GPT-5.6 model, not a separate API model slug. The Responses API also adds persisted reasoning, explicit prompt caching, Programmatic Tool Calling, and multi-agent beta.

Cursor’s first-party frontier model for long-running work beyond software engineering.

  • Released: July 8, 2026
  • Architecture: Mixture-of-experts, jointly trained by Cursor and SpaceXAI
  • Available in: Cursor desktop, web, iOS, CLI, and SDK
  • Pricing: $2 / $6 per 1M input/output tokens; fast variant $4 / $18

When to use: Difficult, long-running tasks that need creative tool use across software, data science, finance, legal work, or general computer interaction. Composer 2.5 remains the better fit for smaller, speed-focused coding loops.

Grok 4.5 is not a replacement for Composer 2.5. Cursor describes them as different model weight classes, says Composer 2.5 will remain available, and plans more models of that smaller size. Grok is the larger, broader frontier model; Composer is the cheaper coding specialist.

Cursor excluded Grok 4.5 from its launch CursorBench comparison because an older snapshot of the Cursor codebase was accidentally present in training data; do not use that benchmark as independent evidence of model quality.

A strong multimodal model with a large context window.

  • Released: February 2026
  • Context window: 1M tokens
  • Key strength: Image, audio, and video analysis with Deep Think mode for complex reasoning.
  • Available in: Cursor (model picker), direct API

When to use: Tasks requiring more than 200K tokens of context, multimodal analysis (diagrams, screenshots, video walkthroughs), or when you need Deep Think reasoning mode.

Pricing: Tiered by prompt size. $2 / $12 per 1M tokens (input/output) for prompts up to 200K input tokens; above that threshold input doubles to $4 and output rises to $18. Budget for the higher tier whenever you are deliberately using the large context window.

Cursor’s smaller, lower-cost coding specialist.

  • Released: May 18, 2026
  • Architecture: Mixture-of-Experts, enhanced with Cursor’s own continued pretraining and reinforcement learning
  • Context window: 200K tokens
  • Key strength: A substantial step up over Composer 2 — better at sustained work on long-running tasks and more reliable at following complex instructions
  • Available in: Cursor and Grok Build

When to use: Fast local iteration in Cursor. Optimized for multi-file edits, code generation, refactoring, and long task chains across hundreds of actions.

Pricing: $0.50 / $2.50 per 1M tokens (standard); $3.00 / $15.00 (fast variant, the default).

Use this decision tree for day-to-day work:

  1. Check the actual account default: Sonnet 5 on standard Anthropic subscription seats, Opus 5 on premium/direct-API accounts and the managed clouds, Sonnet 4.5 on Microsoft Foundry; Terra for everyday Codex work and Sol for the hardest tasks
  2. Complex work on the default not landing? Raise Opus 5’s effort level before changing models — high is the default and xhigh suits most agentic coding. This is usually cheaper than a tier upgrade
  3. Task horizon spans sessions? Then select Fable 5 explicitly for the main session. Pin subagents to Sonnet/Haiku if cost matters; they may inherit the selected parent model rather than auto-routing down
  4. Need frontier long-running work in Cursor? Use Grok 4.5; switch to Composer 2.5 for faster iteration
  5. Need budget savings? Use GPT-5.6 Luna or Haiku 4.5 for bulk/parallel work
  6. Context exceeds 200K? Use Opus 5, Sonnet 5, any GPT-5.6 tier, or Gemini 3.1 Pro (note Gemini’s price step above 200K)
  7. Multimodal analysis? Gemini 3.1 Pro
  8. Everything else? Stay with the default

The examples below assume 80% uncached input and 20% output tokens. They are arithmetic illustrations, not measured task averages; caching, reasoning tokens, tool calls, and host pricing change real cost.

Total tokensOpus 5Sonnet 5*GPT-5.6 SolComposer 2.5
1K~$0.009~$0.0036~$0.010~$0.0009
10K~$0.09~$0.036~$0.10~$0.009
50K~$0.45~$0.18~$0.50~$0.045
100K~$0.90~$0.36~$1.00~$0.09

The Opus 5 column also applies to Opus 4.8, which is priced identically. A Claude Fable 5 request costs exactly 2x that column — $10 / $50 per 1M tokens versus $5 / $25 — and Opus 5 fast mode costs the same as Fable 5 at $10 / $50. * The Sonnet 5 column uses its July 2026 introductory rate.

Cursor meters model use against the usage included in your plan and applies the selected model’s current rate. Grok 4.5, for example, is $2/$6 per MTok in standard mode and $4/$18 in fast mode. Plan allowances and available models can change; verify them in Cursor Settings > Usage and on the official pricing page.

Snapshot verified July 11, 2026. These are launch results published by each model provider; -- means the provider did not publish a directly comparable value in the cited release.

Model and settingSWE-Bench ProDeepSWE 1.1Terminal-Bench 2.1
GPT-5.6 Sol (max)64.6%72.7%88.8%
GPT-5.6 Terra (max)63.4%69.6%87.4%
GPT-5.6 Luna (max)62.7%67.2%84.7%
Grok 4.5 (high)64.7%53.0%83.3%
Claude Sonnet 5 (max)63.2%80.4%

Cursor’s Grok launch chart separately reported Grok 4.5 versus Composer 2.5 at 64.7% vs 54.0% on SWE-Bench Pro, 83.3% vs 73.0% on Terminal-Bench 2.1, and 62.0% vs 18.0% on DeepSWE 1.0. That supports a large capability gap on these launch runs, but not a deprecation claim: Cursor explicitly continues both models. Cursor also notes that some competitor scores in its chart were self-reported.

Artificial Analysis results below were read July 27, 2026. The Intelligence Index is v4.1; the Coding Agent Index is v1.1 and averages DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA. Artificial Analysis supported Anthropic in evaluating Claude Opus 5 ahead of release, so treat its Opus 5 figures as vendor-adjacent rather than fully arm’s-length.

Model + measured systemIntelligence Index v4.1Coding Agent Index v1.1Cost per task
Claude Opus 5 max / Claude Code61joint 1st*$2.03
Claude Fable 5 max + fallback / Claude Code6077$2.75
GPT-5.6 Sol max / Codex5980
Claude Opus 4.8 max / Claude Code5673$1.80
GPT-5.6 Terra max / Codex5577
Grok 4.5 high / Grok Build5476
Claude Sonnet 5 max53$1.53
GPT-5.6 Luna max / Codex5175
Composer 2.5 / Cursor CLI52

* Artificial Analysis places Opus 5 at xhigh with Claude Code in joint first place on the Coding Agent Index alongside GPT-5.6 Sol with Codex, without publishing a separate index value for it. Opus 5 tops the Intelligence Index at 61, narrowly ahead of Fable 5 at 60, and sets the highest GDPval-AA v2 (1,861 Elo versus Fable 5’s ~1,747) and AA-Briefcase (1,720 versus ~1,574) scores measured so far.

The cost-per-task column is the number that matters most for routing: Opus 5 delivers the top Intelligence Index score at roughly 26% lower cost per task than Fable 5. Opus 4.8 is cheaper still per task, but scores five points lower.

The current Composer 2.5 v1.1 breakdown is approximately 16% DeepSWE, 67% Terminal-Bench v2, and 72% SWE-Atlas-QnA, with the same 52 index for Standard and Fast. Standard averaged about $0.08 and 9.7 minutes per task; Fast about $0.55 and 6.8 minutes. The often-cited Composer score of 62 came from the May v1.0 index, before DeepSWE replaced SWE-Bench-Pro-Hard-AA, and is not directly comparable to these v1.1 values.

Grok 4.5 pairs strong independent coding performance with comparatively low measured cost, but its AA-Omniscience result is a useful warning against overgeneralizing: Artificial Analysis reported 52% accuracy and a 54% hallucination rate. Keep verification in the loop for factual work.

  1. Identify your primary tool: Cursor, Claude Code, or Codex

  2. Check the selected/default model: Account/provider-specific in Claude Code (managed-cloud defaults differ), Terra/Sol by Codex workload, or Composer 2.5/Grok 4.5 by Cursor latency and complexity

  3. Evaluate task complexity: Simple tasks do not need the most expensive model

  4. Check context requirements: Workloads exceeding 200K tokens need Opus 5, Sonnet 5, a GPT-5.6 tier, or another verified 1M-context model

  5. Consider budget: Track with /cost (Claude Code), Settings > Usage (Cursor), or Codex dashboard

  6. Adjust as needed: Switch models based on task, not habit

  1. Default to the best model for tasks that matter — architecture, security review, complex debugging
  2. Downgrade for routine work — simple fixes, boilerplate, and formatting do not need an Opus tier
  3. Use speed models for iteration — Composer 2.5 in Cursor for rapid trial-and-error cycles
  4. Route bulk work to Haiku 4.5 — subagents, codemods, and fan-out scans cost a fraction of Opus
  5. Monitor costs weekly — track which models provide the best ROI for your workflow
  6. Stay updated — model capabilities and pricing change frequently. Check the Updates page