Skip to content

Model routing by evidence, not brand

Model routing is a versioned rule that assigns each class of agent work (exploration, implementation, review, bulk edits) to a model and effort level chosen by repository-specific evals. Start on the tool’s default model, tune effort before switching, and change a route only when accepted-task quality holds at lower cost or latency.

Three developers on your team use three different models, and nobody can say which one produced last week’s bad migration. One teammate switched to the most expensive model for everything after a single impressive session. Another runs the cheapest one and quietly reruns failed tasks three times. This page is for the developer who builds the routing table and the tech lead who owns it and signs off every change to it.

Scorecard Q3 · Plan: How do you choose a model and runtime for each task?

Maximum-score answer (3 points): “I route task classes with versioned evals and review quality, latency, failures, and completed-work cost.”

  • A dated routing table in docs/ai/model-routing.md that names one model and effort level per task class, and a fallback.
  • The same routes wired into Claude Code, Codex, and Cursor configuration, so nobody copies a model name into a prompt.
  • A before-and-after eval run that justifies each route with pass rate, retries, latency, and cost per accepted task.
  • Four copy-paste prompts: classify your work, run a candidate, compare results, and draft the routing change.
  • A recovery path for the day a model is retired, renamed, or starts refusing a task class.

Current model names, prices, and context windows live on one page, the AI model comparison guide. This page names models only in its dated example and links there for the numbers.

Why start on the default model and tune effort first?

Section titled “Why start on the default model and tune effort first?”

The rule this site follows is: start on the tool’s default model, tune effort before switching model, and switch model only when your own evals say so. The default is the model each vendor tells you to start with, and effort is a cheaper, reversible lever than a model switch.

As of 2026-09-26 the defaults are:

ToolDefault modelDefault effortCaveat
Claude CodeClaude Opus 5.5mediumFrom v2.1.280 on the latest release channel. The stable channel (v2.1.274 on 2026-09-26) still defaults to older models, and Microsoft Foundry sessions default to Claude Sonnet 4.5.
CodexGPT-6 AstralowBundled default since CLI 0.153.4 (2026-09-04). The server can override it for signed-in accounts.
CursorNot verifiedNot verifiedcursor.com could not be reached on 2026-09-26. Read the model picker in your own account and record what it shows.

Two consequences follow. First, raise effort on the default model (medium to high or xhigh in Claude Code, low to medium or high in Codex) before you evaluate a more expensive model. Second, confirm the release channel and provider before you compare two developers’ results, because the same command can start on different models.

Route by task class, not by person or by mood. A class earns its own route when its success criterion differs from the others. Most repositories need four to six classes. The example below is a starting hypothesis built from the vendors’ own positioning on 2026-09-26, not a result. Replace every cell with what your evals show.

Task classWhat decides the routeExample starting route (2026-09-26)FallbackMeasure
Repository explorationRelevant files found, speedA read-only subagent on Claude Haiku 4.5 or GPT-6 LunaThe session’s default modelRelevant files found vs. a known list, latency
ImplementationDiff accepted after gatesDefault model at default effort: Claude Opus 5.5 or GPT-6 AstraSame model, one effort level higherGate pass rate, retries, cost per accepted task
Long-horizon or hardest reasoningTask finished without human rescueClaude Fable 5.1 chosen by hand, or GPT-6 Astra at maxSplit the task and rerun on the defaultCompletion without intervention, cost per task
ReviewReal defects found, few false positivesA different session, and a different model where evals show it finds moreA human specialistAccepted findings, false-positive rate
Bulk mechanical editsCorrect edits at volumeClaude Sonnet 5, GPT-6 Sol, or GPT-6 LunaDefault model on the failed files onlyGate pass rate per file, cost per file
Regulated dataApproved provider, region, retentionThe policy-approved route onlyHuman-only pathPolicy violations (target: zero)

Two rules keep the table honest. A quality threshold must pass before cost or speed can decide between candidates. And a flagship model can be cheaper per accepted task than a small one if it avoids retries, so compare cost per accepted task, never price per token.

Run this for one task class at a time. Budget at least three runs per task per candidate: a 10-task set with a baseline and one candidate is 60 runs. If you do not have an eval set yet, build one from merged pull requests first.

  1. Freeze the harness. Record the CLI version, CLAUDE.md or AGENTS.md commit, skills, and MCP servers. A route compared across two harness versions tells you nothing about the model.

  2. Pick 10 to 20 tasks for the class. Each task needs a base commit and a grader script that exits 0 on success: tests the agent cannot see, the type checker, the linter, and a scope check on touched files.

  3. Run the baseline and each candidate three times per task, in a disposable worktree. Use the same prompt for every run. One run per task hides the variance that makes a cheap model look good.

  4. Grade with code, then record the sheet below. Pass rate first. Only candidates within your quality threshold of the baseline move on to cost and latency.

  5. Open a pull request that changes docs/ai/model-routing.md and the tool config together. Attach the results sheet. The tech lead reviews the numbers, not the prose, and merges.

To run the candidate matrix from one config instead of shell loops, use promptfoo (npm promptfoo, 0.123.1 on 2026-09-26), which supports Claude Code and Codex as providers: run npx promptfoo@latest init, then promptfoo eval and promptfoo view.

The commands differ by tool. Each run below is headless, stays inside the repository, and writes a machine-readable result.

Terminal window
# Terminal, inside a throwaway worktree for one task
claude -p "$(cat evals/implementation/task-07/prompt.md)" \
--model claude-opus-5-5 --effort high \
--permission-mode acceptEdits \
--allowedTools "Bash(npm test *)" "Bash(npx tsc *)" "Bash(npm run lint *)" \
--max-budget-usd 5 \
--output-format json > runs/task-07-opus55-high-run1.json
./evals/implementation/task-07/check.sh "$PWD"; echo "exit=$?"

The JSON result includes total_cost_usd and a per-model cost breakdown. Both are client-side estimates, so reconcile them with your usage dashboard monthly. Swap --model and --effort for each candidate.

evals/results/2026-09-26-implementation.yaml
date: 2026-09-26
task_class: implementation
harness: { claude_code: 2.1.283, codex: 0.157.1, agents_md_commit: 4f1c2e9 }
tasks: 12
runs_per_task: 3
candidates:
- route: claude-opus-5-5 @ medium # baseline: the default
pass_rate: 0.83
median_retries: 0
median_latency_s: 410
cost_per_accepted_task_usd: 1.90
- route: claude-opus-5-5 @ high
pass_rate: 0.92
median_retries: 0
median_latency_s: 520
cost_per_accepted_task_usd: 2.30
decision: promote "@ high" for implementation; quality gain beats the cost rise
signed_off_by: tech lead, 2026-09-26

The numbers above are placeholders that show the shape; yours come from your runs.

A route that lives only in a document drifts. Put it where the tool reads it, and keep full model IDs there rather than aliases. Aliases such as opus and sonnet resolve to different models on different providers and move when a new model ships.

Set the session model in .claude/settings.json ("model": "claude-opus-5-5"), or leave it unset to follow the default. Route classes through subagents: each subagent file takes a model and an effort field. Give exploration its own project subagent on a cheaper model, then ask for it by name or reference it in CLAUDE.md:

---
name: explore-cheap
description: Read-only codebase search. Use before any edit to find files, call sites, and tests.
tools: Read, Grep, Glob
model: claude-haiku-4-5
---
Return file paths with one line each on why they matter. Never edit files.

Claude Haiku 4.5’s retirement floor is 2026-10-15 (no deprecation notice as of 2026-09-26); rerun the exploration eval when a successor ships.

Other levers: opusplan uses Opus in plan mode and Sonnet for execution; fallbackModel in settings (or --fallback-model) holds an ordered list of fallback models tried when the primary is overloaded or unavailable; and administrators restrict choices with availableModels, plus deniedModels from v2.1.283 (the latest channel). Note that /model saves your pick as the default for new sessions; press s in the picker to switch for the current session only.

How do you know a route is right without reading every diff?

Section titled “How do you know a route is right without reading every diff?”

You judge a route by its evidence, not by reading the code it produced.

  • Graders decide pass or fail. Each eval task has a check.sh with tests hidden from the agent, the type checker, the linter, and a scope check. A route that cannot pass them does not ship, whatever it costs.
  • Production confirms the eval. After a route change, watch the acceptance rate of agent pull requests and the rework rate for that class for two weeks, using the definitions in lifecycle metrics.
  • Every change reruns automatically. A model or effort change in the routing table triggers the continuous eval suite, and a vendor model release triggers the new-model playbook.
  • The pull request records the route. The agent’s evidence bundle names the model and effort that produced the change, so a regression traces back to a route.
  • One person signs off. The tech lead owns docs/ai/model-routing.md and evals/ through CODEOWNERS and is the only one who merges a route change. The developer who proposes it attaches the results sheet.

What breaks in model routing, and how to recover

Section titled “What breaks in model routing, and how to recover”
SymptomCauseRecovery
A pinned model disappears or starts failingThe vendor retired or renamed it. Claude Haiku 4.5, for example, has a retirement floor of 2026-10-15 with no deprecation notice as of 2026-09-26.Point the row at the vendor’s named replacement, rerun that class’s eval set, and date the change. Never guess a successor from marketing copy.
Two developers get different results from the same commandDifferent release channels, providers, or a saved /model choice in either toolCompare /status output in both tools. Pin the model in project settings (.claude/settings.json, or .codex/config.toml in a trusted project), and record the CLI version on the sheet.
The cheap route looks good but costs moreRetries and human repair are not countedRecompute cost per accepted task including reruns and repair time. Promote the route with the lower total.
A Claude Code request switches model mid-taskA safety classifier re-ran a flagged cybersecurity or biology request on a fallback model, or fallbackModel fired on an outageRead the notice in the transcript. If a class trips it often, route that class to the fallback model deliberately and eval it there.
A run’s model cannot be reproducedNo model was pinned, so the tool or account default chose itRecord the model ID on the sheet. Pin an evaluated model for CI and audited runs.
A fallback crosses a data boundaryThe fallback provider is not approved for that dataKeep the task on the approved provider or use the human-only path. Availability never overrides data policy.
Everyone upgrades to the biggest modelOne impressive session, no evalRun the eval for that class. If the flagship wins on cost per accepted task, promote it through the same pull request as any other change.
  • Every important task class has a quality threshold, a current route, and a fallback.
  • Model names appear in one dated routing table and the tool config, not in prompts.
  • Each route links to a results sheet from the same harness version.
  • Effort was tuned on the default model before any other model was tried.
  • At least one fallback has run on the same eval set.
  • Cost per accepted task includes retries and human repair.
  • Permissions, sandboxing, and data boundaries are the same for every candidate.
  • The tech lead owns the table through CODEOWNERS and reviews it at least quarterly.

Before this page, choose a primary AI engineering harness (Q1) and size your plan from measured workload (Q2), and map the ecosystem of skills, MCP servers and plugins. After it, carry the routes into the AI-native SDLC and compare the harnesses in the tool map. Next on the developer track: agent frameworks.

Official references: Claude Code model configuration and subagents; the Codex configuration schema; the @cursor/sdk package.