Skip to content

When a new model ships: a 48-hour evaluation playbook

The 48-hour new-model playbook is a fixed procedure a team runs whenever a vendor ships a coding model: pin the current configuration, re-run the team’s own eval set against the candidate, compare accepted-task rate, cost per accepted task and latency, re-tune effort before prompts, then roll out one agent loop at a time behind written rollback criteria.

Claude Opus 5.5 lands on a Tuesday. By Wednesday, half your team is on it because Claude Code auto-updated, and another lead wants GPT-6 Astra because of a launch chart. This page is for the tech lead who decides by Thursday, the CTO who signs off, and the developer who runs the evals.

  • A two-day timeline with an owner per step, and pinning commands for Claude Code, Codex and Cursor, so a release reaches your team only after you decide.
  • A triage checklist and a decision table that turn release notes and eval results into “switch”, “stay” or “switch for some loops”.
  • Rollback criteria per agent loop and three copy-paste prompts for release day.

This page assumes you already have an internal eval set. If you do not, build one first with the benchmarks and evals guide, which sizes a first set at about 30 tasks; this playbook reuses its tasks and commands. Current model names, prices and defaults live on the model comparison hub, and this page does not repeat them.

Why a new model is recurring work, not a one-off event

Section titled “Why a new model is recurring work, not a one-off event”

In 22 days of September 2026, four vendors shipped seven models that coding teams could pick up, and two tools changed their default model on their own:

Release dateVendorModelWhat changed in the tools
2026-09-01AnthropicClaude Fable 5.1Available in Claude Code from v2.1.257; never a default
2026-09-02 (secondary: The Register)GoogleGemini 3.8 FlashAdded to GitHub Copilot on 2026-09-03
2026-09-03 (secondary: CNBC)OpenAIGPT-6 AstraCodex’s bundled default from CLI 0.153.4 on 2026-09-04, at low effort
2026-09-21 (secondary: MarkTechPost)SpaceXAIGrok 4.7Added to GitHub Copilot on 2026-09-21; day one in Cursor (secondary: MarkTechPost)
2026-09-22OpenAIGPT-6 Sol, GPT-6 LunaIn the Codex model picker from CLI 0.156.1 on 2026-09-23
2026-09-22AnthropicClaude Opus 5.5Claude Code default from v2.1.280 on the latest channel, at medium effort

Two of those rows changed what agents ran without anyone deciding it. A model change is a production change, and your tools can make it for you.

WindowStepOutputOwner
Hour 0–2Freeze the current configuration and triage the releasePinned config commit; triage notesEval owner
Hour 2–12Re-run the eval set: baseline versus candidateTwo Harbor job directoriesEval owner
Hour 12–24Compare the four numbers; re-tune effort, then instructionsComparison table; tuned candidateEval owner, one developer
Hour 24–40Canary the candidate in one agent loopCanary metrics against rollback criteriaTech lead
Hour 40–48Decide per loop and record itUpdated routing record; rollback planTech lead approves; CTO for a vendor change

The operating model explains who owns the eval suite across teams. Forty-eight hours is a target, not a deadline: if the evals show a tie, the answer is “stay and re-test at the next release”.

  1. Freeze the current configuration (hour 0). Pin the model by its full ID, pin the effort, and pin the tool version wherever your tool updates itself. Aliases move: the Claude Code 2.1.283 help text describes opus, sonnet and fable as aliases “for the latest model”, so a config that says opus changed model on 2026-09-22 without a commit. Commit the pins (CI workflow, .claude/settings.json) so that rollback later is a git revert.

  2. Triage the release notes (hour 0–2). Read the vendor’s model page and the tool’s changelog, not the launch post, for five kinds of change:

    • Default effort shifts. Claude Opus 5.5 defaults to medium effort in Claude Code and on the API, while Opus 5 defaulted to high. Anthropic’s effort docs say a request that omits effort “runs one level lower than it did on Claude Opus 5”. In Codex, GPT-6 Astra defaults to low and GPT-6 Sol and Luna to medium.
    • Removed or rejected parameters. On Claude 4.7 and later models, a non-default temperature, top_p or top_k returns a 400 error. Opus 5.5 also rejects disabling thinking and forced tool use. GPT-6 Astra has no none effort, and OpenAI’s upgrade guide says to remove temperature and top_p.
    • Token accounting. Anthropic’s pricing page says Claude 4.7 and later models use a tokenizer that produces about 30% more tokens for the same text. A lower price per token can still mean a higher cost per task.
    • Data and billing terms. Claude Fable 5.1 is a Covered Model with data retention by default and may bill to usage credits, depending on plan. Check them against your data rules first.
    • Retirement of your rollback target. Codex 0.157.1 still carries GPT-5.4 as a retired entry that migrates sessions to GPT-6 Sol automatically. Before you plan a rollback, confirm on the model comparison hub that the model you would roll back to stays available.
  3. Re-run the eval set, baseline versus candidate (hour 2–12). Same tasks, same instructions commit, same harness version; change only the model, with at least three attempts per task. With the Harbor setup from the benchmarks and evals guide, that is two commands per tool, shown in the tabs below. Save a manifest with every job: model ID, effort, claude --version or codex --version, harbor --version, and the commit of your CLAUDE.md, AGENTS.md and rules. The go/no-go prompt reads it from manifest.txt in each job directory.

  4. Compare four numbers, not one (hour 12–18). A pass rate alone picks the wrong model when the candidate wins by spending more. Compare:

    • Accepted-task rate: tasks whose verifier passed and whose patch a reviewer would merge unchanged, checked on a sample of 10 patches by a human or a calibrated judge (model-graded checks shows how to calibrate one). Record the verdicts as reviews.csv (task, accepted yes or no) in each job directory, so the go/no-go prompt can use them.
    • Cost per accepted task: total tokens or dollars for the run divided by accepted tasks. Claude Code’s --output-format json result carries total_cost_usd and duration_ms; codex exec --json emits token usage on each turn.completed event.
    • Latency: median wall time per task, which decides interactive loops.
    • Per-task flips: tasks the candidate newly passes and tasks it newly fails. Six gained and four lost is a different trade from two gained.

    On 30 tasks × 3 attempts, the 95% interval on one pass rate is about ±10 points, but you are comparing two rates, so the interval on the difference is roughly ±14 points. Correlated attempts on one task widen it further, because they are not 90 independent trials. Compute it with a paired comparison per task, for example by bootstrapping over tasks, and treat any difference inside it as a tie. Decide ties on cost, latency or reviewer acceptance.

  5. Re-tune effort before you touch instructions (hour 18–24). Run the candidate at its default effort, one level above and one level below; each tab below shows the Harbor command for the sweep. The Claude Code docs claim “Opus 5.5 at medium matches or exceeds Opus 5 at high”; the sweep checks that on your tasks. Only then audit your agent instructions for old-model workarounds (the second prompt below), remove one at a time, and re-run the tasks it affects. Never change model, effort and instructions in the same run, or you cannot tell which change moved the score.

  6. Canary the candidate in one loop (hour 24–40). An agent loop is one place where an agent runs with its own trigger and reviewer: interactive sessions, a CI fix job, a review bot, a scheduled loop, a cloud agent. Roll out in that order, most supervised first, starting with volunteers in interactive sessions. Write the rollback criteria into the rollout ticket before you flip the first pin.

  7. Decide per loop and record it (hour 40–48). Use the decision table below; the answer is often “switch for some loops”. Record the decision, the evidence and the rollback target in your routing record, as described in model routing, and add any task the candidate failed in an interesting way to the eval set.

Pin and evaluate the candidate in Claude Code, Codex and Cursor

Section titled “Pin and evaluate the candidate in Claude Code, Codex and Cursor”

The steps are the same for every tool; what differs is where the pin lives and what can move it without a commit. Harbor is installed once with uv tool install harbor; load each API key from your secret manager into the environment and never type it on a command line.

Pin the baseline model and the channel. In the project’s .claude/settings.json, set today’s model by full ID. In this example the baseline is Claude Opus 5 and the candidate is Opus 5.5:

.claude/settings.json
{
"model": "claude-opus-5"
}

For the tool version, set "autoUpdatesChannel": "stable" for anyone who should not receive a model default change on release day. On 2026-09-26 stable was v2.1.274, which runs Opus 5 but not Opus 5.5 (that needs v2.1.280+). Only canary volunteers move to the latest channel and switch their pin to claude-opus-5-5. To keep a candidate out of reach until it is approved, see managed policy for availableModels, enforceAvailableModels and, from v2.1.283 (the latest channel), deniedModels.

Pin CI jobs on the command line, which outranks every settings file except managed policy. A claude -p run starts in Manual mode, so grant only the edits and the one test command the job needs, and keep the prompt before --allowedTools, which takes a list:

Terminal window
claude -p "Run the failing test in src/billing and fix the cause, not the test." \
--model claude-opus-5 --effort high --output-format json \
--permission-mode acceptEdits --allowedTools "Bash(npm test *)"

Run the eval pair from the terminal, changing only the model:

Terminal window
harbor run -p evals/tasks -a claude-code -m anthropic/claude-opus-5 -k 3 -n 4
harbor run -p evals/tasks -a claude-code -m anthropic/claude-opus-5-5 -k 3 -n 4

Write a manifest into each job directory Harbor created, filling in the model and effort of that run:

Terminal window
{ echo "model=claude-opus-5-5"; echo "effort=medium"; claude --version; harbor --version; \
echo "instructions=$(git rev-parse HEAD)"; } > jobs/JOB_DIR/manifest.txt

Sweep effort through Harbor’s agent option reasoning_effort, which Harbor 0.23.0 passes to Claude Code as --effort:

Terminal window
harbor run -p evals/tasks -a claude-code -m anthropic/claude-opus-5-5 -k 3 -n 4 \
--ak reasoning_effort=high

Claude Code v2.1.283 accepts low, medium, high, xhigh and max; in a session, use /effort.

Read the four numbers from step 4 against this table. “Noise” means the interval on the difference between the two runs: roughly ±14 points for two runs of 30 tasks × 3 attempts, and wider with correlated attempts.

What your eval showsDecisionNext action
Candidate’s accepted-task rate is higher by more than the noise, at similar costSwitch, loop by loopCanary in interactive sessions first
Accepted-task rates tie; candidate is clearly cheaper per accepted task or fasterSwitch the loops where cost or latency dominatesHigh-volume CI loops first; keep the baseline on long-horizon work
Candidate wins only at a higher effort that costs more per accepted taskSwitch only the loops that need itRecord the effort per loop in the routing record
Everything tiesStayRecord “no change” with the job IDs; re-test at the next release
Candidate loses by more than the noise, or fails a task class you depend onStay and keep the pinAdd the failing tasks to the eval set

A launch chart is not a row in this table. Anthropic itself writes that “at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences”, and the benchmarks guide explains why a public score belongs to a model and harness pair, not to the model.

Rollback criteria to write before the rollout

Section titled “Rollback criteria to write before the rollout”

Write the thresholds into the rollout ticket before the first loop switches, so nobody negotiates them mid-incident. The numbers below are starting points for a team running 20 or more agent tasks a day.

Signal in the canary loopStart with this thresholdAction
Accepted-task rate (verifier pass and reviewer acceptance)Below a baseline rate measured on at least a few hundred historical tasks by more than the canary’s own interval: about ±22 points at 20 canary tasks, ±14 at 50. With a smaller baseline, use the interval on the difference (about ±31 and ±20) and bootstrap as in step 4Revert the pin in that loop
Cost per accepted taskMore than 25% above the baseline for three working daysLower effort one level, re-measure; revert if still over
Median wall time per task in interactive sessionsMore than 50% above the baselineRevert the interactive loop; keep batch loops if they pass
Reviewer rework rate on agent pull requestsMore than 10 points above the baselineSample 10 pull requests; revert if the rework is model-related
Security-gate failure, secret exposure or data-policy breach linked to the changeOne occurrenceRevert immediately and open an incident

Rollback is a revert of the pin commit (.claude/settings.json or the CI workflow; a ~/.codex/config.toml pin is redistributed, not reverted), which any engineer on call can make without a meeting. The rollback pipeline guide covers the same pattern for code; progressive delivery covers staged rollout mechanics; and agent cost tracking covers where the cost signal comes from.

How the switch is proven without reading every diff

Section titled “How the switch is proven without reading every diff”

Nobody reads the candidate’s code line by line. The evidence is verifier-graded evals on your own tasks, a reviewer-acceptance sample of 10 patches per configuration (the only human reading), unchanged CI gates during the canary, and a decision record with job IDs so anyone can re-run it.

Who signs off: the eval owner publishes the comparison; the tech lead approves each loop’s switch and owns the rollback; the CTO approves a change of vendor or a change that affects data terms, such as moving to a Covered Model. Continuous evals runs the same eval set on every change to your agent instructions, so the next release starts from a fresh baseline.

What goes wrong when you evaluate a new model?

Section titled “What goes wrong when you evaluate a new model?”
  • The tool upgraded before you evaluated. Symptom: sessions already run the new model on release day. Recovery: pin the full model ID (claude-opus-5) in .claude/settings.json and CI, which works on any version, and set model explicitly in Codex. Prevention for the next release: move everyone outside the canary to the stable channel.
  • The default effort dropped and nobody noticed. Symptom: “the new model is lazier” on hard tasks. Recovery: compare at matched effort as well as at each model’s default, and set effort explicitly in CI.
  • Cheaper per token, dearer per task. Symptom: the invoice rises after a switch to a lower-priced model. Recovery: compare cost per accepted task, which captures tokenizer changes, extra turns and retries.
  • The eval set is saturated. Symptom: baseline and candidate both pass nearly everything. Recovery: add recent hard tasks and every escaped agent defect; a set that cannot fail cannot choose.
  • Launch-week capacity. Symptom: timeouts and rate-limit errors skew latency and pass rates. Recovery: re-run failed trials after 24 hours before you conclude anything, and use --fallback-model in Claude Code headless jobs.
  • The rollback target is gone. Symptom: reverting the pin fails because the old model retired. Recovery: check retirement in step 2 and keep a second approved model in the routing record.