When a new model ships: a 48-hour evaluation playbook
The 48-hour new-model playbook is a fixed procedure a team runs whenever a vendor ships a coding model: pin the current configuration, re-run the team’s own eval set against the candidate, compare accepted-task rate, cost per accepted task and latency, re-tune effort before prompts, then roll out one agent loop at a time behind written rollback criteria.
Claude Opus 5.5 lands on a Tuesday. By Wednesday, half your team is on it because Claude Code auto-updated, and another lead wants GPT-6 Astra because of a launch chart. This page is for the tech lead who decides by Thursday, the CTO who signs off, and the developer who runs the evals.
What the 48-hour playbook gives your team
Section titled “What the 48-hour playbook gives your team”- A two-day timeline with an owner per step, and pinning commands for Claude Code, Codex and Cursor, so a release reaches your team only after you decide.
- A triage checklist and a decision table that turn release notes and eval results into “switch”, “stay” or “switch for some loops”.
- Rollback criteria per agent loop and three copy-paste prompts for release day.
This page assumes you already have an internal eval set. If you do not, build one first with the benchmarks and evals guide, which sizes a first set at about 30 tasks; this playbook reuses its tasks and commands. Current model names, prices and defaults live on the model comparison hub, and this page does not repeat them.
Why a new model is recurring work, not a one-off event
Section titled “Why a new model is recurring work, not a one-off event”In 22 days of September 2026, four vendors shipped seven models that coding teams could pick up, and two tools changed their default model on their own:
| Release date | Vendor | Model | What changed in the tools |
|---|---|---|---|
| 2026-09-01 | Anthropic | Claude Fable 5.1 | Available in Claude Code from v2.1.257; never a default |
| 2026-09-02 (secondary: The Register) | Gemini 3.8 Flash | Added to GitHub Copilot on 2026-09-03 | |
| 2026-09-03 (secondary: CNBC) | OpenAI | GPT-6 Astra | Codex’s bundled default from CLI 0.153.4 on 2026-09-04, at low effort |
| 2026-09-21 (secondary: MarkTechPost) | SpaceXAI | Grok 4.7 | Added to GitHub Copilot on 2026-09-21; day one in Cursor (secondary: MarkTechPost) |
| 2026-09-22 | OpenAI | GPT-6 Sol, GPT-6 Luna | In the Codex model picker from CLI 0.156.1 on 2026-09-23 |
| 2026-09-22 | Anthropic | Claude Opus 5.5 | Claude Code default from v2.1.280 on the latest channel, at medium effort |
Two of those rows changed what agents ran without anyone deciding it. A model change is a production change, and your tools can make it for you.
The 48-hour plan at a glance
Section titled “The 48-hour plan at a glance”| Window | Step | Output | Owner |
|---|---|---|---|
| Hour 0–2 | Freeze the current configuration and triage the release | Pinned config commit; triage notes | Eval owner |
| Hour 2–12 | Re-run the eval set: baseline versus candidate | Two Harbor job directories | Eval owner |
| Hour 12–24 | Compare the four numbers; re-tune effort, then instructions | Comparison table; tuned candidate | Eval owner, one developer |
| Hour 24–40 | Canary the candidate in one agent loop | Canary metrics against rollback criteria | Tech lead |
| Hour 40–48 | Decide per loop and record it | Updated routing record; rollback plan | Tech lead approves; CTO for a vendor change |
The operating model explains who owns the eval suite across teams. Forty-eight hours is a target, not a deadline: if the evals show a tie, the answer is “stay and re-test at the next release”.
Run the playbook step by step
Section titled “Run the playbook step by step”-
Freeze the current configuration (hour 0). Pin the model by its full ID, pin the effort, and pin the tool version wherever your tool updates itself. Aliases move: the Claude Code 2.1.283 help text describes
opus,sonnetandfableas aliases “for the latest model”, so a config that saysopuschanged model on 2026-09-22 without a commit. Commit the pins (CI workflow,.claude/settings.json) so that rollback later is agit revert. -
Triage the release notes (hour 0–2). Read the vendor’s model page and the tool’s changelog, not the launch post, for five kinds of change:
- Default effort shifts. Claude Opus 5.5 defaults to
mediumeffort in Claude Code and on the API, while Opus 5 defaulted tohigh. Anthropic’s effort docs say a request that omitseffort“runs one level lower than it did on Claude Opus 5”. In Codex, GPT-6 Astra defaults tolowand GPT-6 Sol and Luna tomedium. - Removed or rejected parameters. On Claude 4.7 and later models, a non-default
temperature,top_portop_kreturns a 400 error. Opus 5.5 also rejects disabling thinking and forced tool use. GPT-6 Astra has nononeeffort, and OpenAI’s upgrade guide says to removetemperatureandtop_p. - Token accounting. Anthropic’s pricing page says Claude 4.7 and later models use a tokenizer that produces about 30% more tokens for the same text. A lower price per token can still mean a higher cost per task.
- Data and billing terms. Claude Fable 5.1 is a Covered Model with data retention by default and may bill to usage credits, depending on plan. Check them against your data rules first.
- Retirement of your rollback target. Codex 0.157.1 still carries GPT-5.4 as a retired entry that migrates sessions to GPT-6 Sol automatically. Before you plan a rollback, confirm on the model comparison hub that the model you would roll back to stays available.
- Default effort shifts. Claude Opus 5.5 defaults to
-
Re-run the eval set, baseline versus candidate (hour 2–12). Same tasks, same instructions commit, same harness version; change only the model, with at least three attempts per task. With the Harbor setup from the benchmarks and evals guide, that is two commands per tool, shown in the tabs below. Save a manifest with every job: model ID, effort,
claude --versionorcodex --version,harbor --version, and the commit of yourCLAUDE.md,AGENTS.mdand rules. The go/no-go prompt reads it frommanifest.txtin each job directory. -
Compare four numbers, not one (hour 12–18). A pass rate alone picks the wrong model when the candidate wins by spending more. Compare:
- Accepted-task rate: tasks whose verifier passed and whose patch a reviewer would merge unchanged, checked on a sample of 10 patches by a human or a calibrated judge (model-graded checks shows how to calibrate one). Record the verdicts as
reviews.csv(task, accepted yes or no) in each job directory, so the go/no-go prompt can use them. - Cost per accepted task: total tokens or dollars for the run divided by accepted tasks. Claude Code’s
--output-format jsonresult carriestotal_cost_usdandduration_ms;codex exec --jsonemits token usage on eachturn.completedevent. - Latency: median wall time per task, which decides interactive loops.
- Per-task flips: tasks the candidate newly passes and tasks it newly fails. Six gained and four lost is a different trade from two gained.
On 30 tasks × 3 attempts, the 95% interval on one pass rate is about ±10 points, but you are comparing two rates, so the interval on the difference is roughly ±14 points. Correlated attempts on one task widen it further, because they are not 90 independent trials. Compute it with a paired comparison per task, for example by bootstrapping over tasks, and treat any difference inside it as a tie. Decide ties on cost, latency or reviewer acceptance.
- Accepted-task rate: tasks whose verifier passed and whose patch a reviewer would merge unchanged, checked on a sample of 10 patches by a human or a calibrated judge (model-graded checks shows how to calibrate one). Record the verdicts as
-
Re-tune effort before you touch instructions (hour 18–24). Run the candidate at its default effort, one level above and one level below; each tab below shows the Harbor command for the sweep. The Claude Code docs claim “Opus 5.5 at
mediummatches or exceeds Opus 5 athigh”; the sweep checks that on your tasks. Only then audit your agent instructions for old-model workarounds (the second prompt below), remove one at a time, and re-run the tasks it affects. Never change model, effort and instructions in the same run, or you cannot tell which change moved the score. -
Canary the candidate in one loop (hour 24–40). An agent loop is one place where an agent runs with its own trigger and reviewer: interactive sessions, a CI fix job, a review bot, a scheduled loop, a cloud agent. Roll out in that order, most supervised first, starting with volunteers in interactive sessions. Write the rollback criteria into the rollout ticket before you flip the first pin.
-
Decide per loop and record it (hour 40–48). Use the decision table below; the answer is often “switch for some loops”. Record the decision, the evidence and the rollback target in your routing record, as described in model routing, and add any task the candidate failed in an interesting way to the eval set.
Pin and evaluate the candidate in Claude Code, Codex and Cursor
Section titled “Pin and evaluate the candidate in Claude Code, Codex and Cursor”The steps are the same for every tool; what differs is where the pin lives and what can move it without a commit. Harbor is installed once with uv tool install harbor; load each API key from your secret manager into the environment and never type it on a command line.
Pin the baseline model and the channel. In the project’s .claude/settings.json, set today’s model by full ID. In this example the baseline is Claude Opus 5 and the candidate is Opus 5.5:
{ "model": "claude-opus-5"}For the tool version, set "autoUpdatesChannel": "stable" for anyone who should not receive a model default change on release day. On 2026-09-26 stable was v2.1.274, which runs Opus 5 but not Opus 5.5 (that needs v2.1.280+). Only canary volunteers move to the latest channel and switch their pin to claude-opus-5-5. To keep a candidate out of reach until it is approved, see managed policy for availableModels, enforceAvailableModels and, from v2.1.283 (the latest channel), deniedModels.
Pin CI jobs on the command line, which outranks every settings file except managed policy. A claude -p run starts in Manual mode, so grant only the edits and the one test command the job needs, and keep the prompt before --allowedTools, which takes a list:
claude -p "Run the failing test in src/billing and fix the cause, not the test." \ --model claude-opus-5 --effort high --output-format json \ --permission-mode acceptEdits --allowedTools "Bash(npm test *)"Run the eval pair from the terminal, changing only the model:
harbor run -p evals/tasks -a claude-code -m anthropic/claude-opus-5 -k 3 -n 4harbor run -p evals/tasks -a claude-code -m anthropic/claude-opus-5-5 -k 3 -n 4Write a manifest into each job directory Harbor created, filling in the model and effort of that run:
{ echo "model=claude-opus-5-5"; echo "effort=medium"; claude --version; harbor --version; \ echo "instructions=$(git rev-parse HEAD)"; } > jobs/JOB_DIR/manifest.txtSweep effort through Harbor’s agent option reasoning_effort, which Harbor 0.23.0 passes to Claude Code as --effort:
harbor run -p evals/tasks -a claude-code -m anthropic/claude-opus-5-5 -k 3 -n 4 \ --ak reasoning_effort=highClaude Code v2.1.283 accepts low, medium, high, xhigh and max; in a session, use /effort.
Pin the model and effort in config. Codex 0.154.0 changed fresh sessions to “respect server model defaults unless explicitly overridden”, so for a signed-in ChatGPT account the effective default is decided server-side. An explicit model is the only reliable pin:
model = "gpt-6-astra"model_reasoning_effort = "low"That file lives in each developer’s home directory, outside git, so distribute it as a shared team dotfile or each pin drifts on its own. Put the pin that rollback depends on in the CI workflow, committed with the code (the :workspace permission profile, beta, lets the run write its fix):
codex exec --json -m gpt-6-astra -c model_reasoning_effort=low \ -c default_permissions=':workspace' \ "Run the failing test in src/billing and fix the cause, not the test."Keep the candidate one flag away: --profile NAME (-p) layers $CODEX_HOME/NAME.config.toml on top of the defaults:
model = "gpt-6-sol"model_reasoning_effort = "medium"codex -p candidate exec --json -c default_permissions=':workspace' \ "Run the failing test in src/billing and fix the cause, not the test."Pin the CLI in CI with an exact package version, for example npm install -g @openai/codex@0.157.1, so a new release cannot change the bundled default mid-evaluation. Codex 0.157.0 added migration prompts to GPT-6 Sol or Luna; decline them until the evaluation is done.
Run the eval pair:
harbor run -p evals/tasks -a codex -m openai/gpt-6-astra -k 3 -n 4harbor run -p evals/tasks -a codex -m openai/gpt-6-sol -k 3 -n 4Write a manifest into each job directory:
{ echo "model=gpt-6-sol"; echo "effort=medium"; codex --version; harbor --version; \ echo "instructions=$(git rev-parse HEAD)"; } > jobs/JOB_DIR/manifest.txtSweep effort with the same agent option; Harbor 0.23.0 turns it into -c model_reasoning_effort=high on the codex exec command line:
harbor run -p evals/tasks -a codex -m openai/gpt-6-sol -k 3 -n 4 \ --ak reasoning_effort=highCodex effort levels run from low to max, plus ultra (“Maximum reasoning with automatic task delegation”), which GPT-6 Luna does not offer.
Pin by selection, and verify in your own account. Cursor’s model defaults, model list and prices could not be verified from cursor.com on 2026-09-26, so this page names none. Choose the model explicitly in the model picker, not an automatic choice, and name it in your team’s rules.
Run the eval pair with Harbor’s cursor-cli agent, which requires CURSOR_API_KEY and runs the Cursor CLI in print mode:
harbor run -p evals/tasks -a cursor-cli -m cursor/BASELINE_MODEL -k 3 -n 4harbor run -p evals/tasks -a cursor-cli -m cursor/CANDIDATE_MODEL -k 3 -n 4Replace BASELINE_MODEL and CANDIDATE_MODEL with model IDs your Cursor account lists, and write a manifest.txt into each job directory as in the other tabs. Never carry a score across from Claude Code or Codex.
How to decide from the eval results
Section titled “How to decide from the eval results”Read the four numbers from step 4 against this table. “Noise” means the interval on the difference between the two runs: roughly ±14 points for two runs of 30 tasks × 3 attempts, and wider with correlated attempts.
| What your eval shows | Decision | Next action |
|---|---|---|
| Candidate’s accepted-task rate is higher by more than the noise, at similar cost | Switch, loop by loop | Canary in interactive sessions first |
| Accepted-task rates tie; candidate is clearly cheaper per accepted task or faster | Switch the loops where cost or latency dominates | High-volume CI loops first; keep the baseline on long-horizon work |
| Candidate wins only at a higher effort that costs more per accepted task | Switch only the loops that need it | Record the effort per loop in the routing record |
| Everything ties | Stay | Record “no change” with the job IDs; re-test at the next release |
| Candidate loses by more than the noise, or fails a task class you depend on | Stay and keep the pin | Add the failing tasks to the eval set |
A launch chart is not a row in this table. Anthropic itself writes that “at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences”, and the benchmarks guide explains why a public score belongs to a model and harness pair, not to the model.
Rollback criteria to write before the rollout
Section titled “Rollback criteria to write before the rollout”Write the thresholds into the rollout ticket before the first loop switches, so nobody negotiates them mid-incident. The numbers below are starting points for a team running 20 or more agent tasks a day.
| Signal in the canary loop | Start with this threshold | Action |
|---|---|---|
| Accepted-task rate (verifier pass and reviewer acceptance) | Below a baseline rate measured on at least a few hundred historical tasks by more than the canary’s own interval: about ±22 points at 20 canary tasks, ±14 at 50. With a smaller baseline, use the interval on the difference (about ±31 and ±20) and bootstrap as in step 4 | Revert the pin in that loop |
| Cost per accepted task | More than 25% above the baseline for three working days | Lower effort one level, re-measure; revert if still over |
| Median wall time per task in interactive sessions | More than 50% above the baseline | Revert the interactive loop; keep batch loops if they pass |
| Reviewer rework rate on agent pull requests | More than 10 points above the baseline | Sample 10 pull requests; revert if the rework is model-related |
| Security-gate failure, secret exposure or data-policy breach linked to the change | One occurrence | Revert immediately and open an incident |
Rollback is a revert of the pin commit (.claude/settings.json or the CI workflow; a ~/.codex/config.toml pin is redistributed, not reverted), which any engineer on call can make without a meeting. The rollback pipeline guide covers the same pattern for code; progressive delivery covers staged rollout mechanics; and agent cost tracking covers where the cost signal comes from.
How the switch is proven without reading every diff
Section titled “How the switch is proven without reading every diff”Nobody reads the candidate’s code line by line. The evidence is verifier-graded evals on your own tasks, a reviewer-acceptance sample of 10 patches per configuration (the only human reading), unchanged CI gates during the canary, and a decision record with job IDs so anyone can re-run it.
Who signs off: the eval owner publishes the comparison; the tech lead approves each loop’s switch and owns the rollback; the CTO approves a change of vendor or a change that affects data terms, such as moving to a Covered Model. Continuous evals runs the same eval set on every change to your agent instructions, so the next release starts from a fresh baseline.
Copy-paste prompts for release day
Section titled “Copy-paste prompts for release day”What goes wrong when you evaluate a new model?
Section titled “What goes wrong when you evaluate a new model?”- The tool upgraded before you evaluated. Symptom: sessions already run the new model on release day. Recovery: pin the full model ID (
claude-opus-5) in.claude/settings.jsonand CI, which works on any version, and setmodelexplicitly in Codex. Prevention for the next release: move everyone outside the canary to thestablechannel. - The default effort dropped and nobody noticed. Symptom: “the new model is lazier” on hard tasks. Recovery: compare at matched effort as well as at each model’s default, and set effort explicitly in CI.
- Cheaper per token, dearer per task. Symptom: the invoice rises after a switch to a lower-priced model. Recovery: compare cost per accepted task, which captures tokenizer changes, extra turns and retries.
- The eval set is saturated. Symptom: baseline and candidate both pass nearly everything. Recovery: add recent hard tasks and every escaped agent defect; a set that cannot fail cannot choose.
- Launch-week capacity. Symptom: timeouts and rate-limit errors skew latency and pass rates. Recovery: re-run failed trials after 24 hours before you conclude anything, and use
--fallback-modelin Claude Code headless jobs. - The rollback target is gone. Symptom: reverting the pin fails because the old model retired. Recovery: check retirement in step 2 and keep a second approved model in the routing record.