Vendor risk management — test continuity, not logos
AI vendor risk management keeps critical engineering work running safely when an AI provider changes a model, limit, route, price, term or product. The CTO Scorecard Q25 maximum needs a tested fallback and degraded mode, exportable artifacts, a contract and data-location review, a named owner, and a recovery exercise. A second vendor alone does not earn it.
This page is for the CTO or VP Engineering who answers for delivery when a vendor moves, and for the tech lead who runs the recovery drill. Before this page: AI usage cost governance. Recent weeks show why. On 2026-08-31 Codex retired GPT-5.4 and moved its users to GPT-6 Sol. On 2026-09-04 GPT-6 Astra became Codex’s bundled default. On 2026-09-22 Claude Code’s latest channel switched its default to Claude Opus 5.5. None of those changes waited for your release calendar, and each one moved your cost per change and your review baseline.
What you’ll walk away with from a continuity review
Section titled “What you’ll walk away with from a continuity review”- A service continuity register that names every AI dependency, its owner, its recovery objective, and its fallback.
- Fallback configuration for Claude Code and Codex that fails closed, so a missing agent review never counts as a passing one.
- A recovery exercise you can run in two hours on synthetic data, with the evidence it must produce.
- A change watch that catches model retirements and default changes before they reach your pipeline.
- Acceptance criteria your security and finance reviewers can sign off on without reading every run.
Which answer earns the full Q25 score?
Section titled “Which answer earns the full Q25 score?”The CTO scorecard asks “How do you manage vendor risk (Anthropic / OpenAI / Cursor / Google)?” and scores four answers.
| Answer | Score | What stays unproven |
|---|---|---|
| Single vendor, no plan B | 0 | Everything: nobody knows what stops when the vendor does |
| Single vendor, but we track competitors | 1 | Awareness is not a fallback; nothing has been run |
| Documented fallback or degraded mode, tested on representative workflows | 2 | Artifacts, contracts, data location, and a repeatable exercise |
| Tested fallback and degraded mode, exportable artifacts, contract and data-location review, named owner, recovery exercise | 3 | Nothing structural, provided the exercise repeats |
A single-vendor setup can score 3. The question is whether the stated recovery objective holds under test, not how many logos appear on the invoice. A multi-vendor setup that nobody has exercised scores 1 at best.
What counts as a vendor change?
Section titled “What counts as a vendor change?”An outage is the least likely event. Most disruptions arrive as ordinary product changes. Each dated row below happened, or was announced, in the weeks before this page’s lastUpdated date; the last row is the generic case.
| Change type | Dated example (checked 2026-09-26) | What it breaks | Signal to watch |
|---|---|---|---|
| Model retirement | Codex retired GPT-5.4 on 2026-08-31 and moved users to GPT-6 Sol | Pinned CI jobs, eval baselines | Vendor deprecation page, Codex model catalog |
| Default change | Codex bundled default became GPT-6 Astra on 2026-09-04; Claude Code default became Opus 5.5 from v2.1.280 (the latest channel; stable 2.1.274 had not reached it on 2026-09-26) | Cost per change, review baseline, effort settings | Release notes, stable versus latest channel |
| Model removed from a platform | GitHub Copilot retires four models on 2026-10-02, including Claude Opus 4.7 | Configurations that select those models by name | Copilot model-deprecation history |
| Silent in-request fallback | Claude Code sends cybersecurity-flagged Opus 5.5 requests to Opus 4.8 and biology-flagged ones to Opus 5, since Opus 5.5 (2026-09-22) | Quality on security work, without a visible error | Model named in run logs |
| Product shutdown | The Roo Code extension shut down on 15 May 2026, per its README | The whole workflow and its stored configuration | Repository status, vendor announcements |
| Limits and plans | Anthropic raised five-hour usage limits on Pro, Max, Team and seat-based Enterprise on 2026-09-22 | Capacity planning, in either direction | Plan pages, admin console |
| Route, region, or terms | Any change to retention, subprocessors, or regional processing | Compliance approval for a data class | Contract notices, subprocessor list |
Anthropic commits to at least 60 days’ notice before retiring a model (model-deprecations page, checked 2026-09-26). Sixty days is enough time only if someone reads the notice and has a tested path ready. Current model names, prices, and retirement dates live on one page, the models hub.
Build the continuity plan in six steps
Section titled “Build the continuity plan in six steps”-
Inventory every dependency, not only the API. For each AI workflow, list identity (SSO, seats), model, route (direct API, cloud platform, gateway), integrations (MCP servers, GitHub apps), stored artifacts, billing, and legal terms. A coding assistant that works but cannot sign anyone in is down.
-
Set a recovery objective per workflow. Write down the acceptable outage, the accepted quality drop, the acceptable cost increase, and the manual workload the team can absorb. An optional in-editor assistant can tolerate a day. A CI review gate that blocks every merge cannot.
-
Choose the smallest fallback that meets the objective. Options, from cheapest to most expensive: manual work with a documented procedure, a queue that waits for recovery, a reduced feature, another approved model from the same vendor, the same model through another route, and another vendor. Pick the first option that meets the objective.
-
Approve each fallback for the same data class. A fallback route has its own retention, region, and subprocessor terms. Put it through the same review as the primary in the AI data and compliance policy before you need it. Availability never overrides compliance.
-
Configure the fallback to fail closed. When both the primary and the fallback are unavailable, the workflow must report “incomplete” and route to a human, never pass silently. The per-tool configuration follows below.
-
Exercise it, then repeat on a schedule. Run the recovery exercise below once per quarter and after every vendor notice that touches a critical workflow. File the evidence with the register.
Service continuity register template
Section titled “Service continuity register template”Copy this table into your tooling policy and keep one row per critical workflow. The two rows are filled-in examples.
| Workflow | Criticality | Primary (tool, model, route) | Fallback | Degraded mode | Recovery objective | Data class approved for fallback | Exportable artifacts | Owner | Last exercise |
|---|---|---|---|---|---|---|---|---|---|
| PR review agent in CI | Blocks merges | Claude Code, Opus 5.5, Anthropic API | Sonnet 5 on the same route | Job marks “incomplete”, CODEOWNERS review required | 1 hour to degraded mode, no silent pass | Internal code (yes) | REVIEW.md, prompts, eval set | Platform lead | 2026-09-12 |
| Interactive coding | Optional | Codex, GPT-6 Astra | GPT-6 Sol profile; Claude Code with AGENTS.md | Manual work | 1 working day | Internal code (yes) | AGENTS.md, skills, specs | Head of engineering | 2026-08-29 |
The Exportable artifacts column is where portability becomes concrete. AGENTS.md, Agent Skills, MCP configuration, specifications, and eval sets travel between tools. Which assets stay behind, and how to rehearse a full switch, is the subject of avoiding lock-in.
Configure the fallback in each tool
Section titled “Configure the fallback in each tool”The tools differ in one important way. Claude Code has a built-in model fallback flag. Codex 0.157.1 has none (checked in codex --help and codex exec --help), so the fallback is a second profile that your script calls. For Cursor, this page names no model or default, because none could be verified on cursor.com on 2026-09-26.
Add a model fallback to CI jobs. Claude Code v2.1.283 accepts --fallback-model, which switches to the listed model or models “when the default model is overloaded or not available” and retries the primary at the start of each user turn. Wrap the job so that total failure blocks the merge:
# CI step (terminal). The API key comes from the CI secret store.if ! claude -p "Review the diff against origin/main using the rules in REVIEW.md. Report findings as JSON." \ --model claude-opus-5-5 --fallback-model claude-sonnet-5 \ --allowedTools "Bash(git diff *)" "Bash(git log *)" Read Grep Glob \ --output-format json --max-budget-usd 5 > review.json; then echo "agent-review: INCOMPLETE - agent run failed (model unavailable or budget exhausted), human review required" >&2 exit 1fiBudget exhaustion fails closed too: when --max-budget-usd runs out, claude -p exits 1 with the result subtype error_max_budget_usd, so the message names both causes.
A claude -p run starts in Manual mode, which denies edits and shell commands unless you allow them. Manual mode plus this allow-list lets the job read the diff but not edit files. Record which model answered in the job log: a fallback that nobody sees becomes the new baseline without an eval.
Keep a second route ready. The same Claude models run through Amazon Bedrock, Google Cloud’s Agent Platform (formerly Vertex AI), and Microsoft Foundry. Aliases resolve differently there: on 2026-09-26 the sonnet alias meant Sonnet 5 on the Anthropic API but Sonnet 4.5 on Bedrock, Google Cloud, and Foundry. Pin full model IDs on every route. Route setup and regional controls are covered in where the model runs.
Slow down default changes. Set "autoUpdatesChannel": "stable" for the fleet. The stable channel runs about a week behind latest, which gives you a week to run evals before a new default reaches everyone. On 2026-09-26 stable is v2.1.274, where Opus 5.5 is unavailable (it needs v2.1.280+, the latest channel). Pin the CI runner to a latest version you have evaluated (for example npm @anthropic-ai/claude-code@2.1.283), or keep stable and pin claude-opus-5 until stable reaches 2.1.280.
Create a fallback profile. --profile NAME (-p) layers $CODEX_HOME/NAME.config.toml on top of the base config:
model = "gpt-6-sol"Call it when the primary fails, and fail closed when both fail:
# CI step (terminal). Credentials come from the CI secret store.PROMPT="Review the diff against origin/main using the rules in AGENTS.md. Report findings as JSON."codex exec --json -m gpt-6-astra -c default_permissions=':read-only' "$PROMPT" > review.jsonl \ || codex -p fallback exec --json -c default_permissions=':read-only' "$PROMPT" > review.jsonl \ || { echo "agent-review: INCOMPLETE - model unavailable, human review required" >&2; exit 1; }Permission profiles are beta in Codex 0.157.1. The :read-only profile keeps the review from writing to the checkout.
Watch the model catalog. Codex ships its model catalog in the binary, including each model’s replacement and retirement date. Run this in the terminal after every Codex upgrade and diff the output against last month’s copy:
codex debug models --bundled \ | jq -c '.models[] | select(.upgrade != null) | {slug, replacement: .upgrade.model, retirement_at: .upgrade.retirement_at}'On 0.157.1 it lists five models with a named replacement. Only GPT-5.4 carries a retirement date (2026-08-31).
Treat another approved tool, or manual work, as the fallback. Cursor’s model list, defaults, and plan terms could not be verified from cursor.com on 2026-09-26, so this page names none. Verify them in your own admin console and record them in the register.
Keep the working context in portable files. Specifications, eval sets, and MCP server configuration move to Claude Code or Codex without conversion. Claude Code can also import configuration from Cursor with claude import cursor --dry-run, which shows what it would bring across without writing anything.
Exercise it the same way. The recovery exercise below applies unchanged. The question is whether a Cursor user can finish the representative task in the fallback tool within the recovery objective.
Run the recovery exercise
Section titled “Run the recovery exercise”A recovery exercise proves the register is true. Run it on a branch or a sandbox repository with synthetic data, never against production.
| Scenario | How to simulate it safely | Expected result | Evidence to file |
|---|---|---|---|
| Primary model unavailable | Set the CI job’s --model or -m to a model ID your key is not entitled to, or that does not exist | Codex: the fallback profile runs. Claude Code: on v2.1.285 an unrecognised primary ID engages the fallback; the log shows unrecognized_model and modelUsage in the JSON output names the fallback. Re-run this after each upgrade | Job log, time to green |
| Every model unavailable | Revoke the test job’s API key | Job fails closed with “incomplete”; a human reviewer is requested | Job log, the blocked PR |
| Default model changed | Run the eval set on the newest default before the fleet moves | Scores within the agreed tolerance, or the pin stays | Eval report |
| Tool unavailable for a week | One engineer finishes a representative task in the fallback tool | Task accepted by the normal gates within the objective | PR link, time spent |
| Route or terms change | Remove the fallback route’s approval in the register | Nobody can select it until it is reapproved | Updated register row |
| Artifact export | Export the portable files and restore them in a clean clone | The fallback tool reads the instructions and skills | Restored repository |
Measure recovery time, quality change, cost change, manual workload, and any capability lost. The named owner signs the result. A failed scenario is a useful finding: add it to the AI tooling roadmap with a budget and a date.
How do you prove the fallback still meets the bar?
Section titled “How do you prove the fallback still meets the bar?”A fallback that runs is not yet a fallback that works. You prove it with the same evidence you use for the primary, without reading every diff:
- The gates stay identical. Tests, type checks, lint, security scans, and required reviewers apply unchanged on the fallback path. A fallback that skips a gate is a policy violation, not a degraded mode.
- Evals run on the fallback model. Run your eval set on the fallback before you approve it, and again after every vendor change. The procedure is the same as in the 48-hour new-model playbook. The eval set itself comes from evaluating coding agents.
- Acceptance is measured, not assumed. Compare the share of fallback-produced changes that pass the gates with the primary’s share over the same period.
- Cost stays visible. A fallback to a pricier route shows up in the monthly spend review described in AI usage cost governance.
- Sign-off is named. The workflow owner signs the exercise evidence; security signs the data-class approval for each route.
Copy-paste prompts for a continuity review
Section titled “Copy-paste prompts for a continuity review”Paste these into Claude Code or Codex at the root of the repository that holds your tooling policy. They work the same in both tools.
What breaks in vendor continuity plans?
Section titled “What breaks in vendor continuity plans?”Abstraction without equivalence. A gateway or router can switch endpoints, but it cannot make models, context windows, tool calling, or regional terms equivalent. The sonnet alias example above shows how the same name can mean a different model on another route. Recovery: pin full model IDs and approve each route with its own eval run.
A silent fallback becomes the baseline. Automatic fallback keeps the job green while quality drops, and the team calibrates to the weaker model. Recovery: log the answering model on every run and alert when the fallback handles more than an agreed share of runs in a week.
The agent gate fails open. A review job that exits 0 when the model is unreachable merges unreviewed code. Recovery: apply the fail-closed wrapper above and include “every model unavailable” in each exercise.
The fallback breaks the data policy. A second vendor or region was never approved for the code it now sees. Recovery: remove it from the register until it passes the vendor questionnaire, and check the logs for data already sent.
Personal accounts become the shadow fallback. During an outage, engineers switch to personal plans that your data terms do not cover. Recovery: provide an approved fallback seat in advance and enforce organization login, as described in team accounts.
The exercise ran once. A drill from last year proves nothing about this month’s defaults. Recovery: tie the exercise to the quarterly review and to every vendor notice that touches a critical workflow. If a vendor change does cause an incident, handle it with the agent incident process.
Where to go next with vendor risk
Section titled “Where to go next with vendor risk”For the full set of scorecard answers, return to the CTO scorecard guide.