Skip to content

Vendor risk management — test continuity, not logos

AI vendor risk management keeps critical engineering work running safely when an AI provider changes a model, limit, route, price, term or product. The CTO Scorecard Q25 maximum needs a tested fallback and degraded mode, exportable artifacts, a contract and data-location review, a named owner, and a recovery exercise. A second vendor alone does not earn it.

This page is for the CTO or VP Engineering who answers for delivery when a vendor moves, and for the tech lead who runs the recovery drill. Before this page: AI usage cost governance. Recent weeks show why. On 2026-08-31 Codex retired GPT-5.4 and moved its users to GPT-6 Sol. On 2026-09-04 GPT-6 Astra became Codex’s bundled default. On 2026-09-22 Claude Code’s latest channel switched its default to Claude Opus 5.5. None of those changes waited for your release calendar, and each one moved your cost per change and your review baseline.

What you’ll walk away with from a continuity review

Section titled “What you’ll walk away with from a continuity review”
  • A service continuity register that names every AI dependency, its owner, its recovery objective, and its fallback.
  • Fallback configuration for Claude Code and Codex that fails closed, so a missing agent review never counts as a passing one.
  • A recovery exercise you can run in two hours on synthetic data, with the evidence it must produce.
  • A change watch that catches model retirements and default changes before they reach your pipeline.
  • Acceptance criteria your security and finance reviewers can sign off on without reading every run.

The CTO scorecard asks “How do you manage vendor risk (Anthropic / OpenAI / Cursor / Google)?” and scores four answers.

AnswerScoreWhat stays unproven
Single vendor, no plan B0Everything: nobody knows what stops when the vendor does
Single vendor, but we track competitors1Awareness is not a fallback; nothing has been run
Documented fallback or degraded mode, tested on representative workflows2Artifacts, contracts, data location, and a repeatable exercise
Tested fallback and degraded mode, exportable artifacts, contract and data-location review, named owner, recovery exercise3Nothing structural, provided the exercise repeats

A single-vendor setup can score 3. The question is whether the stated recovery objective holds under test, not how many logos appear on the invoice. A multi-vendor setup that nobody has exercised scores 1 at best.

An outage is the least likely event. Most disruptions arrive as ordinary product changes. Each dated row below happened, or was announced, in the weeks before this page’s lastUpdated date; the last row is the generic case.

Change typeDated example (checked 2026-09-26)What it breaksSignal to watch
Model retirementCodex retired GPT-5.4 on 2026-08-31 and moved users to GPT-6 SolPinned CI jobs, eval baselinesVendor deprecation page, Codex model catalog
Default changeCodex bundled default became GPT-6 Astra on 2026-09-04; Claude Code default became Opus 5.5 from v2.1.280 (the latest channel; stable 2.1.274 had not reached it on 2026-09-26)Cost per change, review baseline, effort settingsRelease notes, stable versus latest channel
Model removed from a platformGitHub Copilot retires four models on 2026-10-02, including Claude Opus 4.7Configurations that select those models by nameCopilot model-deprecation history
Silent in-request fallbackClaude Code sends cybersecurity-flagged Opus 5.5 requests to Opus 4.8 and biology-flagged ones to Opus 5, since Opus 5.5 (2026-09-22)Quality on security work, without a visible errorModel named in run logs
Product shutdownThe Roo Code extension shut down on 15 May 2026, per its READMEThe whole workflow and its stored configurationRepository status, vendor announcements
Limits and plansAnthropic raised five-hour usage limits on Pro, Max, Team and seat-based Enterprise on 2026-09-22Capacity planning, in either directionPlan pages, admin console
Route, region, or termsAny change to retention, subprocessors, or regional processingCompliance approval for a data classContract notices, subprocessor list

Anthropic commits to at least 60 days’ notice before retiring a model (model-deprecations page, checked 2026-09-26). Sixty days is enough time only if someone reads the notice and has a tested path ready. Current model names, prices, and retirement dates live on one page, the models hub.

  1. Inventory every dependency, not only the API. For each AI workflow, list identity (SSO, seats), model, route (direct API, cloud platform, gateway), integrations (MCP servers, GitHub apps), stored artifacts, billing, and legal terms. A coding assistant that works but cannot sign anyone in is down.

  2. Set a recovery objective per workflow. Write down the acceptable outage, the accepted quality drop, the acceptable cost increase, and the manual workload the team can absorb. An optional in-editor assistant can tolerate a day. A CI review gate that blocks every merge cannot.

  3. Choose the smallest fallback that meets the objective. Options, from cheapest to most expensive: manual work with a documented procedure, a queue that waits for recovery, a reduced feature, another approved model from the same vendor, the same model through another route, and another vendor. Pick the first option that meets the objective.

  4. Approve each fallback for the same data class. A fallback route has its own retention, region, and subprocessor terms. Put it through the same review as the primary in the AI data and compliance policy before you need it. Availability never overrides compliance.

  5. Configure the fallback to fail closed. When both the primary and the fallback are unavailable, the workflow must report “incomplete” and route to a human, never pass silently. The per-tool configuration follows below.

  6. Exercise it, then repeat on a schedule. Run the recovery exercise below once per quarter and after every vendor notice that touches a critical workflow. File the evidence with the register.

Copy this table into your tooling policy and keep one row per critical workflow. The two rows are filled-in examples.

WorkflowCriticalityPrimary (tool, model, route)FallbackDegraded modeRecovery objectiveData class approved for fallbackExportable artifactsOwnerLast exercise
PR review agent in CIBlocks mergesClaude Code, Opus 5.5, Anthropic APISonnet 5 on the same routeJob marks “incomplete”, CODEOWNERS review required1 hour to degraded mode, no silent passInternal code (yes)REVIEW.md, prompts, eval setPlatform lead2026-09-12
Interactive codingOptionalCodex, GPT-6 AstraGPT-6 Sol profile; Claude Code with AGENTS.mdManual work1 working dayInternal code (yes)AGENTS.md, skills, specsHead of engineering2026-08-29

The Exportable artifacts column is where portability becomes concrete. AGENTS.md, Agent Skills, MCP configuration, specifications, and eval sets travel between tools. Which assets stay behind, and how to rehearse a full switch, is the subject of avoiding lock-in.

The tools differ in one important way. Claude Code has a built-in model fallback flag. Codex 0.157.1 has none (checked in codex --help and codex exec --help), so the fallback is a second profile that your script calls. For Cursor, this page names no model or default, because none could be verified on cursor.com on 2026-09-26.

Add a model fallback to CI jobs. Claude Code v2.1.283 accepts --fallback-model, which switches to the listed model or models “when the default model is overloaded or not available” and retries the primary at the start of each user turn. Wrap the job so that total failure blocks the merge:

Terminal window
# CI step (terminal). The API key comes from the CI secret store.
if ! claude -p "Review the diff against origin/main using the rules in REVIEW.md. Report findings as JSON." \
--model claude-opus-5-5 --fallback-model claude-sonnet-5 \
--allowedTools "Bash(git diff *)" "Bash(git log *)" Read Grep Glob \
--output-format json --max-budget-usd 5 > review.json; then
echo "agent-review: INCOMPLETE - agent run failed (model unavailable or budget exhausted), human review required" >&2
exit 1
fi

Budget exhaustion fails closed too: when --max-budget-usd runs out, claude -p exits 1 with the result subtype error_max_budget_usd, so the message names both causes.

A claude -p run starts in Manual mode, which denies edits and shell commands unless you allow them. Manual mode plus this allow-list lets the job read the diff but not edit files. Record which model answered in the job log: a fallback that nobody sees becomes the new baseline without an eval.

Keep a second route ready. The same Claude models run through Amazon Bedrock, Google Cloud’s Agent Platform (formerly Vertex AI), and Microsoft Foundry. Aliases resolve differently there: on 2026-09-26 the sonnet alias meant Sonnet 5 on the Anthropic API but Sonnet 4.5 on Bedrock, Google Cloud, and Foundry. Pin full model IDs on every route. Route setup and regional controls are covered in where the model runs.

Slow down default changes. Set "autoUpdatesChannel": "stable" for the fleet. The stable channel runs about a week behind latest, which gives you a week to run evals before a new default reaches everyone. On 2026-09-26 stable is v2.1.274, where Opus 5.5 is unavailable (it needs v2.1.280+, the latest channel). Pin the CI runner to a latest version you have evaluated (for example npm @anthropic-ai/claude-code@2.1.283), or keep stable and pin claude-opus-5 until stable reaches 2.1.280.

A recovery exercise proves the register is true. Run it on a branch or a sandbox repository with synthetic data, never against production.

ScenarioHow to simulate it safelyExpected resultEvidence to file
Primary model unavailableSet the CI job’s --model or -m to a model ID your key is not entitled to, or that does not existCodex: the fallback profile runs. Claude Code: on v2.1.285 an unrecognised primary ID engages the fallback; the log shows unrecognized_model and modelUsage in the JSON output names the fallback. Re-run this after each upgradeJob log, time to green
Every model unavailableRevoke the test job’s API keyJob fails closed with “incomplete”; a human reviewer is requestedJob log, the blocked PR
Default model changedRun the eval set on the newest default before the fleet movesScores within the agreed tolerance, or the pin staysEval report
Tool unavailable for a weekOne engineer finishes a representative task in the fallback toolTask accepted by the normal gates within the objectivePR link, time spent
Route or terms changeRemove the fallback route’s approval in the registerNobody can select it until it is reapprovedUpdated register row
Artifact exportExport the portable files and restore them in a clean cloneThe fallback tool reads the instructions and skillsRestored repository

Measure recovery time, quality change, cost change, manual workload, and any capability lost. The named owner signs the result. A failed scenario is a useful finding: add it to the AI tooling roadmap with a budget and a date.

How do you prove the fallback still meets the bar?

Section titled “How do you prove the fallback still meets the bar?”

A fallback that runs is not yet a fallback that works. You prove it with the same evidence you use for the primary, without reading every diff:

  • The gates stay identical. Tests, type checks, lint, security scans, and required reviewers apply unchanged on the fallback path. A fallback that skips a gate is a policy violation, not a degraded mode.
  • Evals run on the fallback model. Run your eval set on the fallback before you approve it, and again after every vendor change. The procedure is the same as in the 48-hour new-model playbook. The eval set itself comes from evaluating coding agents.
  • Acceptance is measured, not assumed. Compare the share of fallback-produced changes that pass the gates with the primary’s share over the same period.
  • Cost stays visible. A fallback to a pricier route shows up in the monthly spend review described in AI usage cost governance.
  • Sign-off is named. The workflow owner signs the exercise evidence; security signs the data-class approval for each route.

Copy-paste prompts for a continuity review

Section titled “Copy-paste prompts for a continuity review”

Paste these into Claude Code or Codex at the root of the repository that holds your tooling policy. They work the same in both tools.

Abstraction without equivalence. A gateway or router can switch endpoints, but it cannot make models, context windows, tool calling, or regional terms equivalent. The sonnet alias example above shows how the same name can mean a different model on another route. Recovery: pin full model IDs and approve each route with its own eval run.

A silent fallback becomes the baseline. Automatic fallback keeps the job green while quality drops, and the team calibrates to the weaker model. Recovery: log the answering model on every run and alert when the fallback handles more than an agreed share of runs in a week.

The agent gate fails open. A review job that exits 0 when the model is unreachable merges unreviewed code. Recovery: apply the fail-closed wrapper above and include “every model unavailable” in each exercise.

The fallback breaks the data policy. A second vendor or region was never approved for the code it now sees. Recovery: remove it from the register until it passes the vendor questionnaire, and check the logs for data already sent.

Personal accounts become the shadow fallback. During an outage, engineers switch to personal plans that your data terms do not cover. Recovery: provide an approved fallback seat in advance and enforce organization login, as described in team accounts.

The exercise ran once. A drill from last year proves nothing about this month’s defaults. Recovery: tie the exercise to the quarterly review and to every vendor notice that touches a critical workflow. If a vendor change does cause an incident, handle it with the agent incident process.

For the full set of scorecard answers, return to the CTO scorecard guide.