Skip to content

Avoiding lock-in: portable context, skills and pipelines

Lock-in with AI coding agents sits in configuration, not in code. AGENTS.md, Agent Skills, MCP servers, specifications and eval suites move between Claude Code, Codex and Cursor with little work. Cloud-agent environments, vendor review bots, hooks, plugin packaging and managed policy do not. Keep the substance in the repository, keep vendor files thin, and rehearse the exit on a schedule.

Your engineering organization standardized on one coding agent a year ago. Since then the team has written 40 skills, a review bot tuned to your house rules, a cloud-agent environment that builds the monorepo, and a hook that blocks secret reads. Now a competitor’s model tops your evals, the vendor has changed its plans twice, and the board asks a fair question: what would it cost to move, and how do you know?

This page is for the CTO who owns vendor strategy and the executive who signs the contract. It sits in the Spend & vendors group after vendor risk management, which classifies dependencies by failure impact, and before buying AI coding tools, where the findings from your exit rehearsal become contract questions.

  • An inventory table of every agent asset, sorted into “travels”, “travels with a converter” and “stays behind”, with the version each claim was checked against.
  • A decision table for when accepting a vendor-specific feature is worth it, written so an executive can apply it.
  • A repository layout and a CI check (tested on 2026-09-26) that fails when the portable layer drifts between tools.
  • An exit-rehearsal checklist with pass criteria and a record template, so switching cost is a measured number rather than a guess.
  • Three copy-paste prompts and the failure modes you meet when you actually try to switch.

Which agent assets travel between vendors?

Section titled “Which agent assets travel between vendors?”

An asset travels when a second tool reads the same file, or the same protocol, without a rewrite. Five kinds of asset meet that bar today, and each one has a trap that the vendor announcements leave out.

AssetOpen standard or formatWho reads it natively (checked 2026-09-26)The trap
Project instructionsAGENTS.md, stewarded by the Agentic AI Foundation (AAIF) under the Linux Foundation; “used by over 60k open-source projects”, per agents.md, September 2026Codex reads AGENTS.md natively. Claude Code reads it only when the project has no CLAUDE.md, from v2.1.277 (first-party) or v2.1.281 (Bedrock, Google Cloud, Foundry, LLM gateways), both on the latest channel; stable (2.1.274) does not yetA repository with both files gets only CLAUDE.md in Claude Code. Make @AGENTS.md the first line of CLAUDE.md so one file stays the source. Cursor reads its own Rules; generate them from the same source (see the Cursor tab)
SkillsAgent Skills, an open standard since 2025-12-18: a folder with a SKILL.md whose name is at most 64 characters and description at most 1,024; the specification’s showcase listed 46 clients on 2026-09-26Claude Code (.claude/skills/), Codex and Cursor (.agents/skills/, per the skills CLI 1.7.0 README)The format travels; the folder does not. A skill copied without its sibling skills, or without the plugin hooks that trigger it, can load and do nothing
Tool connectionsModel Context Protocol (MCP), donated by Anthropic to the AAIF on 2025-12-09; current specification 2026-07-28All three tools act as MCP clientsThe server travels; its registration does not (.mcp.json against config.toml). Claude Code negotiates the 2026-07-28 protocol by default, while Codex 0.157.1 does not (mcp_2026_07_28 is off), so a server that drops the older protocol breaks in Codex
Specifications and decisionsMarkdown in the repository: specs, acceptance criteria, architecture decision recordsEvery agent, as ordinary filesNone, as long as they live in Git and not in a vendor’s project or memory feature
Tests, gates and evalsYour CI, plus an eval harness that runs more than one agentEvery agent, through CI; promptfoo 0.123.1 has providers for the Claude Agent SDK and the Codex SDKAn eval suite that calls one vendor’s API directly is itself lock-in. So is a gate that exists only as a vendor bot’s check run

Two points follow for an executive. First, the open standards are governed by a foundation rather than by one vendor, which lowers the risk that a format disappears with a price change. Second, “travels” still means someone runs a converter or a second install command. The exit rehearsal later on this page measures that effort.

Everything below is valuable, and everything below is rebuilt by hand when you change vendors. The goal is not to avoid these features. The goal is to keep the substance in the repository and let each vendor file be a thin wrapper around it.

AssetHow each vendor stores itWhat to keep in the repository instead
Cloud-agent environmentsClaude Code cloud environments (network access level, environment variables, setup script) are configured per environment in Claude Code’s cloud settings; Codex cloud environments and Cursor Cloud Agents with Builds are configured in each vendor’s own settingsOne scripts/agent-setup.sh that installs dependencies and seeds test data. Every vendor’s setup step calls that script and nothing else
Review botsClaude Code’s managed Code Review takes guidance from CLAUDE.md or REVIEW.md and posts inline findings but never approves or blocks a pull request; Codex review reads custom review rules from AGENTS.md (checked 2026-08-28); Cursor’s Bugbot is configured in CursorReview criteria in one Markdown file that each bot’s config references, and the merge decision in your own required CI checks, never in a bot’s verdict
HooksClaude Code 2.1.283 has 33 hook events; Codex 0.157.1 has 12, and the two event sets do not line up; Cursor hooks are processes that exchange JSON over stdioThe hook logic as a script in the repository (scripts/guard-secrets.sh). Each tool’s hook registration is a short adapter that only calls the script
PluginsThree manifests: .claude-plugin/plugin.json, .codex-plugin/plugin.json with .agents/plugins/marketplace.json, and Cursor’s .cursor-plugin/ layoutThe skills and MCP servers inside the plugin. Codex 0.157.1 can install from a .claude-plugin/marketplace.json, but which Claude-only components then activate was not verified
Managed policy and permissionsmanaged-settings.json for Claude Code, requirements.toml for Codex, the admin dashboard for CursorA tool-neutral intent table, rendered into each file, as described in enforcing one policy across every coding agent
Memory and session historyClaude Code auto memory; Codex memories (off by default); each vendor’s cloud session historyDecisions and lessons written back into AGENTS.md, skills or ADRs. What lives only in memory is lost at exit
Model-specific tuningPrompts, effort settings and model choices tuned to one model familyAn eval suite that tells you, per task, whether the replacement model is good enough. See the models hub for what is current

Apply this table to every new agent feature a team wants to adopt. It gives the CTO a default answer and gives the executive the question to ask. Copy it into your repository as docs/agent-policy/lock-in.md, where the third prompt below expects it.

The feature…Default decisionCondition that keeps the exit cheapQuestion an executive asks
Reads an open format from the repository (AGENTS.md, skills, MCP)AdoptNone beyond the CI check below“Is the source file in Git?”
Adds a vendor wrapper around repository content (plugin, hook registration, cloud setup step)AdoptThe wrapper calls repository scripts and holds no logic of its own“How many lines would we rewrite?”
Holds rules or data only in the vendor’s system (review-bot rules in a dashboard, knowledge in memory)Adopt only with an exportA scheduled export to the repository, checked in the exit rehearsal“Can we export it today, and have we tried?”
Makes a merge or release decisionAdvise onlyYour own required checks decide; the vendor bot comments“What merges if this vendor is down tomorrow?”
Needs a vendor-only model for a production workflowAdopt with a fallbackA tested second model on the same evals, per vendor risk management“What is our degraded mode, and when did we last run it?”

The repository layout is the same for every team. Only the registration commands differ by tool.

repo/
├── AGENTS.md # the one instruction source
├── CLAUDE.md # line 1: @AGENTS.md, then Claude Code-only notes
├── .agents/skills/ # shared skills (Codex, Cursor)
├── .claude/skills/ # generated copy for Claude Code; CI checks it matches
├── .mcp.json # Claude Code project MCP servers
├── .codex/config.toml # Codex project MCP servers, same names
├── docs/specs/ docs/adr/ # specifications and decisions
├── docs/review-rules.md # the criteria every review bot references
├── evals/ # tasks and checks that run against any agent
└── scripts/
├── agent-setup.sh # called by every cloud-agent environment
├── guard-secrets.sh # called by every tool's hook adapter
└── check-portability.sh # the CI check below

Install shared skills for all three tools in one command with the skills CLI (npm skills 1.7.0), run in the repository root:

Terminal window
npx skills add anthropics/skills --skill skill-creator -a claude-code -a codex -a cursor

Put @AGENTS.md on the first line of CLAUDE.md so Claude Code reads the shared rules on both release channels, then register MCP servers at project scope so the file is committed:

Terminal window
claude mcp add --transport http --scope project sentry https://mcp.sentry.dev/mcp

To see what a move from another agent would carry over, Claude Code 2.1.283 has an importer for Codex, Gemini CLI and Cursor configuration. Run it in dry-run mode first:

Terminal window
claude import codex --dry-run

It brings in instruction files, MCP servers, commands, subagents and skills. Hooks, permissions and cloud environments are not in that list, so plan to rebuild them.

A rules generator such as rulesync or Ruler earns its place once three or more agents need the same rules, MCP configuration and skills. With plain rules only, AGENTS.md plus the @AGENTS.md import covers Claude Code and Codex with nothing to regenerate. See one rules source for every agent for the trade-offs.

How do you prove the setup stays portable?

Section titled “How do you prove the setup stays portable?”

A portable layer decays one convenient edit at a time: a rule added to CLAUDE.md only, a skill installed for one tool, an MCP server registered on one laptop. Two checks and one report catch that without anyone reading the config.

  1. Run a drift check on every pull request. This script ran against a test repository on 2026-09-26: it passed on a synchronized layout and failed with both messages after one skill was edited and one MCP server removed. It needs bash, jq and diff.

    #!/usr/bin/env bash
    # scripts/check-portability.sh: fail CI when the portable agent layer drifts between tools.
    set -euo pipefail
    fail=0
    # 1. One instruction source: CLAUDE.md imports AGENTS.md instead of forking it.
    if [ -f CLAUDE.md ] && [ "$(head -n1 CLAUDE.md)" != '@AGENTS.md' ]; then
    echo "CLAUDE.md does not start with @AGENTS.md"; fail=1
    fi
    # 2. One skill set: the Claude Code copy matches the shared .agents/skills copy.
    if ! diff -r .agents/skills .claude/skills > /dev/null; then
    echo "Skills differ between .agents/skills and .claude/skills"; fail=1
    fi
    # 3. One MCP list: the same server names for Claude Code and Codex.
    claude_mcp=$(jq -r '.mcpServers | keys[]' .mcp.json | sort | xargs)
    codex_mcp=$(sed -nE 's/^\[mcp_servers\.([A-Za-z0-9_-]+)\]$/\1/p' .codex/config.toml | sort | xargs)
    if [ "$claude_mcp" != "$codex_mcp" ]; then
    echo "MCP servers differ: .mcp.json [$claude_mcp] vs .codex/config.toml [$codex_mcp]"; fail=1
    fi
    [ "$fail" -eq 0 ] && echo "Portable layer in sync"
    exit "$fail"
  2. Run the same eval suite against two agents every month. Keep the tasks in evals/ and run them through a harness with providers for more than one agent. promptfoo 0.123.1 lists anthropic:claude-agent-sdk and openai:codex-sdk as providers, so one promptfooconfig.yaml produces a side-by-side pass rate. Note that promptfoo’s README states it is now part of OpenAI and remains MIT-licensed; keep the task definitions in plain files so the harness itself can be swapped. The evals for coding agents page covers task design.

  3. Report the trend, not the snapshot. The platform team reports two numbers each quarter: the drift check’s failure count and the second agent’s pass rate on the eval suite. A widening gap is early warning that your harness is tuning itself to one vendor.

The platform team owns the check and the eval run; the CTO signs off on the quarterly numbers, as set out in the operating model.

An exit rehearsal is a timed exercise in which one team delivers real work with the secondary agent, using only what is in the repository. It turns “we could switch” into a measured cost. Run it once a quarter, or before any contract renewal.

  1. Pick the scope. One team of three to five engineers, one working day, two or three backlog items of normal size with acceptance criteria already written. Choose items that touch your MCP servers and at least one skill.
  2. Freeze the primary tool for the rehearsal. The team works only in the secondary agent. They may not copy settings from the primary tool’s user directory; everything must come from Git or from a documented install command.
  3. Rebuild only what the repository cannot provide, and time it. Log every minute spent on setup: MCP registration, skill installs, hook adapters, the cloud environment, review-bot configuration.
  4. Deliver through the normal gates. Each item goes through the same required CI checks, eval suite and review route as any other change. No gate is waived for the rehearsal.
  5. Export what the vendor holds. Try to export review-bot rules, cloud-environment settings and any memory or session history the team relies on. Record what could not be exported.
  6. Record the result and set the next action. Fill in the template below. Every row marked “rebuilt by hand” becomes a backlog item to move that substance into the repository.

Record each rehearsal in the policy repository with this template:

FieldRecord
Date, team, primary agent and version, secondary agent and version
Items attempted / merged through normal gates
Setup time before the first agent task started (minutes)
Eval-suite pass rate: primary agent / secondary agent
Assets that worked unchanged (instructions, skills, MCP, specs, evals)
Assets rebuilt by hand, with minutes each
Assets that could not be exported from the vendor
Gates that depended on a vendor bot
Follow-up items, owners and due dates

The rehearsal passes when all four hold: every attempted item merged through the normal gates, setup took less than half a working day, the secondary agent’s eval pass rate is within the tolerance you set in advance, and no gate depended on the primary vendor’s bot. A failed rehearsal is a useful result: it names the next items for the backlog. The CTO signs off the record, and the executive reads the setup time and the follow-up list as the current switching cost.

Copy-paste prompts for a portability audit

Section titled “Copy-paste prompts for a portability audit”

What breaks when you try to switch agent vendors?

Section titled “What breaks when you try to switch agent vendors?”

The second tool ignores your instructions. Claude Code on the stable channel (2.1.274 on 2026-09-26) does not fall back to AGENTS.md, and even on latest it reads only CLAUDE.md when both files exist. Recovery: put @AGENTS.md on the first line of CLAUDE.md and let the drift check enforce it.

Skills load but never fire. A skill copied with npx skills add lacks the session-start hook that its plugin version installs, or it delegates to a sibling skill you did not copy. Recovery: install the whole skill set, or the vendor’s plugin version where one exists, and add a smoke task to the eval suite that must trigger the skill.

An MCP server connects in one tool and fails in another. The server name differs between .mcp.json and .codex/config.toml, or the server supports only the 2026-07-28 protocol, which Codex 0.157.1 does not negotiate. Recovery: keep names identical (the drift check does this) and test each server in every client before you add it to the approved list.

Review coverage silently disappears. The old vendor’s review bot was the only reviewer for a class of change, and nobody noticed until it was switched off. Recovery: move review criteria into docs/review-rules.md, make the merge decision a required CI check, and use the evidence bundle so a pull request is mergeable without a single vendor’s verdict.

The cloud agent cannot build the repository. The setup lived in a vendor’s environment form, not in Git. Recovery: move every setup step into scripts/agent-setup.sh, and have each vendor’s environment call only that script.

Knowledge leaves with the vendor. Decisions lived in agent memory or session history. Recovery: write lessons back into AGENTS.md, skills or ADRs as part of the definition of done, and export session history before a contract ends.

The new model underperforms on your work. Prompts and effort settings were tuned to the old model. Recovery: compare both agents on the eval suite before you switch, and set the tolerance in advance so the decision is not argued after the fact.

Where to go next with lock-in and portability

Section titled “Where to go next with lock-in and portability”