Tooling policy — approved workflows and measured exceptions
An AI tooling policy names the approved service and configuration for each engineering workflow, the data and authority boundary of each, and the owner who signs it off. The CTO Scorecard Q2 maximum needs an approved-workflow matrix, controlled default services, and time-bounded exceptions that end with evidence and cleanup. A list of approved logos does not earn it.
This page is for the CTO or VP Engineering who owns AI tooling decisions, and for the platform or DevEx lead who maintains the policy. The usual starting point: the wiki says “we use Cursor”, the platform team runs Claude Code on company seats, two engineers use Codex through personal ChatGPT plans, and nobody reviewed the MCP servers that read the issue tracker. Every new tool request becomes a debate, and nothing that was tried ever gets switched off. Before you start, read the CTO scorecard answer key for where Q2 sits among the other questions, and the AI usage policy, which defines the data classes this page refers to.
What the tooling policy gives you
Section titled “What the tooling policy gives you”- An approved-workflow matrix that answers “which tool, with which data, allowed to change what” for each common task.
- A service record per tool and plan, dated, that security and finance can check against the real configuration.
- Default tools enforced by managed settings, not by a wiki page.
- An exception record and a CI check that expires every exception unless its evidence passes a decision gate.
- A proof table your reviewers can sign off without auditing every laptop.
Which Q2 answer describes your organization?
Section titled “Which Q2 answer describes your organization?”The CTO scorecard asks “What’s the tooling policy (Claude Code / Cursor / Codex / others)?” and offers four answers. The table shows what each one leaves open.
| Answer | Score | What stays ungoverned |
|---|---|---|
| Shadow AI: everyone picks their own, no list | 0 | Data routes, access, support, cost, and incident response |
| A “preferred” tool, but in practice a mix | 1 | Everything outside the preferred tool, which is where the risk accumulates |
| A short list with explicit “when to use which” | 2 | Enforcement, data boundaries per workflow, and a way to try or retire a tool |
| Approved workflow matrix, controlled default services, and time-bounded exceptions with evidence and cleanup | 3 | Nothing structural, provided the proof table below passes |
The step from 2 to 3 is the move from a recommendation to a control. A short list tells engineers what to prefer. A policy at level 3 decides what each tool may touch, enforces the default on company machines, and gives every other tool a dated path in or out. Google’s DORA team names a “clear and communicated AI stance” as the first of seven capabilities in its AI Capabilities Model (Google Cloud blog, 2025-09-23), and its follow-up puts the reason plainly: “Ambiguity creates risk. A clear policy provides the psychological safety developers need to experiment effectively” (Google Cloud blog, 2025-12-10).
Build the tooling policy in six steps
Section titled “Build the tooling policy in six steps”-
Inventory what engineers actually run. Cover IDE extensions, CLIs, desktop apps, cloud agents, CI actions, MCP servers, plugins and skills, API keys, and personal plans. The next section shows how to collect it per tool. Mark every unknown instead of guessing.
-
Classify the workflows. Group real use into five to eight workflows: interactive repository work, PR review, CI repair, cloud agents that open PRs, prototypes, internal-system access through MCP, and production operation. A workflow, not a tool, is the unit you approve.
-
Write a service record for each tool and plan. Record product, plan, identity controls, data retention, region, allowed models, MCP and plugin sources, version floor, owner, and the date you verified it. The template is below. The plan matters more than the brand: Codex is included in the Plus, Pro, Business, Edu, and Enterprise ChatGPT plans, so “Codex is approved” says nothing until you name the workspace.
-
Choose the defaults and enforce them. Pick one default per workflow, and a second only where the workflows differ materially. Then encode the default in each tool’s managed settings, so the policy holds on a laptop nobody audits.
-
Open the exception path. Any tool outside the matrix enters through an exception record with a hypothesis, a data class, an owner, a cost cap, and an expiry of 90 days or less. A CI check fails when a record expires without a decision.
-
Review on a cadence and retire. Every quarter, re-run the inventory, re-verify each service record, close expired exceptions, and remove tools with no completed work. Re-verify immediately when a vendor changes a plan, a default model, or a data-handling feature.
How do you inventory the AI tools engineers already use?
Section titled “How do you inventory the AI tools engineers already use?”Start with the machines, because expense claims and the IdP only show what someone paid for or signed in to. The commands below read each tool’s own view of its configuration. The same inventory works in every repository, but MCP servers can be configured per project, so run it from each repository root, not once per laptop.
claude --version # compare with the version floor in the service recordclaude mcp list # every MCP server, including plugin-provided and project .mcp.json onesclaude plugin list --json # installed pluginsclaude doctor # installation health; reads settings in the current directoryclaude mcp list health-checks every approved server, which starts local stdio servers such as npx or uvx commands. Run it only in repositories you trust. Servers from an unapproved project .mcp.json are listed as pending approval and are not started. Inside a session, /status shows the setting sources, including whether managed settings are loaded. Checked against Claude Code 2.1.283.
codex --version # compare with the version floor in the service recordcodex mcp list --json # configured MCP servers with transport and auth statuscodex plugin list --json # installed plugins (add --available to include uninstalled marketplace plugins)codex doctor # installation, config, auth, and runtime healthIn the Codex TUI, /debug-config shows each config layer and which source set each requirement, which tells you whether the managed requirements.toml reached the machine. Checked against Codex CLI 0.157.1.
Cursor keeps team membership, usage, and settings in its admin dashboard. cursor.com was unreachable from the environment this page was checked in on 2026-09-26, so no Cursor command or setting name appears here as fact. Close these questions in Cursor’s current admin documentation and date the answers:
- Which usage report lists active members, and can you export it?
- Can you list the MCP servers and extensions a member has configured, or only enforce an allowed set?
- Which privacy setting applies to team members, and can members change it?
Collect the output with a script that each engineer runs, or that your device-management tool runs, and commit the results to the policy repository. Add the organization-wide signals the machines cannot show: expense claims for AI plans, SSO sign-ins to AI vendors, GitHub and GitLab apps installed on your organization, and CI workflows that call an agent action.
#!/usr/bin/env bash# ai-inventory.sh: run from a repository root; prints the agent tooling this machine and repo configureset -uecho "== host: $(hostname) repo: $(basename "$PWD") date: $(date +%F)"for cli in claude codex; do command -v "$cli" >/dev/null && "$cli" --version; doneif command -v claude >/dev/null; then echo "== claude mcp list"; claude mcp list echo "== claude plugin list"; claude plugin list --jsonfiif command -v codex >/dev/null; then echo "== codex mcp list"; codex mcp list --json echo "== codex plugin list"; codex plugin list --jsonfiThe approved-workflow matrix
Section titled “The approved-workflow matrix”This is the artifact that earns the Q2 point. Copy it into your policy repository and replace the example values. Each row approves a workflow with a boundary, not a product.
| Workflow | Default service and configuration | Highest data class | Authority | Required evidence | Owner |
|---|---|---|---|---|---|
| Interactive repository work | Claude Code, Codex, or Cursor on company seats, managed policy v1.4 | Internal source | Local branch; no push to protected branches | Plan, diff, passing checks | DevEx lead |
| PR review | Approved review agent on the Git host | Internal source | Comments only | Review comments linked to the PR | Platform lead |
| CI repair and headless runs | anthropics/claude-code-action or openai/codex-action, pinned to a full commit SHA of a v1 release, with read-only tools unless the job must write | Internal source, no production secrets | Draft PR | CI logs, the evidence bundle | Platform lead |
| Cloud agents that open PRs | One approved cloud agent per Git organization | Internal source | Draft PR behind required checks | Same checks as a human PR | Platform lead |
| Prototypes | Any default tool in a sandbox repository | Synthetic or public data | No production deploy | Graduation review | Product engineering manager |
| Internal systems through MCP | Servers on the MCP allowlist only | As approved per server | Read-only unless the server record says otherwise | Server record and audit log | Security lead |
| Production operation | No direct coding-agent authority by default | Production policy | A named human gate | Release and rollback record | Service owner |
Three rules keep the matrix short and honest:
- Approve the workflow boundary first, then the tools that fit it. Claude Code, Codex, and Cursor can all sit in the first row, because the row’s controls, not the vendor, carry the risk. The primary harness guide shows how a developer proves a setup meets the row.
- Link, do not copy, volatile facts. Plan prices live on the pricing analysis and model versions on the models hub. The service record stores the date you verified them.
- Let the matrix point at the rules, not restate them. Data classes and acceptable use come from the AI usage policy, retention and residency from your data privacy policy, and MCP authorization from MCP security.
Service record template
Section titled “Service record template”One file per tool and plan, in the same repository as the matrix. The record is what a reviewer compares with the real configuration, so every field must be checkable.
service: Claude Codeplan: Claude Team, Premium seats for platform engineers, Standard for everyone elsecontract_owner: head-of-platform@example.comsurfaces: [CLI, VS Code extension, desktop app] # anything not listed is not approvedidentity: { sso: required, login_binding: forceLoginOrgUUID in managed settings }data: { retention: FROM_CONTRACT, training_use: FROM_CONTRACT, region: FROM_CONTRACT }models: governed by availableModels in managed policy v1.4; default follows the models hubmcp_allowlist: policy/mcp-allowlist.yamlplugin_sources: [anthropics/claude-plugins-official, acme/agent-plugins]version_floor: "2.1.274" # the stable release channel on 2026-09-26workflows_approved: [interactive-repo-work, ci-repair]verified_on: 2026-09-26next_review: 2026-12-26exit_plan: export settings, skills, and hooks from the repository; revoke seats and API keysThe FROM_CONTRACT values are placeholders, not a statement of Anthropic’s terms: fill them from your contract and the vendor’s current plan documentation, and cite the document and its date. The full questionnaire for a new vendor is in buying AI coding tools.
Enforce the default tools on company machines
Section titled “Enforce the default tools on company machines”A policy that lives only in a wiki is advice. Each tool reads a managed file that users cannot override, and that file is where the matrix becomes a control. The complete cross-tool file, with every key explained, is on enforcing one policy across every coding agent. The tabs below show only the keys that implement the matrix.
Deploy through the claude.ai admin console, MDM, or /etc/claude-code/managed-settings.json on Linux.
{ "forceLoginMethod": "claudeai", "forceLoginOrgUUID": "YOUR_ORG_UUID", "availableModels": ["opus", "sonnet"], "enforceAvailableModels": true, "allowedMcpServers": [ { "serverUrl": "https://api.githubcopilot.com/*" }, { "serverCommand": ["npx", "@playwright/mcp@0.0.82"] } ], "allowManagedMcpServersOnly": true, "strictKnownMarketplaces": [ { "source": "github", "repo": "anthropics/claude-plugins-official" }, { "source": "github", "repo": "acme/agent-plugins" } ], "requiredMinimumVersion": "2.1.274"}forceLoginMethod with forceLoginOrgUUID closes the personal-plan route: a personal Pro or Max account cannot sign in. enforceAvailableModels makes the Default model option obey the list, and allowManagedMcpServersOnly stops users from adding servers. A serverCommand entry must match the configured command exactly, so pin the package version: an @latest entry approves whatever npm resolves on the day. The version shown is the current @playwright/mcp release on 2026-09-26. Release channels matter for the version floor: on 2026-09-26 the latest channel was 2.1.283 and stable was 2.1.274, and the two started sessions on different default models. Record which channel the service record approves.
Put constraints in requirements.toml (/etc/codex/requirements.toml on Linux, or delivered by MDM or the workspace), never in config.toml, which holds user defaults.
allowed_login_methods = ["chatgpt"]allowed_chatgpt_workspaces = ["CHATGPT_WORKSPACE_ID"]
[mcp_servers.github.identity]url = "https://api.githubcopilot.com/mcp/"
[marketplaces]restrict_to_allowed_sources = true
[marketplaces.allowed_sources.acme]source = "git"url = "https://github.com/acme/agent-plugins.git"
[models.new_thread]model = "MODEL_FROM_THE_MODELS_HUB"allowed_login_methods = ["chatgpt"] with a workspace ID closes the personal-plan route. MCP requirements match the server name and identity together, so publish the exact names. With restrict_to_allowed_sources = true, only marketplaces listed under allowed_sources load, so list every source the service record approves, or plugins are blocked without a message to the engineer. requirements.toml has no model allowlist key in Codex CLI 0.157.1 (checked against config_requirements.rs on 2026-09-26); it can replace the model catalog with model_catalog_json, which shapes the picker but is not documented as a hard allowlist. It can pin the default for new threads with [models.new_thread] model = "…" (also model_reasoning_effort and service_tier) and constrain model_provider; for an allowlist, use the workspace or a gateway as described in where the model runs.
Cursor applies team settings from its admin dashboard, and the @cursor/sdk 1.0.32 package lists team and mdm setting sources, so team-level and MDM-delivered settings exist. The individual setting names could not be verified on 2026-09-26. Before you mark Cursor as a default in the matrix, confirm in Cursor’s admin documentation that you can restrict sign-in to your team, restrict MCP servers and models, and enforce the privacy setting, and record each answer with its date.
Exception record template
Section titled “Exception record template”An exception is a small, funded experiment, not a permission slip. It names what it tries to prove, what it may touch, what it costs, and when it ends. The example tests Codex cloud tasks for a team whose default is a local agent.
id: EX-2026-014owner: payments-lead@example.comtool: Codex cloud tasksplan: company ChatGPT Business workspace (no personal plans)hypothesis: > Codex cloud tasks take a triaged dependency-bump issue to a green draft PR within one working day for the payments service, with no rise in reverted PRs.representative_tasks: 20 dependency bumps and 10 flaky-test fixes from the current backlogdata_class: internal-source # no customer data, no production credentialsauthority: draft PR only; required checks and a human approval before mergecontrols: [company SSO, workspace-bound login, MCP allowlist v1.4, no production secrets]cost_cap_usd: 1500start: 2026-10-01expires: 2026-11-30success_gate: 21 of the 30 tasks merged without rework; revert rate at or below the team baselinestop_gate: any data-class breach, or fewer than 12 tasks merged by 2026-11-01evidence: policy/evidence/EX-2026-014/ # PR links, CI results, invoice linescleanup: [remove the GitHub app from repositories, delete environment secrets, remove seats, remove local config]decision: pending # pending | graduate | extend-once | stopThe thresholds are examples; set yours from the team’s own baseline. Design the trial itself with pilot design so the result means something, and fund anything larger than one team as a bet on the AI tooling roadmap.
Run this check in CI on the policy repository, on every pull request and on a daily schedule. It fails when a record is incomplete, has a decision outside the four allowed values, runs longer than 90 days, expires without a decision, or stops without verified cleanup. An extension is a new record whose extended_from names the original; the check refuses to extend that record again, so extend-once holds without relying on the reviewer.
#!/usr/bin/env python3"""Fail CI when a tooling exception is incomplete, too long, or expired without a decision.
Usage: python3 check_exceptions.py policy/exceptions/*.yaml (requires PyYAML)"""import sysfrom datetime import date
import yaml
REQUIRED = ["id", "owner", "tool", "hypothesis", "data_class", "authority", "start", "expires", "success_gate", "stop_gate", "cleanup"]MAX_DAYS = 90DECISIONS = {"pending", "graduate", "extend-once", "stop"}
failures = []for path in sys.argv[1:]: with open(path) as f: rec = yaml.safe_load(f) or {} missing = [k for k in REQUIRED if not rec.get(k)] if missing: failures.append(f"{path}: missing {', '.join(missing)}") continue start, expires = rec["start"], rec["expires"] if not isinstance(start, date) or not isinstance(expires, date): failures.append(f"{path}: start and expires must be YYYY-MM-DD dates") continue decision = rec.get("decision", "pending") if decision not in DECISIONS: failures.append(f"{path}: decision {decision!r} is not one of {sorted(DECISIONS)}") continue if decision == "extend-once" and rec.get("extended_from"): failures.append(f"{path}: already an extension of {rec['extended_from']}; decide graduate or stop") if (expires - start).days > MAX_DAYS: failures.append(f"{path}: runs {(expires - start).days} days (limit {MAX_DAYS})") if expires < date.today() and decision == "pending": failures.append(f"{path}: expired {expires} with no decision") if decision == "stop" and not rec.get("cleanup_verified"): failures.append(f"{path}: stopped but cleanup_verified is not set")
print("\n".join(failures) or f"{len(sys.argv) - 1} exception records OK")sys.exit(1 if failures else 0)Default, specialist exception, or retirement?
Section titled “Default, specialist exception, or retirement?”When two tools overlap, decide on evidence from completed work, not on enthusiasm. Score both on the same representative tasks and apply the first row that matches.
| Evidence | Decision |
|---|---|
| The tool fails a control the workflow row requires (login binding, MCP allowlist, data terms) | Not approvable for that row, whatever its results |
| It completes the representative tasks no better than the default | Retire, or never admit |
| It is clearly better for one workflow and no worse on controls | Specialist default for that workflow only |
| It is better across most workflows, with equal controls and acceptable cost | Candidate to replace the default, as a roadmap bet with a migration plan |
Portability lowers the cost of every row. Context in AGENTS.md, skills in the open Agent Skills format, and checks in CI move with you between tools. Codex reads AGENTS.md natively, and Claude Code reads it when a project has no CLAUDE.md (v2.1.277 and later, latest channel). The rest of the checklist is on avoiding lock-in.
Copy-paste prompts for tooling policy work
Section titled “Copy-paste prompts for tooling policy work”Run these in Claude Code, Codex, or Cursor with the policy repository and the inventory output in the working directory. They work the same way in all three tools.
How do you prove the tooling policy is in force?
Section titled “How do you prove the tooling policy is in force?”You do not need to inspect every laptop. You need checks that fail loudly when the policy drifts, and a named person who signs each one off.
| Control | Evidence | Pass condition | Signs off |
|---|---|---|---|
| Matrix coverage | Quarterly inventory gap report | Every workflow in use maps to a matrix row | CTO |
| Service records | One dated record per tool and plan | Each record verified within the last 90 days and matching the admin console | Platform lead |
| Default enforcement | A test on a managed machine: add an unlisted MCP server, sign in with a personal account | Both attempts are rejected, and the rejection is recorded | Security lead |
| Exceptions | check_exceptions.py in CI | Green on the default branch; no expired record without a decision | CTO |
| Retirement | Cleanup evidence for each stopped exception or retired tool | Seats, apps, secrets, and local configuration removed | IT or identity owner |
| Engineer clarity | Five engineers asked which tool and data class a common task uses | All five answer from the matrix | DevEx lead |
Adoption and outcome evidence for the approved tools belongs to the AI metrics panel. The policy’s own job is narrower: showing that every tool in use is approved for its workflow and that nothing expired stays live.
What breaks in an AI tooling policy?
Section titled “What breaks in an AI tooling policy?”Approving logos instead of workflows. “Codex is approved” lets an engineer use it on a personal Plus plan with none of your data terms. Recovery: rewrite each approval as a matrix row plus a service record that names the plan and workspace, then bind logins with managed settings, as described in team accounts.
Exceptions that never end. A pilot from spring is still running in autumn, with seats, a GitHub app, and secrets nobody owns. Recovery: backfill an exception record for each one with an expiry within 30 days, run check_exceptions.py in CI, and treat any stopped exception without cleanup_verified as open work.
A single mandated tool. The policy forces one tool, a team with a real need goes around it, and the shadow use returns. Recovery: keep one default per workflow, and route the need through an exception so it can prove itself on evidence.
Service records that go stale. A vendor changes a plan, a default model, or a retention rule, and the record still shows the old terms. For example, Claude Code’s two release channels started sessions on different default models on 2026-09-26. Recovery: date every record, re-verify on a 90-day cadence and on every vendor announcement, and pin the release channel in the record.
Unpinned extensions and CLIs. Approved tools ship through registries that have been compromised. A malicious script was injected into the Amazon Q Developer extension for VS Code, version 1.84.0 (AWS advisory, 2025-07-26), and an unauthorized npm publish of cline@2.3.0 carried a modified postinstall script (Cline advisory, 2026-02-17). Recovery: set a version floor in managed settings where the tool supports one, install from an internal mirror or a reviewed update channel, and add registry advisories for approved tools to the security team’s watch list.
No exit plan. A tool disappears and the team loses its context. The Roo Code extension shut down in May 2026, per its own README. Recovery: keep context, skills, and checks in the repository in portable formats, write an exit_plan in every service record, and rehearse it for your riskiest vendor with vendor risk management.
Where to go next with the tooling policy
Section titled “Where to go next with the tooling policy”The policy decides what is approved. Account governance makes the approval real, and the roadmap decides what gets tried next.