Skip to content

Tooling policy — approved workflows and measured exceptions

An AI tooling policy names the approved service and configuration for each engineering workflow, the data and authority boundary of each, and the owner who signs it off. The CTO Scorecard Q2 maximum needs an approved-workflow matrix, controlled default services, and time-bounded exceptions that end with evidence and cleanup. A list of approved logos does not earn it.

This page is for the CTO or VP Engineering who owns AI tooling decisions, and for the platform or DevEx lead who maintains the policy. The usual starting point: the wiki says “we use Cursor”, the platform team runs Claude Code on company seats, two engineers use Codex through personal ChatGPT plans, and nobody reviewed the MCP servers that read the issue tracker. Every new tool request becomes a debate, and nothing that was tried ever gets switched off. Before you start, read the CTO scorecard answer key for where Q2 sits among the other questions, and the AI usage policy, which defines the data classes this page refers to.

  • An approved-workflow matrix that answers “which tool, with which data, allowed to change what” for each common task.
  • A service record per tool and plan, dated, that security and finance can check against the real configuration.
  • Default tools enforced by managed settings, not by a wiki page.
  • An exception record and a CI check that expires every exception unless its evidence passes a decision gate.
  • A proof table your reviewers can sign off without auditing every laptop.

Which Q2 answer describes your organization?

Section titled “Which Q2 answer describes your organization?”

The CTO scorecard asks “What’s the tooling policy (Claude Code / Cursor / Codex / others)?” and offers four answers. The table shows what each one leaves open.

AnswerScoreWhat stays ungoverned
Shadow AI: everyone picks their own, no list0Data routes, access, support, cost, and incident response
A “preferred” tool, but in practice a mix1Everything outside the preferred tool, which is where the risk accumulates
A short list with explicit “when to use which”2Enforcement, data boundaries per workflow, and a way to try or retire a tool
Approved workflow matrix, controlled default services, and time-bounded exceptions with evidence and cleanup3Nothing structural, provided the proof table below passes

The step from 2 to 3 is the move from a recommendation to a control. A short list tells engineers what to prefer. A policy at level 3 decides what each tool may touch, enforces the default on company machines, and gives every other tool a dated path in or out. Google’s DORA team names a “clear and communicated AI stance” as the first of seven capabilities in its AI Capabilities Model (Google Cloud blog, 2025-09-23), and its follow-up puts the reason plainly: “Ambiguity creates risk. A clear policy provides the psychological safety developers need to experiment effectively” (Google Cloud blog, 2025-12-10).

  1. Inventory what engineers actually run. Cover IDE extensions, CLIs, desktop apps, cloud agents, CI actions, MCP servers, plugins and skills, API keys, and personal plans. The next section shows how to collect it per tool. Mark every unknown instead of guessing.

  2. Classify the workflows. Group real use into five to eight workflows: interactive repository work, PR review, CI repair, cloud agents that open PRs, prototypes, internal-system access through MCP, and production operation. A workflow, not a tool, is the unit you approve.

  3. Write a service record for each tool and plan. Record product, plan, identity controls, data retention, region, allowed models, MCP and plugin sources, version floor, owner, and the date you verified it. The template is below. The plan matters more than the brand: Codex is included in the Plus, Pro, Business, Edu, and Enterprise ChatGPT plans, so “Codex is approved” says nothing until you name the workspace.

  4. Choose the defaults and enforce them. Pick one default per workflow, and a second only where the workflows differ materially. Then encode the default in each tool’s managed settings, so the policy holds on a laptop nobody audits.

  5. Open the exception path. Any tool outside the matrix enters through an exception record with a hypothesis, a data class, an owner, a cost cap, and an expiry of 90 days or less. A CI check fails when a record expires without a decision.

  6. Review on a cadence and retire. Every quarter, re-run the inventory, re-verify each service record, close expired exceptions, and remove tools with no completed work. Re-verify immediately when a vendor changes a plan, a default model, or a data-handling feature.

How do you inventory the AI tools engineers already use?

Section titled “How do you inventory the AI tools engineers already use?”

Start with the machines, because expense claims and the IdP only show what someone paid for or signed in to. The commands below read each tool’s own view of its configuration. The same inventory works in every repository, but MCP servers can be configured per project, so run it from each repository root, not once per laptop.

Terminal window
claude --version # compare with the version floor in the service record
claude mcp list # every MCP server, including plugin-provided and project .mcp.json ones
claude plugin list --json # installed plugins
claude doctor # installation health; reads settings in the current directory

claude mcp list health-checks every approved server, which starts local stdio servers such as npx or uvx commands. Run it only in repositories you trust. Servers from an unapproved project .mcp.json are listed as pending approval and are not started. Inside a session, /status shows the setting sources, including whether managed settings are loaded. Checked against Claude Code 2.1.283.

Collect the output with a script that each engineer runs, or that your device-management tool runs, and commit the results to the policy repository. Add the organization-wide signals the machines cannot show: expense claims for AI plans, SSO sign-ins to AI vendors, GitHub and GitLab apps installed on your organization, and CI workflows that call an agent action.

#!/usr/bin/env bash
# ai-inventory.sh: run from a repository root; prints the agent tooling this machine and repo configure
set -u
echo "== host: $(hostname) repo: $(basename "$PWD") date: $(date +%F)"
for cli in claude codex; do command -v "$cli" >/dev/null && "$cli" --version; done
if command -v claude >/dev/null; then
echo "== claude mcp list"; claude mcp list
echo "== claude plugin list"; claude plugin list --json
fi
if command -v codex >/dev/null; then
echo "== codex mcp list"; codex mcp list --json
echo "== codex plugin list"; codex plugin list --json
fi

This is the artifact that earns the Q2 point. Copy it into your policy repository and replace the example values. Each row approves a workflow with a boundary, not a product.

WorkflowDefault service and configurationHighest data classAuthorityRequired evidenceOwner
Interactive repository workClaude Code, Codex, or Cursor on company seats, managed policy v1.4Internal sourceLocal branch; no push to protected branchesPlan, diff, passing checksDevEx lead
PR reviewApproved review agent on the Git hostInternal sourceComments onlyReview comments linked to the PRPlatform lead
CI repair and headless runsanthropics/claude-code-action or openai/codex-action, pinned to a full commit SHA of a v1 release, with read-only tools unless the job must writeInternal source, no production secretsDraft PRCI logs, the evidence bundlePlatform lead
Cloud agents that open PRsOne approved cloud agent per Git organizationInternal sourceDraft PR behind required checksSame checks as a human PRPlatform lead
PrototypesAny default tool in a sandbox repositorySynthetic or public dataNo production deployGraduation reviewProduct engineering manager
Internal systems through MCPServers on the MCP allowlist onlyAs approved per serverRead-only unless the server record says otherwiseServer record and audit logSecurity lead
Production operationNo direct coding-agent authority by defaultProduction policyA named human gateRelease and rollback recordService owner

Three rules keep the matrix short and honest:

  • Approve the workflow boundary first, then the tools that fit it. Claude Code, Codex, and Cursor can all sit in the first row, because the row’s controls, not the vendor, carry the risk. The primary harness guide shows how a developer proves a setup meets the row.
  • Link, do not copy, volatile facts. Plan prices live on the pricing analysis and model versions on the models hub. The service record stores the date you verified them.
  • Let the matrix point at the rules, not restate them. Data classes and acceptable use come from the AI usage policy, retention and residency from your data privacy policy, and MCP authorization from MCP security.

One file per tool and plan, in the same repository as the matrix. The record is what a reviewer compares with the real configuration, so every field must be checkable.

policy/services/claude-code-team.yaml
service: Claude Code
plan: Claude Team, Premium seats for platform engineers, Standard for everyone else
contract_owner: head-of-platform@example.com
surfaces: [CLI, VS Code extension, desktop app] # anything not listed is not approved
identity: { sso: required, login_binding: forceLoginOrgUUID in managed settings }
data: { retention: FROM_CONTRACT, training_use: FROM_CONTRACT, region: FROM_CONTRACT }
models: governed by availableModels in managed policy v1.4; default follows the models hub
mcp_allowlist: policy/mcp-allowlist.yaml
plugin_sources: [anthropics/claude-plugins-official, acme/agent-plugins]
version_floor: "2.1.274" # the stable release channel on 2026-09-26
workflows_approved: [interactive-repo-work, ci-repair]
verified_on: 2026-09-26
next_review: 2026-12-26
exit_plan: export settings, skills, and hooks from the repository; revoke seats and API keys

The FROM_CONTRACT values are placeholders, not a statement of Anthropic’s terms: fill them from your contract and the vendor’s current plan documentation, and cite the document and its date. The full questionnaire for a new vendor is in buying AI coding tools.

Enforce the default tools on company machines

Section titled “Enforce the default tools on company machines”

A policy that lives only in a wiki is advice. Each tool reads a managed file that users cannot override, and that file is where the matrix becomes a control. The complete cross-tool file, with every key explained, is on enforcing one policy across every coding agent. The tabs below show only the keys that implement the matrix.

Deploy through the claude.ai admin console, MDM, or /etc/claude-code/managed-settings.json on Linux.

{
"forceLoginMethod": "claudeai",
"forceLoginOrgUUID": "YOUR_ORG_UUID",
"availableModels": ["opus", "sonnet"],
"enforceAvailableModels": true,
"allowedMcpServers": [
{ "serverUrl": "https://api.githubcopilot.com/*" },
{ "serverCommand": ["npx", "@playwright/mcp@0.0.82"] }
],
"allowManagedMcpServersOnly": true,
"strictKnownMarketplaces": [
{ "source": "github", "repo": "anthropics/claude-plugins-official" },
{ "source": "github", "repo": "acme/agent-plugins" }
],
"requiredMinimumVersion": "2.1.274"
}

forceLoginMethod with forceLoginOrgUUID closes the personal-plan route: a personal Pro or Max account cannot sign in. enforceAvailableModels makes the Default model option obey the list, and allowManagedMcpServersOnly stops users from adding servers. A serverCommand entry must match the configured command exactly, so pin the package version: an @latest entry approves whatever npm resolves on the day. The version shown is the current @playwright/mcp release on 2026-09-26. Release channels matter for the version floor: on 2026-09-26 the latest channel was 2.1.283 and stable was 2.1.274, and the two started sessions on different default models. Record which channel the service record approves.

An exception is a small, funded experiment, not a permission slip. It names what it tries to prove, what it may touch, what it costs, and when it ends. The example tests Codex cloud tasks for a team whose default is a local agent.

policy/exceptions/EX-2026-014.yaml
id: EX-2026-014
owner: payments-lead@example.com
tool: Codex cloud tasks
plan: company ChatGPT Business workspace (no personal plans)
hypothesis: >
Codex cloud tasks take a triaged dependency-bump issue to a green draft PR within
one working day for the payments service, with no rise in reverted PRs.
representative_tasks: 20 dependency bumps and 10 flaky-test fixes from the current backlog
data_class: internal-source # no customer data, no production credentials
authority: draft PR only; required checks and a human approval before merge
controls: [company SSO, workspace-bound login, MCP allowlist v1.4, no production secrets]
cost_cap_usd: 1500
start: 2026-10-01
expires: 2026-11-30
success_gate: 21 of the 30 tasks merged without rework; revert rate at or below the team baseline
stop_gate: any data-class breach, or fewer than 12 tasks merged by 2026-11-01
evidence: policy/evidence/EX-2026-014/ # PR links, CI results, invoice lines
cleanup: [remove the GitHub app from repositories, delete environment secrets, remove seats, remove local config]
decision: pending # pending | graduate | extend-once | stop

The thresholds are examples; set yours from the team’s own baseline. Design the trial itself with pilot design so the result means something, and fund anything larger than one team as a bet on the AI tooling roadmap.

Run this check in CI on the policy repository, on every pull request and on a daily schedule. It fails when a record is incomplete, has a decision outside the four allowed values, runs longer than 90 days, expires without a decision, or stops without verified cleanup. An extension is a new record whose extended_from names the original; the check refuses to extend that record again, so extend-once holds without relying on the reviewer.

#!/usr/bin/env python3
"""Fail CI when a tooling exception is incomplete, too long, or expired without a decision.
Usage: python3 check_exceptions.py policy/exceptions/*.yaml (requires PyYAML)
"""
import sys
from datetime import date
import yaml
REQUIRED = ["id", "owner", "tool", "hypothesis", "data_class", "authority",
"start", "expires", "success_gate", "stop_gate", "cleanup"]
MAX_DAYS = 90
DECISIONS = {"pending", "graduate", "extend-once", "stop"}
failures = []
for path in sys.argv[1:]:
with open(path) as f:
rec = yaml.safe_load(f) or {}
missing = [k for k in REQUIRED if not rec.get(k)]
if missing:
failures.append(f"{path}: missing {', '.join(missing)}")
continue
start, expires = rec["start"], rec["expires"]
if not isinstance(start, date) or not isinstance(expires, date):
failures.append(f"{path}: start and expires must be YYYY-MM-DD dates")
continue
decision = rec.get("decision", "pending")
if decision not in DECISIONS:
failures.append(f"{path}: decision {decision!r} is not one of {sorted(DECISIONS)}")
continue
if decision == "extend-once" and rec.get("extended_from"):
failures.append(f"{path}: already an extension of {rec['extended_from']}; decide graduate or stop")
if (expires - start).days > MAX_DAYS:
failures.append(f"{path}: runs {(expires - start).days} days (limit {MAX_DAYS})")
if expires < date.today() and decision == "pending":
failures.append(f"{path}: expired {expires} with no decision")
if decision == "stop" and not rec.get("cleanup_verified"):
failures.append(f"{path}: stopped but cleanup_verified is not set")
print("\n".join(failures) or f"{len(sys.argv) - 1} exception records OK")
sys.exit(1 if failures else 0)

Default, specialist exception, or retirement?

Section titled “Default, specialist exception, or retirement?”

When two tools overlap, decide on evidence from completed work, not on enthusiasm. Score both on the same representative tasks and apply the first row that matches.

EvidenceDecision
The tool fails a control the workflow row requires (login binding, MCP allowlist, data terms)Not approvable for that row, whatever its results
It completes the representative tasks no better than the defaultRetire, or never admit
It is clearly better for one workflow and no worse on controlsSpecialist default for that workflow only
It is better across most workflows, with equal controls and acceptable costCandidate to replace the default, as a roadmap bet with a migration plan

Portability lowers the cost of every row. Context in AGENTS.md, skills in the open Agent Skills format, and checks in CI move with you between tools. Codex reads AGENTS.md natively, and Claude Code reads it when a project has no CLAUDE.md (v2.1.277 and later, latest channel). The rest of the checklist is on avoiding lock-in.

Copy-paste prompts for tooling policy work

Section titled “Copy-paste prompts for tooling policy work”

Run these in Claude Code, Codex, or Cursor with the policy repository and the inventory output in the working directory. They work the same way in all three tools.

How do you prove the tooling policy is in force?

Section titled “How do you prove the tooling policy is in force?”

You do not need to inspect every laptop. You need checks that fail loudly when the policy drifts, and a named person who signs each one off.

ControlEvidencePass conditionSigns off
Matrix coverageQuarterly inventory gap reportEvery workflow in use maps to a matrix rowCTO
Service recordsOne dated record per tool and planEach record verified within the last 90 days and matching the admin consolePlatform lead
Default enforcementA test on a managed machine: add an unlisted MCP server, sign in with a personal accountBoth attempts are rejected, and the rejection is recordedSecurity lead
Exceptionscheck_exceptions.py in CIGreen on the default branch; no expired record without a decisionCTO
RetirementCleanup evidence for each stopped exception or retired toolSeats, apps, secrets, and local configuration removedIT or identity owner
Engineer clarityFive engineers asked which tool and data class a common task usesAll five answer from the matrixDevEx lead

Adoption and outcome evidence for the approved tools belongs to the AI metrics panel. The policy’s own job is narrower: showing that every tool in use is approved for its workflow and that nothing expired stays live.

Approving logos instead of workflows. “Codex is approved” lets an engineer use it on a personal Plus plan with none of your data terms. Recovery: rewrite each approval as a matrix row plus a service record that names the plan and workspace, then bind logins with managed settings, as described in team accounts.

Exceptions that never end. A pilot from spring is still running in autumn, with seats, a GitHub app, and secrets nobody owns. Recovery: backfill an exception record for each one with an expiry within 30 days, run check_exceptions.py in CI, and treat any stopped exception without cleanup_verified as open work.

A single mandated tool. The policy forces one tool, a team with a real need goes around it, and the shadow use returns. Recovery: keep one default per workflow, and route the need through an exception so it can prove itself on evidence.

Service records that go stale. A vendor changes a plan, a default model, or a retention rule, and the record still shows the old terms. For example, Claude Code’s two release channels started sessions on different default models on 2026-09-26. Recovery: date every record, re-verify on a 90-day cadence and on every vendor announcement, and pin the release channel in the record.

Unpinned extensions and CLIs. Approved tools ship through registries that have been compromised. A malicious script was injected into the Amazon Q Developer extension for VS Code, version 1.84.0 (AWS advisory, 2025-07-26), and an unauthorized npm publish of cline@2.3.0 carried a modified postinstall script (Cline advisory, 2026-02-17). Recovery: set a version floor in managed settings where the tool supports one, install from an internal mirror or a reviewed update channel, and add registry advisories for approved tools to the security team’s watch list.

No exit plan. A tool disappears and the team loses its context. The Roo Code extension shut down in May 2026, per its own README. Recovery: keep context, skills, and checks in the repository in portable formats, write an exit_plan in every service record, and rehearse it for your riskiest vendor with vendor risk management.

The policy decides what is approved. Account governance makes the approval real, and the roadmap decides what gets tried next.