AI usage cost governance
AI usage cost governance is the set of owners, budgets, spend limits, and model policies that keep Claude Code, Codex, and Cursor spend visible, attributed to teams, and tied to accepted work. It runs on each tool’s own telemetry and admin controls, and it judges spend by cost per accepted change, never by tokens per engineer.
Finance messages you on the first of the month: the AI tooling line doubled again, and they want to know why. You can name the three tools, but not which teams drove the increase, whether a flagship model ran on routine work at maximum effort, or whether any of that spend produced merged code. This page is for the CTO or engineering director who owns that line, and for the tech leads who will run the controls week to week.
What this cost-governance setup gives you
Section titled “What this cost-governance setup gives you”- A spend source and a spend cap for every way your teams reach a model: plan seats, API workspaces, cloud providers, and CI.
- A managed-settings file that turns on Claude Code telemetry for every developer, tags spend by team, caps effort, and restricts models.
- A model and effort policy that follows the vendors’ own defaults and names who may approve an escalation.
- A cost policy template, a monthly review checklist, and four copy-paste prompts for spend reviews, outlier analysis, the full cost of one workflow, and alert design.
- A way to prove the controls work without reading every session: reconciliation to the invoice, alert fire drills, and cost paired with quality.
Where does coding-agent spend come from?
Section titled “Where does coding-agent spend come from?”Agent spend is metered usage, not a seat fee, so the bill follows behaviour. Three drivers account for most of it:
- Model and effort. The model sets the price per token and the effort level sets how many thinking tokens each request burns. Anthropic’s costs page notes that thinking tokens bill as output and “can be tens of thousands of tokens per request”. Prices per model live on the models hub, not here.
- Context that is never cleared. Claude Code sends the whole conversation with every request. A session left open all day, or one resumed after the cache expired, re-reads its full history on each turn.
- Fan-out. Every subagent, workflow agent, and teammate sends its own requests. Anthropic’s costs page (checked 2026-09-26) says agent teams use “approximately 7x more tokens than standard sessions when teammates run in plan mode”.
For a baseline, Anthropic reports that across enterprise deployments Claude Code averages “around $13 per developer per active day and $150-250 per developer per month”, with 90% of users below $30 per active day (Anthropic, Claude Code costs docs, checked 2026-09-26). That is a vendor figure for one tool. Measure your own teams for a month before you set a cap.
Tokens are only part of the cost. Governance decisions need three layers, each with its own owner and decision:
| Layer | What it includes | Decision it drives |
|---|---|---|
| Portfolio | Contracts, seats, platform and integration work | Renew, consolidate, or retire a tool |
| Workflow | Model usage, CI and runtime, review, rework, failed attempts, incidents | Expand, narrow, or redesign a workflow |
| Change | Run cost and verification cost of one change, where the data allows | Diagnose outliers; never rank people |
The formula that joins those layers, cost per accepted change, lives on the economics of agent-built software. This page covers the controls that feed it.
Step 1: find the spend source and the cap for each access route
Section titled “Step 1: find the spend source and the cap for each access route”Cost controls depend on how a team reaches the model, not only on which tool it uses. A team on Claude Team seats, a CI job on an API key, and a platform team on Amazon Bedrock each have a different meter. List every route first; the workflows genuinely differ by tool.
Checked on 2026-09-26 against Anthropic’s costs documentation:
| Access route | Where you see spend | Where you cap it |
|---|---|---|
| Claude Team or Enterprise seats | Spend report in org analytics: estimated spend per user and per model, CSV, updated daily. Covers usage-credit spend only | The seat allowance is the default ceiling. With usage credits on, set spend limits at organization, group, or member level in Admin settings > Usage |
| Claude Console (API) | Console usage page and Claude Code dashboard, per member | Workspace spend limits. Claude Code creates a workspace named “Claude Code” on first Console login |
| Amazon Bedrock, Google Cloud’s Agent Platform (formerly Vertex AI), Microsoft Foundry | Your cloud billing console | Your cloud’s budget controls, or per-user spend limits on a self-hosted Claude apps gateway |
CI and scripts (claude -p) | OpenTelemetry, or the JSON output of each run | --max-budget-usd per run (works only with --print) |
OpenTelemetry export works on every route and streams per-user token and cost metrics into your own stack in near real time (on cloud providers, a Claude apps gateway can emit them too). Step 2 configures it. To choose and budget a cloud route or gateway, see model hosting.
Codex draws on ChatGPT plan allowances (five-hour and weekly windows, plus credits) or on an API key, and the two bill differently:
- In a session:
/statusshows estimated thread credits or cost for eligible workspaces (Codex CLI 0.148.0 or later), and/usageshows account usage (0.156.0 or later). - Plan usage: workspace admins read consumption in the ChatGPT workspace admin view. OpenAI’s help pages describe an Enterprise Codex analytics dashboard and an Analytics API (secondary: search extracts of OpenAI’s docs, which were unreachable on 2026-09-26).
- API-key usage: CLI or CI runs authenticated with an API key bill to that API project, not to the ChatGPT plan. Check which budget controls your OpenAI organization offers before relying on them.
- CI:
codex exec --jsonprints every event as JSONL. Store the output with the PR number so each automated run has a cost trail.
Decide per team whether Codex runs on plan allowances or API billing, and record the choice. Mixing both inside one team splits its spend across two bills.
Cursor’s plans, billing modes, and admin controls could not be re-verified from cursor.com on 2026-09-26, so no Cursor price or limit is stated here. Search extracts of Cursor’s docs (secondary) describe a team analytics dashboard and an Analytics API on Teams and Enterprise.
What to do regardless of plan: find the spend view and any spend limit in your admin dashboard, record each team’s plan and billing mode, and confirm how Bugbot reviews are billed. Anything billed per use belongs in the workflow layer, even inside a plan.
Tool plan prices are compared on pricing analysis. Contract terms, seat ownership, and offboarding are covered on team accounts.
Step 2: attribute spend to teams, not people
Section titled “Step 2: attribute spend to teams, not people”Aggregate spend is a number to worry about; attributed spend is a number you can act on. For Claude Code, deliver telemetry and attribution through managed settings so every developer reports the same way. Claude Code ignores the OpenTelemetry exporter variables in a repository’s .claude/settings.json, so a repository cannot turn telemetry on or redirect it.
{ "env": { "CLAUDE_CODE_ENABLE_TELEMETRY": "1", "OTEL_METRICS_EXPORTER": "otlp", "OTEL_LOGS_EXPORTER": "otlp", "OTEL_EXPORTER_OTLP_PROTOCOL": "grpc", "OTEL_EXPORTER_OTLP_ENDPOINT": "http://collector.internal.example:4317", "OTEL_RESOURCE_ATTRIBUTES": "department=engineering,team.id=payments,cost_center=eng-142", "OTEL_METRICS_INCLUDE_ACCOUNT_UUID": "false", "OTEL_METRICS_INCLUDE_SESSION_ID": "false" }}Ship one file per team, differing only in team.id and cost_center. The metric that matters for spend is claude_code.cost.usage (USD, labelled by model); claude_code.token.usage splits tokens by type (input, output, cache read, cache creation). Sum cost by team.id and model and you have per-team spend without building anything. Collector setup and dashboards are on agent observability.
Three details decide whether the numbers are usable:
- No spaces in attribute values.
OTEL_RESOURCE_ATTRIBUTESrejects whitespace, double quotes, commas, semicolons, and backslashes. Writecost_center=eng_platform, nevercost_center=Eng Platform. - Report at contracted rates. Claude Code computes cost at list price. Set the
modelPricingmanaged setting (Claude Code v2.1.242 or later) to your contract’s rates, orclaude_code.cost.usagewill not match the invoice. - Proportionate identity. The two
INCLUDE_*lines dropuser.account_uuid,user.account_id, andsession.idfrom metrics.user.email(when signed in with a Claude account) and the anonymoususer.idare always attached; drop or hashuser.emailin your collector (for example, an OpenTelemetry attributes processor) if policy is team-level only. Keep team-level attribution as the default and collect developer-level data only for a reviewed, legitimate purpose, such as a developer who opts in to personal optimization. Run identity-level telemetry past your privacy or works-council review first.
For Codex, the [otel] table in ~/.codex/config.toml exports logs, traces, and metrics to the same collector. The field names below are read from Codex source (OtelConfigToml, Codex CLI 0.157.1, checked 2026-09-26):
# ~/.codex/config.toml — Codex CLI 0.157.1[otel]environment = "prod"log_user_prompt = falseexporter = { otlp-http = { endpoint = "http://collector.internal.example:4318/v1/logs", protocol = "binary" } }metrics_exporter and trace_exporter take the same otlp-http or otlp-grpc values. Team attribution for Codex usually comes from separate ChatGPT workspaces or API projects per team. For Cursor, use the per-team view in the admin dashboard. Local trackers such as ccusage answer “what did I spend on this laptop?” and are covered on agent cost tracking.
Step 3: set a model and effort policy
Section titled “Step 3: set a model and effort policy”Routing is the lever you control directly, but the direction matters: the vendors now default to capable models at moderate effort. The rule this site follows is start on the tool’s default model, tune effort before switching model, and switch model only when your own evals say so. As of 2026-09-26:
| Work | Start with | When cost dominates | Escalate deliberately to |
|---|---|---|---|
| Hard agentic coding | Claude Opus 5.5 (Claude Code default, medium effort); GPT-6 Astra (Codex default) | Raise effort only on failure, not by habit | Claude Fable 5.1, chosen by hand with /model fable; never a default |
| Everyday feature work and review | Opus 5.5 at default effort; GPT-6 Sol in Codex | Claude Sonnet 5 | Opus 5.5 at high effort |
| High-volume, simple edits, subagent fan-out | Claude Haiku 4.5; GPT-6 Luna | Same | The session default |
On Claude Code’s stable release channel (2.1.274 on 2026-09-26), Opus 5.5 is not yet available and older defaults still apply. Prices and context windows for every model are on the models hub.
A policy nobody can find is not a policy. Encode it where each tool enforces or reads it:
Add the model and effort keys to the same managed settings file as Step 2:
{ "availableModels": ["opus", "sonnet", "haiku"], "enforceAvailableModels": true, "maxEffortLevel": "high"}availableModelsrestricts every place a user can pick a model, including subagent frontmatter andCLAUDE_CODE_SUBAGENT_MODEL. Leavingfableout blocks Claude Fable 5.1 for everyone who receives this file; give a team that needs it a different file.enforceAvailableModels(v2.1.175 or later) extends the allowlist to the Default option. Without it, a user who picks Default gets the account’s runtime default even when it is not on the list.maxEffortLevel(v2.1.267 or later) caps effort from every source, including/effort,--effort, and skill or subagent frontmatter. The lowest cap across settings scopes applies. A team approved forxhighormaxgets its own managed file with a higher cap, the same way as for Fable.
Subagents inherit the session’s model unless their frontmatter sets one, so a /model switch to a pricier model also reprices every inheriting subagent. Pin model: haiku on subagents that search, read logs, or run tests.
Set the team default and effort in config.toml, and keep the expensive setup one flag away as a profile:
# ~/.codex/config.toml — the team's everyday default, once evals support itmodel = "gpt-6-sol"model_reasoning_effort = "medium"# ~/.codex/deep.config.toml — run with: codex --profile deepmodel = "gpt-6-astra"model_reasoning_effort = "high"service_tier = "default" # standard routing, not Fast--profile deep layers $CODEX_HOME/deep.config.toml on top of the base config (Codex CLI 0.157.1). Admin constraints go in requirements.toml. Its 0.157.1 source has no model allowlist key (it can pin a custom model_catalog_json, which we have not tested), so treat the model choice as a default, not an enforced rule. GPT-6 Sol and Luna run on the Fast tier by default (“1.5x speed”, no extra-usage note in the 0.157.1 catalog). GPT-6 Astra’s Fast tier is marked “2x speed, increased usage”, so do not enable it in the deep profile; service_tier = "default" requests standard routing explicitly (the key accepts default, priority, or flex in 0.157.1 source).
Cursor’s model defaults and pools could not be verified on 2026-09-26, so enforce the policy through the admin dashboard’s model settings if your plan offers them, and make the default visible with a project rule (see Cursor’s Rules documentation):
Cost policy for agents in this repository:- Stay on the team default model for feature work and review.- Use the cheapest listed model for renames, formatting and single-file edits.- Before switching to a more expensive model or a long-running cloud agent, state the reason in one sentence in the chat and link the task.Step 4: set budgets and alerts that fail safely
Section titled “Step 4: set budgets and alerts that fail safely”Hard blocks that fire in the middle of an incident teach teams to route around you. Favour visibility and alerts, and put hard caps only where a runaway would hurt.
-
Size budgets per team, not per head. Base each team’s monthly budget on a month of measured spend and its current workload. Keep three budget lines apart: experimentation, production operation, and incident response. An incident must never hit a cap.
-
Alert before any cap. Notify the team lead at 75% of budget and the engineering manager at 90%. For Claude Code, alert on
claude_code.cost.usagesummed byteam.idover a rolling window in your metrics backend. -
Set hard caps as a backstop. Put the organization or workspace cap at 150-200% of a normal month so that it fires only on an anomaly, such as a looping automation. Use per-member limits in Admin settings > Usage for contractors and trial accounts.
-
Cap every unattended run. Give each CI job its own ceiling, for example
claude -p --max-budget-usd 5 "Review this diff for injection risks". A capped run stops; make sure the job reports that it stopped and keeps its partial output as an artifact. -
Alert on cost without output. Flag sessions or pipelines with repeated failed attempts, and spend with no merged change in the same week. Those are the signals of a broken workflow, not of a heavy user.
Write the policy down so the controls have an owner. Adopt this template as-is and fill in the brackets:
AI usage cost policy — [ORGANIZATION], version [N], owner: [NAME, ROLE]
1. Unit of decision: cost per accepted change (merged, not reverted within 30 days).2. Budgets: per team, per month, in three lines — experimentation, production operation, incident response. Incident response has no hard cap.3. Alerts: team lead at 75%, engineering manager at 90%. Org cap at 175% of the trailing three-month average, reviewed quarterly.4. Models: tool defaults apply. Escalation to Claude Fable 5.1 or maximum effort needs a named task and is reviewed monthly. Subagents doing search, logs or tests run on the cheapest allowed model.5. Attribution: team and cost center only; user.email is dropped or hashed in the collector. Developer-level data requires opt-in or a documented, reviewed purpose. Spend is never used to rate individuals.6. Review: monthly spend review with finance; quarterly budget reallocation.7. Evidence kept: allocation logic, currency, taxes, time window, and every estimated or missing data source, marked as such.Copy-paste prompts for cost reviews
Section titled “Copy-paste prompts for cost reviews”Run these in Claude Code, Codex, or Cursor with the exported spend report or telemetry query result attached. They work the same way in all three tools.
Developers can apply the same discipline per session. Clearing context between unrelated tasks, delegating verbose work to subagents, and scoping the files in play are covered on the cost of context.
How do you prove the cost controls work?
Section titled “How do you prove the cost controls work?”You do not verify spend by reading sessions. You verify it with checks that fail loudly when a control breaks, and each has an owner:
- Reconcile to the invoice monthly. Compare summed
claude_code.cost.usageand the vendor spend reports with the invoice. WithmodelPricingin force,/usagemarks the session total “at your organization’s configured rates”; any gap beyond your agreed tolerance is a data defect to fix before the review. Owner: the FinOps partner. - Run an alert fire drill each quarter. Lower one team’s threshold below its current spend and confirm the message reaches the channel. An alert that has never fired is untested. Owner: the platform team.
- Check telemetry coverage. Count distinct
user.idvalues (anonymous) reportingclaude_code.session.countand compare them with your seat count, so the check works without email-level identity.user.idis per installation, so treat the comparison as approximate: fewer IDs than active seats is the signal to chase, because it means someone runs without managed settings. - Pair every cost view with quality. Review cost next to accepted lead time, reverts, escaped defects, and review load from the AI metrics panel. A saving that raises reverts is not a saving.
- Trace any aggregate to source. The dashboard must drill from a team total to the source export and the decision it drove. Missing and estimated data stay visible, never silently zero.
Run the same monthly review in this order:
- Pull the month’s export: OpenTelemetry totals per team plus each vendor’s spend report.
- Reconcile the totals with the invoice and log any gap beyond tolerance as a data defect.
- Run the monthly spend and routing review prompt on the reconciled data.
- Run the outlier prompt on the highest-cost sessions and workflows.
- Check each cost change against the quality metrics for the same teams.
- Record the decisions (default changes, budget moves, fixes) with an owner and a date.
The engineering manager signs off each team’s monthly numbers; the CTO signs off the quarterly reallocation with finance.
What breaks when you govern AI spend?
Section titled “What breaks when you govern AI spend?”- Dashboards stay empty while spend continues. Telemetry was set in a repository’s
.claude/settings.json, which Claude Code ignores for exporter variables, orCLAUDE_CODE_ENABLE_TELEMETRYis missing. Recovery: deliver the variables through managed settings, then test on one machine withOTEL_METRICS_EXPORTER=consoleandOTEL_METRIC_EXPORT_INTERVAL=10000; aclaude_code.session.countdatapoint prints within about ten seconds. - Metrics never reach the collector. The endpoint and protocol disagree: gRPC usually listens on port 4317 and HTTP on 4318. Recovery: match
OTEL_EXPORTER_OTLP_PROTOCOLto the port, and grep the output ofclaude --debug-file <path>for[3P telemetry]. - Team attribution comes back blank. A space or quote in
OTEL_RESOURCE_ATTRIBUTESinvalidates the value. Recovery: use underscores or camelCase and re-deploy. - Telemetry cost does not match the invoice. Claude Code reports list price by default, and seat-allowance usage is not metered in dollars. Recovery: set
modelPricing, and record allowance consumption as its own line. - The model allowlist leaks through Default.
availableModelsalone leaves the Default option on the account’s runtime default. Recovery: addenforceAvailableModels: truein the same managed source. - Subagent fan-out quietly costs flagship rates. Subagents inherit the session model. Recovery: pin
model: haikuon read-heavy subagents and on subscription plans, review the subagent share in the/usagebreakdown. - Codex CI runs show no metrics. Recovery: if
codex execruns show no metrics in your collector, keep thecodex exec --jsonoutput per run as the cost record until they do. - A hard cap stops incident response. Recovery: move incident work to its own uncapped budget line and review it after the incident, not during it.
- Spend becomes a performance metric. High spend can mean valuable hard work or repeated failure; low spend can mean efficiency or no adoption. Recovery: remove per-person rankings from every dashboard and diagnose workflows instead.
- The policy names last quarter’s models. In September 2026 alone, Anthropic (Fable 5.1, Opus 5.5), OpenAI (GPT-6), Google (Gemini 3.8 Flash), and xAI (Grok 4.7) shipped new models (Gemini and Grok dates: secondary, The Register and MarkTechPost). Recovery: date the model table, check it against the models hub monthly, and re-run your evals before changing a default.