Skip to content

AI usage cost governance

AI usage cost governance is the set of owners, budgets, spend limits, and model policies that keep Claude Code, Codex, and Cursor spend visible, attributed to teams, and tied to accepted work. It runs on each tool’s own telemetry and admin controls, and it judges spend by cost per accepted change, never by tokens per engineer.

Finance messages you on the first of the month: the AI tooling line doubled again, and they want to know why. You can name the three tools, but not which teams drove the increase, whether a flagship model ran on routine work at maximum effort, or whether any of that spend produced merged code. This page is for the CTO or engineering director who owns that line, and for the tech leads who will run the controls week to week.

  • A spend source and a spend cap for every way your teams reach a model: plan seats, API workspaces, cloud providers, and CI.
  • A managed-settings file that turns on Claude Code telemetry for every developer, tags spend by team, caps effort, and restricts models.
  • A model and effort policy that follows the vendors’ own defaults and names who may approve an escalation.
  • A cost policy template, a monthly review checklist, and four copy-paste prompts for spend reviews, outlier analysis, the full cost of one workflow, and alert design.
  • A way to prove the controls work without reading every session: reconciliation to the invoice, alert fire drills, and cost paired with quality.

Agent spend is metered usage, not a seat fee, so the bill follows behaviour. Three drivers account for most of it:

  • Model and effort. The model sets the price per token and the effort level sets how many thinking tokens each request burns. Anthropic’s costs page notes that thinking tokens bill as output and “can be tens of thousands of tokens per request”. Prices per model live on the models hub, not here.
  • Context that is never cleared. Claude Code sends the whole conversation with every request. A session left open all day, or one resumed after the cache expired, re-reads its full history on each turn.
  • Fan-out. Every subagent, workflow agent, and teammate sends its own requests. Anthropic’s costs page (checked 2026-09-26) says agent teams use “approximately 7x more tokens than standard sessions when teammates run in plan mode”.

For a baseline, Anthropic reports that across enterprise deployments Claude Code averages “around $13 per developer per active day and $150-250 per developer per month”, with 90% of users below $30 per active day (Anthropic, Claude Code costs docs, checked 2026-09-26). That is a vendor figure for one tool. Measure your own teams for a month before you set a cap.

Tokens are only part of the cost. Governance decisions need three layers, each with its own owner and decision:

LayerWhat it includesDecision it drives
PortfolioContracts, seats, platform and integration workRenew, consolidate, or retire a tool
WorkflowModel usage, CI and runtime, review, rework, failed attempts, incidentsExpand, narrow, or redesign a workflow
ChangeRun cost and verification cost of one change, where the data allowsDiagnose outliers; never rank people

The formula that joins those layers, cost per accepted change, lives on the economics of agent-built software. This page covers the controls that feed it.

Step 1: find the spend source and the cap for each access route

Section titled “Step 1: find the spend source and the cap for each access route”

Cost controls depend on how a team reaches the model, not only on which tool it uses. A team on Claude Team seats, a CI job on an API key, and a platform team on Amazon Bedrock each have a different meter. List every route first; the workflows genuinely differ by tool.

Checked on 2026-09-26 against Anthropic’s costs documentation:

Access routeWhere you see spendWhere you cap it
Claude Team or Enterprise seatsSpend report in org analytics: estimated spend per user and per model, CSV, updated daily. Covers usage-credit spend onlyThe seat allowance is the default ceiling. With usage credits on, set spend limits at organization, group, or member level in Admin settings > Usage
Claude Console (API)Console usage page and Claude Code dashboard, per memberWorkspace spend limits. Claude Code creates a workspace named “Claude Code” on first Console login
Amazon Bedrock, Google Cloud’s Agent Platform (formerly Vertex AI), Microsoft FoundryYour cloud billing consoleYour cloud’s budget controls, or per-user spend limits on a self-hosted Claude apps gateway
CI and scripts (claude -p)OpenTelemetry, or the JSON output of each run--max-budget-usd per run (works only with --print)

OpenTelemetry export works on every route and streams per-user token and cost metrics into your own stack in near real time (on cloud providers, a Claude apps gateway can emit them too). Step 2 configures it. To choose and budget a cloud route or gateway, see model hosting.

Tool plan prices are compared on pricing analysis. Contract terms, seat ownership, and offboarding are covered on team accounts.

Step 2: attribute spend to teams, not people

Section titled “Step 2: attribute spend to teams, not people”

Aggregate spend is a number to worry about; attributed spend is a number you can act on. For Claude Code, deliver telemetry and attribution through managed settings so every developer reports the same way. Claude Code ignores the OpenTelemetry exporter variables in a repository’s .claude/settings.json, so a repository cannot turn telemetry on or redirect it.

{
"env": {
"CLAUDE_CODE_ENABLE_TELEMETRY": "1",
"OTEL_METRICS_EXPORTER": "otlp",
"OTEL_LOGS_EXPORTER": "otlp",
"OTEL_EXPORTER_OTLP_PROTOCOL": "grpc",
"OTEL_EXPORTER_OTLP_ENDPOINT": "http://collector.internal.example:4317",
"OTEL_RESOURCE_ATTRIBUTES": "department=engineering,team.id=payments,cost_center=eng-142",
"OTEL_METRICS_INCLUDE_ACCOUNT_UUID": "false",
"OTEL_METRICS_INCLUDE_SESSION_ID": "false"
}
}

Ship one file per team, differing only in team.id and cost_center. The metric that matters for spend is claude_code.cost.usage (USD, labelled by model); claude_code.token.usage splits tokens by type (input, output, cache read, cache creation). Sum cost by team.id and model and you have per-team spend without building anything. Collector setup and dashboards are on agent observability.

Three details decide whether the numbers are usable:

  • No spaces in attribute values. OTEL_RESOURCE_ATTRIBUTES rejects whitespace, double quotes, commas, semicolons, and backslashes. Write cost_center=eng_platform, never cost_center=Eng Platform.
  • Report at contracted rates. Claude Code computes cost at list price. Set the modelPricing managed setting (Claude Code v2.1.242 or later) to your contract’s rates, or claude_code.cost.usage will not match the invoice.
  • Proportionate identity. The two INCLUDE_* lines drop user.account_uuid, user.account_id, and session.id from metrics. user.email (when signed in with a Claude account) and the anonymous user.id are always attached; drop or hash user.email in your collector (for example, an OpenTelemetry attributes processor) if policy is team-level only. Keep team-level attribution as the default and collect developer-level data only for a reviewed, legitimate purpose, such as a developer who opts in to personal optimization. Run identity-level telemetry past your privacy or works-council review first.

For Codex, the [otel] table in ~/.codex/config.toml exports logs, traces, and metrics to the same collector. The field names below are read from Codex source (OtelConfigToml, Codex CLI 0.157.1, checked 2026-09-26):

# ~/.codex/config.toml — Codex CLI 0.157.1
[otel]
environment = "prod"
log_user_prompt = false
exporter = { otlp-http = { endpoint = "http://collector.internal.example:4318/v1/logs", protocol = "binary" } }

metrics_exporter and trace_exporter take the same otlp-http or otlp-grpc values. Team attribution for Codex usually comes from separate ChatGPT workspaces or API projects per team. For Cursor, use the per-team view in the admin dashboard. Local trackers such as ccusage answer “what did I spend on this laptop?” and are covered on agent cost tracking.

Routing is the lever you control directly, but the direction matters: the vendors now default to capable models at moderate effort. The rule this site follows is start on the tool’s default model, tune effort before switching model, and switch model only when your own evals say so. As of 2026-09-26:

WorkStart withWhen cost dominatesEscalate deliberately to
Hard agentic codingClaude Opus 5.5 (Claude Code default, medium effort); GPT-6 Astra (Codex default)Raise effort only on failure, not by habitClaude Fable 5.1, chosen by hand with /model fable; never a default
Everyday feature work and reviewOpus 5.5 at default effort; GPT-6 Sol in CodexClaude Sonnet 5Opus 5.5 at high effort
High-volume, simple edits, subagent fan-outClaude Haiku 4.5; GPT-6 LunaSameThe session default

On Claude Code’s stable release channel (2.1.274 on 2026-09-26), Opus 5.5 is not yet available and older defaults still apply. Prices and context windows for every model are on the models hub.

A policy nobody can find is not a policy. Encode it where each tool enforces or reads it:

Add the model and effort keys to the same managed settings file as Step 2:

{
"availableModels": ["opus", "sonnet", "haiku"],
"enforceAvailableModels": true,
"maxEffortLevel": "high"
}
  • availableModels restricts every place a user can pick a model, including subagent frontmatter and CLAUDE_CODE_SUBAGENT_MODEL. Leaving fable out blocks Claude Fable 5.1 for everyone who receives this file; give a team that needs it a different file.
  • enforceAvailableModels (v2.1.175 or later) extends the allowlist to the Default option. Without it, a user who picks Default gets the account’s runtime default even when it is not on the list.
  • maxEffortLevel (v2.1.267 or later) caps effort from every source, including /effort, --effort, and skill or subagent frontmatter. The lowest cap across settings scopes applies. A team approved for xhigh or max gets its own managed file with a higher cap, the same way as for Fable.

Subagents inherit the session’s model unless their frontmatter sets one, so a /model switch to a pricier model also reprices every inheriting subagent. Pin model: haiku on subagents that search, read logs, or run tests.

Step 4: set budgets and alerts that fail safely

Section titled “Step 4: set budgets and alerts that fail safely”

Hard blocks that fire in the middle of an incident teach teams to route around you. Favour visibility and alerts, and put hard caps only where a runaway would hurt.

  1. Size budgets per team, not per head. Base each team’s monthly budget on a month of measured spend and its current workload. Keep three budget lines apart: experimentation, production operation, and incident response. An incident must never hit a cap.

  2. Alert before any cap. Notify the team lead at 75% of budget and the engineering manager at 90%. For Claude Code, alert on claude_code.cost.usage summed by team.id over a rolling window in your metrics backend.

  3. Set hard caps as a backstop. Put the organization or workspace cap at 150-200% of a normal month so that it fires only on an anomaly, such as a looping automation. Use per-member limits in Admin settings > Usage for contractors and trial accounts.

  4. Cap every unattended run. Give each CI job its own ceiling, for example claude -p --max-budget-usd 5 "Review this diff for injection risks". A capped run stops; make sure the job reports that it stopped and keeps its partial output as an artifact.

  5. Alert on cost without output. Flag sessions or pipelines with repeated failed attempts, and spend with no merged change in the same week. Those are the signals of a broken workflow, not of a heavy user.

Write the policy down so the controls have an owner. Adopt this template as-is and fill in the brackets:

AI usage cost policy — [ORGANIZATION], version [N], owner: [NAME, ROLE]
1. Unit of decision: cost per accepted change (merged, not reverted within 30 days).
2. Budgets: per team, per month, in three lines — experimentation, production
operation, incident response. Incident response has no hard cap.
3. Alerts: team lead at 75%, engineering manager at 90%. Org cap at 175% of the
trailing three-month average, reviewed quarterly.
4. Models: tool defaults apply. Escalation to Claude Fable 5.1 or maximum effort
needs a named task and is reviewed monthly. Subagents doing search, logs or
tests run on the cheapest allowed model.
5. Attribution: team and cost center only; user.email is dropped or hashed in
the collector. Developer-level data requires opt-in or a documented,
reviewed purpose. Spend is never used to rate individuals.
6. Review: monthly spend review with finance; quarterly budget reallocation.
7. Evidence kept: allocation logic, currency, taxes, time window, and every
estimated or missing data source, marked as such.

Run these in Claude Code, Codex, or Cursor with the exported spend report or telemetry query result attached. They work the same way in all three tools.

Developers can apply the same discipline per session. Clearing context between unrelated tasks, delegating verbose work to subagents, and scoping the files in play are covered on the cost of context.

You do not verify spend by reading sessions. You verify it with checks that fail loudly when a control breaks, and each has an owner:

  • Reconcile to the invoice monthly. Compare summed claude_code.cost.usage and the vendor spend reports with the invoice. With modelPricing in force, /usage marks the session total “at your organization’s configured rates”; any gap beyond your agreed tolerance is a data defect to fix before the review. Owner: the FinOps partner.
  • Run an alert fire drill each quarter. Lower one team’s threshold below its current spend and confirm the message reaches the channel. An alert that has never fired is untested. Owner: the platform team.
  • Check telemetry coverage. Count distinct user.id values (anonymous) reporting claude_code.session.count and compare them with your seat count, so the check works without email-level identity. user.id is per installation, so treat the comparison as approximate: fewer IDs than active seats is the signal to chase, because it means someone runs without managed settings.
  • Pair every cost view with quality. Review cost next to accepted lead time, reverts, escaped defects, and review load from the AI metrics panel. A saving that raises reverts is not a saving.
  • Trace any aggregate to source. The dashboard must drill from a team total to the source export and the decision it drove. Missing and estimated data stay visible, never silently zero.

Run the same monthly review in this order:

  1. Pull the month’s export: OpenTelemetry totals per team plus each vendor’s spend report.
  2. Reconcile the totals with the invoice and log any gap beyond tolerance as a data defect.
  3. Run the monthly spend and routing review prompt on the reconciled data.
  4. Run the outlier prompt on the highest-cost sessions and workflows.
  5. Check each cost change against the quality metrics for the same teams.
  6. Record the decisions (default changes, budget moves, fixes) with an owner and a date.

The engineering manager signs off each team’s monthly numbers; the CTO signs off the quarterly reallocation with finance.

  • Dashboards stay empty while spend continues. Telemetry was set in a repository’s .claude/settings.json, which Claude Code ignores for exporter variables, or CLAUDE_CODE_ENABLE_TELEMETRY is missing. Recovery: deliver the variables through managed settings, then test on one machine with OTEL_METRICS_EXPORTER=console and OTEL_METRIC_EXPORT_INTERVAL=10000; a claude_code.session.count datapoint prints within about ten seconds.
  • Metrics never reach the collector. The endpoint and protocol disagree: gRPC usually listens on port 4317 and HTTP on 4318. Recovery: match OTEL_EXPORTER_OTLP_PROTOCOL to the port, and grep the output of claude --debug-file <path> for [3P telemetry].
  • Team attribution comes back blank. A space or quote in OTEL_RESOURCE_ATTRIBUTES invalidates the value. Recovery: use underscores or camelCase and re-deploy.
  • Telemetry cost does not match the invoice. Claude Code reports list price by default, and seat-allowance usage is not metered in dollars. Recovery: set modelPricing, and record allowance consumption as its own line.
  • The model allowlist leaks through Default. availableModels alone leaves the Default option on the account’s runtime default. Recovery: add enforceAvailableModels: true in the same managed source.
  • Subagent fan-out quietly costs flagship rates. Subagents inherit the session model. Recovery: pin model: haiku on read-heavy subagents and on subscription plans, review the subagent share in the /usage breakdown.
  • Codex CI runs show no metrics. Recovery: if codex exec runs show no metrics in your collector, keep the codex exec --json output per run as the cost record until they do.
  • A hard cap stops incident response. Recovery: move incident work to its own uncapped budget line and review it after the incident, not during it.
  • Spend becomes a performance metric. High spend can mean valuable hard work or repeated failure; low spend can mean efficiency or no adoption. Recovery: remove per-person rankings from every dashboard and diagnose workflows instead.
  • The policy names last quarter’s models. In September 2026 alone, Anthropic (Fable 5.1, Opus 5.5), OpenAI (GPT-6), Google (Gemini 3.8 Flash), and xAI (Grok 4.7) shipped new models (Gemini and Grok dates: secondary, The Register and MarkTechPost). Recovery: date the model table, check it against the models hub monthly, and re-run your evals before changing a default.