Skip to content

What Did the Agents Cost? ccusage, /usage, Monitors and Gateway Budgets

Agent cost tracking works at three levels. Built-in commands (Claude Code /usage, Codex /status and /usage) show one session or plan. Local trackers such as ccusage total one machine’s cost. Capping a team’s spend takes budgets in the request path: plan or Console spend limits, or a gateway with per-key budgets such as LiteLLM or Anthropic’s Claude apps gateway.

A developer ran three agents in parallel over a weekend, the API invoice jumped, and finance wants per-team numbers you do not have. This page is for the developer who needs to see their own spend today and for the tech lead who has to put a budget on a team without slowing it down.

What you get from this cost-tracking setup

Section titled “What you get from this cost-tracking setup”
  • A map of which tool answers which cost question, checked on 2026-09-26 against Claude Code 2.1.283, Codex CLI 0.157.1 and ccusage 20.0.24.
  • Your own monthly cost in one command, and what that figure means on a subscription versus an API key.
  • A live status line and a 5-hour window monitor, so a long agent run does not surprise you.
  • A per-team budget with LiteLLM virtual keys for Claude Code and Codex, plus two alternatives.
  • A ten-minute monthly cost report, three copy-paste prompts, and the billing trap that makes a gateway count nothing.

This page covers the tools. Organization-wide policy (budget ownership, cost-centre attribution, overruns) lives on AI usage cost governance, and the method for turning spend into cost per accepted change lives on the economics page.

Start from the question. Every row was checked on 2026-09-26 against the installed CLIs, the Claude Code costs and gateway spend limits docs, the Codex source at tag rust-v0.157.1, and each project’s GitHub repository.

QuestionToolScopeCan it cap spend?
What did this session cost, and how close am I to my plan limit?Claude Code /usage; Codex /status and /usageOne session, one planNo
What did I spend this month across every agent on my laptop?ccusageOne machine, 18 agentsNo
When will my 5-hour window run out?ccusage blocks --active, Claude-Code-Usage-MonitorOne machine, Claude CodeNo (warns only)
How many tokens per repository, even across worktrees?tokscaleOne machine, many agentsNo
What did each developer spend, in near real time?Claude Code OpenTelemetry, vendor spend reportsWhole teamNo
Can I stop one unattended CI run from overspending?claude -p --max-budget-usdOne headless runYes, per run
Can I give each team a monthly budget?LiteLLM virtual keys, Claude apps gateway, your Console or plan spend limits; Cloudflare AI Gateway tracks spend per gateway, rate limits onlyEvery request through the gateway or accountYes

Adoption as of 2026-09-26, from GitHub and the package registries: ccusage has 18.7k stars (npm 20.0.24), LiteLLM 59.6k (PyPI 1.102.1), Portkey’s gateway 13.1k (npm 1.15.2), Claude-Code-Usage-Monitor 8.7k (PyPI claude-monitor 4.0.0) and tokscale 5.5k (npm 4.17.0). None is a plugin or an MCP server, so none adds tokens to the agent’s context.

For a baseline before you set budgets: Anthropic’s costs page (checked 2026-09-26) reports that across enterprise deployments Claude Code averages “around $13 per developer per active day and $150-250 per developer per month”, with 90% of users below $30 per active day. That is a vendor figure for Claude Code only; measure your own team for a month before setting a cap.

Type /usage in a session. The Session block shows the session’s total cost, API and wall-clock duration, code changes and cost by model.

What the dollar figure means depends on how you sign in:

  • Pro, Max, Team or Enterprise seat: the Session cost is a local list-price figure; for Pro and Max, Anthropic’s costs page says it “isn’t relevant for billing purposes”. What counts is the plan usage bars on the same screen, with recent usage attributed to skills, subagents, plugins and MCP servers (d and w switch between 24 hours and 7 days).
  • API key or cloud provider: the Session cost estimates real per-token spend; the authoritative figure is on the Console usage page or your cloud billing console.

Since v2.1.211 the totals reset on /clear. Since v2.1.251 a Prompt cache (main) line shows the cache-hit share of input tokens and the cache misses, the usual source of unexplained spend in long sessions.

Report a month across every agent with ccusage

Section titled “Report a month across every agent with ccusage”

ccusage prices the usage logs that coding agents write on your machine; no server, no sign-up:

Terminal window
npx ccusage@latest monthly

The report opens with a Detected: line naming the agents found (for example Detected: Claude, Codex), then a row per month with a sub-row per agent. Columns: Month, Agent, Models, Input, Output, Cache Create, Cache Read, Total Tokens and Cost (USD). On most agent-heavy machines Cache Read is by far the largest token column, which is why a long session costs more than its prompts suggest.

The commands you will use most, all checked against ccusage --help 20.0.24:

Terminal window
npx ccusage@latest # default: daily report, all detected agents
npx ccusage@latest monthly --last 3 # the last three months only
npx ccusage@latest claude daily --instances # Claude Code, split by project
npx ccusage@latest claude monthly --breakdown # Claude Code, split by model
npx ccusage@latest blocks --active # current 5-hour block with a projection
npx ccusage@latest monthly --since 2026-09-01 --json > sept.json

The unified monthly has no --breakdown flag (the per-agent subcommands and blocks do); --by-agent adds per-agent breakdowns to the unified JSON.

For Claude Code, ccusage has a status-line mode (marked beta in 20.0.24). Add it to ~/.claude/settings.json:

~/.claude/settings.json
{
"statusLine": { "type": "command", "command": "npx -y ccusage@20.0.24 statusline" }
}

Pin the version: the status line runs on every refresh, and an unpinned npx -y ccusage would fetch whatever npm publishes next. It runs locally and costs no context; --cost-source cc shows Claude Code’s own session cost instead of ccusage’s calculation.

Two alternatives, both local:

Terminal window
uv tool install claude-monitor # PyPI package; npm "claude-monitor" is a different project
claude-monitor --plan pro # live TUI of token burn against your plan window
npx tokscale@latest models --light --group-by workspace,model --merge-worktrees --month

Keep claude-monitor in a split pane during a long run. The tokscale line gives one row per repository even across parallel git worktrees, which neither ccusage nor /usage does.

For a headless job in CI, Claude Code 2.1.283 has --max-budget-usd <amount>: “Maximum dollar amount to spend on API calls (only works with —print)”.

Terminal window
claude -p "Fix the failing test in packages/billing and run the suite" \
--max-budget-usd 5 \
--permission-mode acceptEdits \
--allowedTools "Bash(npm test *)"

claude -p starts in Manual permission mode, so without --permission-mode and --allowedTools the run can neither edit files nor run the suite; headless agents in CI covers the rest of the setup. The budget caps one run, not a team, using the same list-price estimate as /usage. No equivalent flag could be found in codex exec --help 0.157.1.

Local tools report; they cannot stop anyone. A cap needs every request to pass something that counts spend before forwarding it. Pick the first row that fits.

You runUseWhyWatch out for
Claude Team or Enterprise seatsThe plan’s own controls: seat allowance, usage credits and spend limits in Admin settings > UsageNo infrastructure; per-user spend report with CSV exportUsage inside the seat allowance is not metered in dollars
Anthropic API through the ConsoleWorkspace spend limits in the ConsoleBuilt in; Claude Code gets its own workspaceOne “Claude Code” workspace for Console logins; per-team limits need team API keys issued from separate workspaces
Claude Code on Bedrock, Claude Platform on AWS, Google Cloud or FoundryClaude apps gateway (claude gateway --config gateway.yaml)Per-user, per-group and org caps by day, week or month; SSO; ships in the claude binaryNeeds Postgres; enforcement fails open by default if the store is down. Also accepts the Anthropic API as an upstream, so Console shops can use it for per-user SSO caps
Several providers, Claude Code and Codex togetherLiteLLM proxyVirtual keys per developer, budgets per key and per team, one bill across providersSelf-hosted; pin the version (see step 1 below)
Cloudflare already in front of your trafficCloudflare AI GatewayManaged; logs, caching, rate limits and spend per gatewayReports spend per gateway; no verified dollar cap
A gateway for your own LLM features tooPortkey gateway (npx @portkey-ai/gateway, port 8787)Routing, fallbacks, guardrailsClaude Code setup unverified (2026-09-26)

Avoid Helicone for new work: in maintenance mode since March 2026 after Mintlify acquired it (secondary: press and blog sources).

Set a per-team budget with LiteLLM virtual keys

Section titled “Set a per-team budget with LiteLLM virtual keys”

The workflow below gives the payments team $400 a month, each developer on it a personal $150 cap, and produces a monthly report. It uses Anthropic models behind LiteLLM; the same keys work for Codex through an OpenAI-format provider. The air-gapped variant with a local model is on gateways and local models.

  1. Install a pinned LiteLLM. On 2026-03-24, litellm 1.82.7 and 1.82.8 on PyPI were malicious credential stealers; PyPI removed both, and the maintainers track the incident in BerriAI/litellm#24518. Never install an unpinned proxy on a host that holds provider keys:

    Terminal window
    uv tool install 'litellm[proxy]==1.102.1'

    For a server image, compile a hashed requirements file and install with pip --require-hashes. The npm package litellm is an unrelated JavaScript port.

  2. Describe the models and the database. Virtual keys and budgets need Postgres (DATABASE_URL) and a master key. The provider key and the master key come from your secret store:

    litellm-config.yaml
    model_list:
    - model_name: claude-opus-5-5
    litellm_params:
    model: anthropic/claude-opus-5-5
    api_key: os.environ/ANTHROPIC_API_KEY
    input_cost_per_token: 0.000004 # $4 per million input tokens
    output_cost_per_token: 0.00002 # $20 per million output tokens
    cache_read_input_token_cost: 0.0000002 # $0.20 per million cached input tokens
    - model_name: claude-haiku-4-5
    litellm_params:
    model: anthropic/claude-haiku-4-5
    api_key: os.environ/ANTHROPIC_API_KEY
    general_settings:
    master_key: os.environ/LITELLM_MASTER_KEY
    database_url: os.environ/DATABASE_URL
    Terminal window
    litellm --config litellm-config.yaml # listens on port 4000

    The Opus 5.5 entry carries its own list prices from the models hub because the price map bundled with LiteLLM 1.102.1 predates Opus 5.5. Without them, a proxy that cannot fetch the remote map (an egress-restricted host, or LITELLM_LOCAL_MODEL_COST_MAP=True) records $0 for every Opus 5.5 request.

    Terminate TLS in front of the proxy, because keys travel in headers; the examples use https://llm-gw.internal. The Haiku entry serves Claude Code’s background work (conversation summaries and similar), which step 5 pins to it. List every model your developers pick with /model; unlisted models fail.

  3. Create the team with a budget. max_budget is in US dollars and budget_duration resets it:

    Terminal window
    curl -sS https://llm-gw.internal/team/new \
    -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
    -H "Content-Type: application/json" \
    -d '{"team_alias": "payments", "max_budget": 400, "soft_budget": 300, "budget_duration": "1mo"}'

    The response contains a team_id; keep it. soft_budget blocks nothing; it only triggers LiteLLM’s Slack or email alerts, if configured. In LiteLLM 1.102.1, 1mo (and 30d) resets at midnight on the first of each month, in UTC unless you set timezone under litellm_settings to your finance calendar’s zone.

  4. Issue one key per developer, inside the team. A key’s own cap stops one developer’s weekend fleet from spending the whole team budget:

    Terminal window
    curl -sS https://llm-gw.internal/key/generate \
    -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
    -H "Content-Type: application/json" \
    -d '{"team_id": "TEAM_ID", "key_alias": "payments-anna", "max_budget": 150, "budget_duration": "1mo", "models": ["claude-opus-5-5", "claude-haiku-4-5"]}'

    TEAM_ID is the value from step 3. Hand the returned key to the developer through your secret manager, never in chat. When a key or its team passes max_budget, LiteLLM rejects further requests with HTTP 429 (budget_exceeded) until the period resets.

  5. Point the agents at the gateway.

    Put the base URL in managed settings so no developer can forget it, and let each developer supply their own key:

    Managed settings: env block
    {
    "env": {
    "ANTHROPIC_BASE_URL": "https://llm-gw.internal",
    "ANTHROPIC_DEFAULT_HAIKU_MODEL": "claude-haiku-4-5"
    }
    }

    Each developer exports ANTHROPIC_AUTH_TOKEN from their secret store, or you configure an apiKeyHelper that reads it from your vault. Without it, the trap above applies. Managed settings are covered on managed policy.

  6. Prove the budget counts before anyone relies on it. On one laptop, start a fresh session, run /status and confirm the base URL is the gateway and the credential source is the environment variable, not the claude.ai login. Send one prompt. Then read the key’s spend:

    Terminal window
    curl -sS "https://llm-gw.internal/key/info?key=$(printf %s "$DEVELOPER_KEY" | sha256sum | cut -d' ' -f1)" \
    -H "Authorization: Bearer $LITELLM_MASTER_KEY"

    The command sends the SHA-256 hash of the step-4 key (LiteLLM 1.102.1 accepts it), so the raw key stays out of URLs and logs. Spend must be above zero and close to the Session cost in /usage. If it is zero, LiteLLM has no price for that model: add input_cost_per_token and output_cost_per_token to the model entry, using the rates on the models hub. Finally, set a test key to a $0.01 budget and confirm the next request is rejected.

  7. Report monthly. On the first working day, the tech lead pulls each team’s budget and spend with GET /team/info?team_id=TEAM_ID and sets it beside the vendor bill; the prompt below automates this.

Route through Cloudflare AI Gateway or cap with the Claude apps gateway

Section titled “Route through Cloudflare AI Gateway or cap with the Claude apps gateway”

Cloudflare AI Gateway needs no server. Cloudflare’s Claude Code integration page (checked 2026-09-26) uses three variables; read the token from your secret store:

Terminal window
export ANTHROPIC_BASE_URL="https://gateway.ai.cloudflare.com/v1/$CF_ACCOUNT_ID/$CF_GATEWAY_ID/anthropic"
export ANTHROPIC_API_KEY="$CF_AIG_TOKEN"
export ANTHROPIC_CUSTOM_HEADERS="cf-aig-authorization: Bearer $CF_AIG_TOKEN"
claude

The gateway must be authenticated and the token needs the Run permission; the Anthropic credential comes from Cloudflare’s stored keys, Unified Billing or your own key. Create one gateway per team for per-team spend. The sources we checked show spend reporting and rate limits but no dollar cap, so enforce budgets elsewhere.

The Claude apps gateway sets caps through its admin API once you enable the admin: block (with a write key) in gateway.yaml; amounts are strings of whole US cents, so "40000" is $400:

Terminal window
curl -sS https://claude-gateway.internal.example.com/v1/organizations/spend_limits \
-H "x-api-key: $GATEWAY_ADMIN_WRITE_KEY" \
-H "Content-Type: application/json" \
-d '{"scope": {"type": "rbac_group", "rbac_group_id": "payments"}, "amount": "40000", "period": "monthly"}'

A group cap is a per-seat default each member inherits, not a shared pool, and caps reset on UTC calendar boundaries. Over the cap, a developer gets 429 with billing_error and the reset time; Claude Code v2.1.225 or later warns at 75% and 95%, and v2.1.251 or later adds a Spend limit bar to /usage. This call lists the month’s top spenders for the report (-g stops curl globbing the brackets; a key from admin.read_keys is enough):

Terminal window
curl -sS -g "https://claude-gateway.internal.example.com/v1/organizations/spend_limits/effective?sort=spend_desc&period[]=monthly" \
-H "x-api-key: $GATEWAY_ADMIN_READ_KEY"

Copy-paste prompts for agent cost tracking

Section titled “Copy-paste prompts for agent cost tracking”

How do you know the cost numbers are right?

Section titled “How do you know the cost numbers are right?”

Every figure here is an estimate until reconciled. Check at setup and after every agent or gateway upgrade:

  • Reconcile against the bill monthly. Compare the gateway’s team spend, or the sum of ccusage and /usage, with the Console usage page, the Team or Enterprise spend report, or your cloud billing console. A gap above your tolerance means a machine or CI job that bypasses the gateway, a model the gateway prices at zero, or cache tokens priced wrongly.
  • Report at contracted rates where you can. Claude Code computes /usage, the status line and OpenTelemetry cost at list price. From v2.1.242 an administrator can set the modelPricing managed setting to your contracted rates; /usage then marks the total at your organization's configured rates. The Claude apps gateway has an equivalent pricing.overrides block.
  • Test the stop, not only the count. Keep one test key at a tiny budget and check monthly that it is rejected.
  • Name the owners. The gateway operator owns the smoke test and the version pin, the tech lead the monthly report and budget changes, and finance the reconciliation against the invoice.
SymptomCauseRecovery
Gateway shows no spend while developers workANTHROPIC_BASE_URL set without a gateway credential, so requests still use the claude.ai loginSet ANTHROPIC_AUTH_TOKEN or an apiKeyHelper; confirm with /status
Key spend stays at $0 after real requestsLiteLLM has no price entry for a new modelAdd input_cost_per_token and output_cost_per_token to the model entry; recheck /key/info
Gateway total well below the invoiceCache-read and cache-write tokens not priced, or CI jobs use a direct API keyCompare token columns with ccusage for one day; move CI secrets to gateway keys
Claude Code requests fail with “model not found” behind the gatewayA /model choice or the model pinned in ANTHROPIC_DEFAULT_HAIKU_MODEL is not in model_listAdd the model to model_list, or pin ANTHROPIC_DEFAULT_HAIKU_MODEL to a listed one
Budget resets hours before or after finance’s month closesLiteLLM resets at midnight UTC by defaultSet timezone under litellm_settings; the Claude apps gateway always resets on UTC calendar boundaries
Claude apps gateway lets requests through during an incidentPostgres unreachable, and enforcement fails open by defaultSet enforcement.fail_closed_on_error: true if unmetered spend is worse than an outage

To cut the spend these tools reveal, see context cost optimization, Claude Code cost control and Codex cost management. Plan prices are compared on cost comparison and ROI; the section hub is measuring agentic engineering.