What Did the Agents Cost? ccusage, /usage, Monitors and Gateway Budgets
Agent cost tracking works at three levels. Built-in commands (Claude Code /usage, Codex /status and /usage) show one session or plan. Local trackers such as ccusage total one machine’s cost. Capping a team’s spend takes budgets in the request path: plan or Console spend limits, or a gateway with per-key budgets such as LiteLLM or Anthropic’s Claude apps gateway.
A developer ran three agents in parallel over a weekend, the API invoice jumped, and finance wants per-team numbers you do not have. This page is for the developer who needs to see their own spend today and for the tech lead who has to put a budget on a team without slowing it down.
What you get from this cost-tracking setup
Section titled “What you get from this cost-tracking setup”- A map of which tool answers which cost question, checked on 2026-09-26 against Claude Code 2.1.283, Codex CLI 0.157.1 and ccusage 20.0.24.
- Your own monthly cost in one command, and what that figure means on a subscription versus an API key.
- A live status line and a 5-hour window monitor, so a long agent run does not surprise you.
- A per-team budget with LiteLLM virtual keys for Claude Code and Codex, plus two alternatives.
- A ten-minute monthly cost report, three copy-paste prompts, and the billing trap that makes a gateway count nothing.
This page covers the tools. Organization-wide policy (budget ownership, cost-centre attribution, overruns) lives on AI usage cost governance, and the method for turning spend into cost per accepted change lives on the economics page.
Which tool answers which cost question?
Section titled “Which tool answers which cost question?”Start from the question. Every row was checked on 2026-09-26 against the installed CLIs, the Claude Code costs and gateway spend limits docs, the Codex source at tag rust-v0.157.1, and each project’s GitHub repository.
| Question | Tool | Scope | Can it cap spend? |
|---|---|---|---|
| What did this session cost, and how close am I to my plan limit? | Claude Code /usage; Codex /status and /usage | One session, one plan | No |
| What did I spend this month across every agent on my laptop? | ccusage | One machine, 18 agents | No |
| When will my 5-hour window run out? | ccusage blocks --active, Claude-Code-Usage-Monitor | One machine, Claude Code | No (warns only) |
| How many tokens per repository, even across worktrees? | tokscale | One machine, many agents | No |
| What did each developer spend, in near real time? | Claude Code OpenTelemetry, vendor spend reports | Whole team | No |
| Can I stop one unattended CI run from overspending? | claude -p --max-budget-usd | One headless run | Yes, per run |
| Can I give each team a monthly budget? | LiteLLM virtual keys, Claude apps gateway, your Console or plan spend limits; Cloudflare AI Gateway tracks spend per gateway, rate limits only | Every request through the gateway or account | Yes |
Adoption as of 2026-09-26, from GitHub and the package registries: ccusage has 18.7k stars (npm 20.0.24), LiteLLM 59.6k (PyPI 1.102.1), Portkey’s gateway 13.1k (npm 1.15.2), Claude-Code-Usage-Monitor 8.7k (PyPI claude-monitor 4.0.0) and tokscale 5.5k (npm 4.17.0). None is a plugin or an MCP server, so none adds tokens to the agent’s context.
For a baseline before you set budgets: Anthropic’s costs page (checked 2026-09-26) reports that across enterprise deployments Claude Code averages “around $13 per developer per active day and $150-250 per developer per month”, with 90% of users below $30 per active day. That is a vendor figure for Claude Code only; measure your own team for a month before setting a cap.
See your own spend in two minutes
Section titled “See your own spend in two minutes”Type /usage in a session. The Session block shows the session’s total cost, API and wall-clock duration, code changes and cost by model.
What the dollar figure means depends on how you sign in:
- Pro, Max, Team or Enterprise seat: the Session cost is a local list-price figure; for Pro and Max, Anthropic’s costs page says it “isn’t relevant for billing purposes”. What counts is the plan usage bars on the same screen, with recent usage attributed to skills, subagents, plugins and MCP servers (
dandwswitch between 24 hours and 7 days). - API key or cloud provider: the Session cost estimates real per-token spend; the authoritative figure is on the Console usage page or your cloud billing console.
Since v2.1.211 the totals reset on /clear. Since v2.1.251 a Prompt cache (main) line shows the cache-hit share of input tokens and the cache misses, the usual source of unexplained spend in long sessions.
Codex 0.157.1 has two commands. /status shows “current session configuration and token usage”, and /usage lets you “view account usage or use a usage limit reset” (descriptions from slash_command.rs at rust-v0.157.1). On a ChatGPT plan, usage counts against 5-hour and weekly allowances rather than a dollar bill.
For a dollar estimate across days, ccusage reads Codex’s local logs:
npx ccusage@latest codex daily # also: monthly, sessionThe separate @ccusage/codex package is deprecated.
We could not verify a per-session cost command for Cursor on 2026-09-26. Two routes give numbers:
-
tokscale pulls usage from Cursor’s usage-export API, not local files. Sign in to the Cursor desktop app first:
Terminal window npx tokscale@latest cursor login --name work # reuses the desktop app's sessionnpx tokscale@latest cursor sync --json # caches the export under ~/.config/tokscale/cursor-cache/npx tokscale@latest models --light --client cursor --month # print a table of Cursor usageCursor rows carry no workspace, so they land in
Unknown workspace. -
Cursor’s team dashboard and Analytics API (secondary sources) show usage to Teams and Enterprise admins; confirm what your plan exposes before reporting on them.
ccusage 20.0.24 does not list Cursor among its 18 sources.
Report a month across every agent with ccusage
Section titled “Report a month across every agent with ccusage”ccusage prices the usage logs that coding agents write on your machine; no server, no sign-up:
npx ccusage@latest monthlyThe report opens with a Detected: line naming the agents found (for example Detected: Claude, Codex), then a row per month with a sub-row per agent. Columns: Month, Agent, Models, Input, Output, Cache Create, Cache Read, Total Tokens and Cost (USD). On most agent-heavy machines Cache Read is by far the largest token column, which is why a long session costs more than its prompts suggest.
The commands you will use most, all checked against ccusage --help 20.0.24:
npx ccusage@latest # default: daily report, all detected agentsnpx ccusage@latest monthly --last 3 # the last three months onlynpx ccusage@latest claude daily --instances # Claude Code, split by projectnpx ccusage@latest claude monthly --breakdown # Claude Code, split by modelnpx ccusage@latest blocks --active # current 5-hour block with a projectionnpx ccusage@latest monthly --since 2026-09-01 --json > sept.jsonThe unified monthly has no --breakdown flag (the per-agent subcommands and blocks do); --by-agent adds per-agent breakdowns to the unified JSON.
How do I watch agent spend live?
Section titled “How do I watch agent spend live?”For Claude Code, ccusage has a status-line mode (marked beta in 20.0.24). Add it to ~/.claude/settings.json:
{ "statusLine": { "type": "command", "command": "npx -y ccusage@20.0.24 statusline" }}Pin the version: the status line runs on every refresh, and an unpinned npx -y ccusage would fetch whatever npm publishes next. It runs locally and costs no context; --cost-source cc shows Claude Code’s own session cost instead of ccusage’s calculation.
Two alternatives, both local:
uv tool install claude-monitor # PyPI package; npm "claude-monitor" is a different projectclaude-monitor --plan pro # live TUI of token burn against your plan window
npx tokscale@latest models --light --group-by workspace,model --merge-worktrees --monthKeep claude-monitor in a split pane during a long run. The tokscale line gives one row per repository even across parallel git worktrees, which neither ccusage nor /usage does.
Cap one unattended run
Section titled “Cap one unattended run”For a headless job in CI, Claude Code 2.1.283 has --max-budget-usd <amount>: “Maximum dollar amount to spend on API calls (only works with —print)”.
claude -p "Fix the failing test in packages/billing and run the suite" \ --max-budget-usd 5 \ --permission-mode acceptEdits \ --allowedTools "Bash(npm test *)"claude -p starts in Manual permission mode, so without --permission-mode and --allowedTools the run can neither edit files nor run the suite; headless agents in CI covers the rest of the setup. The budget caps one run, not a team, using the same list-price estimate as /usage. No equivalent flag could be found in codex exec --help 0.157.1.
Choose a gateway for team budgets
Section titled “Choose a gateway for team budgets”Local tools report; they cannot stop anyone. A cap needs every request to pass something that counts spend before forwarding it. Pick the first row that fits.
| You run | Use | Why | Watch out for |
|---|---|---|---|
| Claude Team or Enterprise seats | The plan’s own controls: seat allowance, usage credits and spend limits in Admin settings > Usage | No infrastructure; per-user spend report with CSV export | Usage inside the seat allowance is not metered in dollars |
| Anthropic API through the Console | Workspace spend limits in the Console | Built in; Claude Code gets its own workspace | One “Claude Code” workspace for Console logins; per-team limits need team API keys issued from separate workspaces |
| Claude Code on Bedrock, Claude Platform on AWS, Google Cloud or Foundry | Claude apps gateway (claude gateway --config gateway.yaml) | Per-user, per-group and org caps by day, week or month; SSO; ships in the claude binary | Needs Postgres; enforcement fails open by default if the store is down. Also accepts the Anthropic API as an upstream, so Console shops can use it for per-user SSO caps |
| Several providers, Claude Code and Codex together | LiteLLM proxy | Virtual keys per developer, budgets per key and per team, one bill across providers | Self-hosted; pin the version (see step 1 below) |
| Cloudflare already in front of your traffic | Cloudflare AI Gateway | Managed; logs, caching, rate limits and spend per gateway | Reports spend per gateway; no verified dollar cap |
| A gateway for your own LLM features too | Portkey gateway (npx @portkey-ai/gateway, port 8787) | Routing, fallbacks, guardrails | Claude Code setup unverified (2026-09-26) |
Avoid Helicone for new work: in maintenance mode since March 2026 after Mintlify acquired it (secondary: press and blog sources).
Set a per-team budget with LiteLLM virtual keys
Section titled “Set a per-team budget with LiteLLM virtual keys”The workflow below gives the payments team $400 a month, each developer on it a personal $150 cap, and produces a monthly report. It uses Anthropic models behind LiteLLM; the same keys work for Codex through an OpenAI-format provider. The air-gapped variant with a local model is on gateways and local models.
-
Install a pinned LiteLLM. On 2026-03-24,
litellm1.82.7 and 1.82.8 on PyPI were malicious credential stealers; PyPI removed both, and the maintainers track the incident inBerriAI/litellm#24518. Never install an unpinned proxy on a host that holds provider keys:Terminal window uv tool install 'litellm[proxy]==1.102.1'For a server image, compile a hashed requirements file and install with
pip --require-hashes. The npm packagelitellmis an unrelated JavaScript port. -
Describe the models and the database. Virtual keys and budgets need Postgres (
DATABASE_URL) and a master key. The provider key and the master key come from your secret store:litellm-config.yaml model_list:- model_name: claude-opus-5-5litellm_params:model: anthropic/claude-opus-5-5api_key: os.environ/ANTHROPIC_API_KEYinput_cost_per_token: 0.000004 # $4 per million input tokensoutput_cost_per_token: 0.00002 # $20 per million output tokenscache_read_input_token_cost: 0.0000002 # $0.20 per million cached input tokens- model_name: claude-haiku-4-5litellm_params:model: anthropic/claude-haiku-4-5api_key: os.environ/ANTHROPIC_API_KEYgeneral_settings:master_key: os.environ/LITELLM_MASTER_KEYdatabase_url: os.environ/DATABASE_URLTerminal window litellm --config litellm-config.yaml # listens on port 4000The Opus 5.5 entry carries its own list prices from the models hub because the price map bundled with LiteLLM 1.102.1 predates Opus 5.5. Without them, a proxy that cannot fetch the remote map (an egress-restricted host, or
LITELLM_LOCAL_MODEL_COST_MAP=True) records $0 for every Opus 5.5 request.Terminate TLS in front of the proxy, because keys travel in headers; the examples use
https://llm-gw.internal. The Haiku entry serves Claude Code’s background work (conversation summaries and similar), which step 5 pins to it. List every model your developers pick with/model; unlisted models fail. -
Create the team with a budget.
max_budgetis in US dollars andbudget_durationresets it:Terminal window curl -sS https://llm-gw.internal/team/new \-H "Authorization: Bearer $LITELLM_MASTER_KEY" \-H "Content-Type: application/json" \-d '{"team_alias": "payments", "max_budget": 400, "soft_budget": 300, "budget_duration": "1mo"}'The response contains a
team_id; keep it.soft_budgetblocks nothing; it only triggers LiteLLM’s Slack or email alerts, if configured. In LiteLLM 1.102.1,1mo(and30d) resets at midnight on the first of each month, in UTC unless you settimezoneunderlitellm_settingsto your finance calendar’s zone. -
Issue one key per developer, inside the team. A key’s own cap stops one developer’s weekend fleet from spending the whole team budget:
Terminal window curl -sS https://llm-gw.internal/key/generate \-H "Authorization: Bearer $LITELLM_MASTER_KEY" \-H "Content-Type: application/json" \-d '{"team_id": "TEAM_ID", "key_alias": "payments-anna", "max_budget": 150, "budget_duration": "1mo", "models": ["claude-opus-5-5", "claude-haiku-4-5"]}'TEAM_IDis the value from step 3. Hand the returned key to the developer through your secret manager, never in chat. When a key or its team passesmax_budget, LiteLLM rejects further requests with HTTP 429 (budget_exceeded) until the period resets. -
Point the agents at the gateway.
Put the base URL in managed settings so no developer can forget it, and let each developer supply their own key:
Managed settings: env block {"env": {"ANTHROPIC_BASE_URL": "https://llm-gw.internal","ANTHROPIC_DEFAULT_HAIKU_MODEL": "claude-haiku-4-5"}}Each developer exports
ANTHROPIC_AUTH_TOKENfrom their secret store, or you configure anapiKeyHelperthat reads it from your vault. Without it, the trap above applies. Managed settings are covered on managed policy.Add a provider to
~/.codex/config.tomland select it:~/.codex/config.toml model_provider = "litellm"[model_providers.litellm]name = "LiteLLM"base_url = "https://llm-gw.internal/v1"env_key = "LITELLM_API_KEY"wire_api = "responses"env_keynames the environment variable that holds the developer’s virtual key. Codex speaks the Responses API, so the gateway must serve/v1/responsesfor the models you list. Add the OpenAI models your team uses tomodel_listin step 2, then run one short task as a smoke test.We found no verified way to route Cursor’s agent through your gateway (2026-09-26). Add spend from Cursor’s team dashboard (secondary) to the monthly report by hand, as a separate source.
-
Prove the budget counts before anyone relies on it. On one laptop, start a fresh session, run
/statusand confirm the base URL is the gateway and the credential source is the environment variable, not the claude.ai login. Send one prompt. Then read the key’s spend:Terminal window curl -sS "https://llm-gw.internal/key/info?key=$(printf %s "$DEVELOPER_KEY" | sha256sum | cut -d' ' -f1)" \-H "Authorization: Bearer $LITELLM_MASTER_KEY"The command sends the SHA-256 hash of the step-4 key (LiteLLM 1.102.1 accepts it), so the raw key stays out of URLs and logs. Spend must be above zero and close to the Session cost in
/usage. If it is zero, LiteLLM has no price for that model: addinput_cost_per_tokenandoutput_cost_per_tokento the model entry, using the rates on the models hub. Finally, set a test key to a $0.01 budget and confirm the next request is rejected. -
Report monthly. On the first working day, the tech lead pulls each team’s budget and spend with
GET /team/info?team_id=TEAM_IDand sets it beside the vendor bill; the prompt below automates this.
Route through Cloudflare AI Gateway or cap with the Claude apps gateway
Section titled “Route through Cloudflare AI Gateway or cap with the Claude apps gateway”Cloudflare AI Gateway needs no server. Cloudflare’s Claude Code integration page (checked 2026-09-26) uses three variables; read the token from your secret store:
export ANTHROPIC_BASE_URL="https://gateway.ai.cloudflare.com/v1/$CF_ACCOUNT_ID/$CF_GATEWAY_ID/anthropic"export ANTHROPIC_API_KEY="$CF_AIG_TOKEN"export ANTHROPIC_CUSTOM_HEADERS="cf-aig-authorization: Bearer $CF_AIG_TOKEN"claudeThe gateway must be authenticated and the token needs the Run permission; the Anthropic credential comes from Cloudflare’s stored keys, Unified Billing or your own key. Create one gateway per team for per-team spend. The sources we checked show spend reporting and rate limits but no dollar cap, so enforce budgets elsewhere.
The Claude apps gateway sets caps through its admin API once you enable the admin: block (with a write key) in gateway.yaml; amounts are strings of whole US cents, so "40000" is $400:
curl -sS https://claude-gateway.internal.example.com/v1/organizations/spend_limits \ -H "x-api-key: $GATEWAY_ADMIN_WRITE_KEY" \ -H "Content-Type: application/json" \ -d '{"scope": {"type": "rbac_group", "rbac_group_id": "payments"}, "amount": "40000", "period": "monthly"}'A group cap is a per-seat default each member inherits, not a shared pool, and caps reset on UTC calendar boundaries. Over the cap, a developer gets 429 with billing_error and the reset time; Claude Code v2.1.225 or later warns at 75% and 95%, and v2.1.251 or later adds a Spend limit bar to /usage. This call lists the month’s top spenders for the report (-g stops curl globbing the brackets; a key from admin.read_keys is enough):
curl -sS -g "https://claude-gateway.internal.example.com/v1/organizations/spend_limits/effective?sort=spend_desc&period[]=monthly" \ -H "x-api-key: $GATEWAY_ADMIN_READ_KEY"Copy-paste prompts for agent cost tracking
Section titled “Copy-paste prompts for agent cost tracking”How do you know the cost numbers are right?
Section titled “How do you know the cost numbers are right?”Every figure here is an estimate until reconciled. Check at setup and after every agent or gateway upgrade:
- Reconcile against the bill monthly. Compare the gateway’s team spend, or the sum of ccusage and
/usage, with the Console usage page, the Team or Enterprise spend report, or your cloud billing console. A gap above your tolerance means a machine or CI job that bypasses the gateway, a model the gateway prices at zero, or cache tokens priced wrongly. - Report at contracted rates where you can. Claude Code computes
/usage, the status line and OpenTelemetry cost at list price. From v2.1.242 an administrator can set themodelPricingmanaged setting to your contracted rates;/usagethen marks the totalat your organization's configured rates. The Claude apps gateway has an equivalentpricing.overridesblock. - Test the stop, not only the count. Keep one test key at a tiny budget and check monthly that it is rejected.
- Name the owners. The gateway operator owns the smoke test and the version pin, the tech lead the monthly report and budget changes, and finance the reconciliation against the invoice.
What breaks when you track agent costs?
Section titled “What breaks when you track agent costs?”| Symptom | Cause | Recovery |
|---|---|---|
| Gateway shows no spend while developers work | ANTHROPIC_BASE_URL set without a gateway credential, so requests still use the claude.ai login | Set ANTHROPIC_AUTH_TOKEN or an apiKeyHelper; confirm with /status |
| Key spend stays at $0 after real requests | LiteLLM has no price entry for a new model | Add input_cost_per_token and output_cost_per_token to the model entry; recheck /key/info |
| Gateway total well below the invoice | Cache-read and cache-write tokens not priced, or CI jobs use a direct API key | Compare token columns with ccusage for one day; move CI secrets to gateway keys |
| Claude Code requests fail with “model not found” behind the gateway | A /model choice or the model pinned in ANTHROPIC_DEFAULT_HAIKU_MODEL is not in model_list | Add the model to model_list, or pin ANTHROPIC_DEFAULT_HAIKU_MODEL to a listed one |
| Budget resets hours before or after finance’s month closes | LiteLLM resets at midnight UTC by default | Set timezone under litellm_settings; the Claude apps gateway always resets on UTC calendar boundaries |
| Claude apps gateway lets requests through during an incident | Postgres unreachable, and enforcement fails open by default | Set enforcement.fail_closed_on_error: true if unmetered spend is worse than an outage |
Where to go next with agent cost tracking
Section titled “Where to go next with agent cost tracking”To cut the spend these tools reveal, see context cost optimization, Claude Code cost control and Codex cost management. Plan prices are compared on cost comparison and ROI; the section hub is measuring agentic engineering.