Skip to content

Size an AI plan from measured workload

The right AI coding plan is the cheapest tier that finishes a representative month of agent work at the accepted quality, with few limit interruptions, at the parallelism the team can actually review, and within its data and identity rules. Plan fit is measured from usage logs and accepted pull requests, then re-checked each quarter, because allowances and prices change.

It is Wednesday afternoon, your Claude Code session stops with a usage-limit message halfway through a migration, and the obvious fix is to click Upgrade. Or the opposite: the team pays for the top tier on every seat, and nobody can say whether the extra allowance ever finished a task that the cheaper plan would not. This page is for the developer who pays for or requests their own plan and for the tech lead who approves seat tiers for a team.

Scorecard Q2 · Plan: Which plan best matches your actual workload (limits, session length, parallel agents)?

Maximum-score answer (3 points): “A quarterly plan review uses representative tasks, concurrency, completed-work cost, interruption data, and governance needs.”

  • A one-month usage record built from /usage, ccusage, and an interruption log, instead of a memory of the worst afternoon.
  • A rule that separates a real capacity shortage from a workflow problem that no plan would fix.
  • A decision table that maps each pattern in the record to a plan change, a routing change, or no change.
  • A dated decision record with an owner, a downgrade trigger, and a next review date.
  • Three copy-paste prompts that turn the logs into a recommendation you can check.

Plan prices and what each tier includes live on one page, AI coding tool plan prices. This page names tiers but never repeats their prices, so it stays correct when a price moves.

Collect six signals over a representative month. The first one is a gate: a cheaper plan that lowers quality is not cheaper.

SignalHow to measure itDecision it supports
Accepted-task qualityAgent pull requests merged without material rework, and CI gate pass rateWhether the models the plan gives you do the job at all
Limit interruptionsSessions stopped by a five-hour or weekly limit, logged with the date and the taskWhether more included usage has any value
Cost per accepted taskSubscription share plus any metered usage, divided by accepted tasksWhether a cheaper or a larger tier is more economical
Safe concurrencyParallel agent sessions you can review in a day without stale diffs or merge conflictsWhether parallel capacity is usable or only available
Wait timeTime from a limit message to work resumingWhether waiting for a reset hurts delivery
GovernanceIdentity, data retention, audit, and admin requirements for the code you touchWhether a personal plan is eligible at all

Two rules follow. First, compare cost per accepted task, never the price per token or per seat: a small plan that forces reruns and manual repair can cost more per merged change than a larger one. Second, cap concurrency at review capacity. Every parallel agent on a subscription draws on the same five-hour and weekly windows, so three agents whose output waits two days for review burn the allowance without shipping anything.

Where do you read usage and limits in each tool?

Section titled “Where do you read usage and limits in each tool?”

Each tool reports usage differently, and a dollar figure in a usage view is not always your bill. Read the view in your own account on the day you record it.

Type /usage in a session. On a Pro, Max, Team, or Enterprise plan it shows plan usage bars, activity stats, and a breakdown of what counts against your limits. /usage-credits manages paid usage beyond the plan; it needs a claude.ai subscription login, not an API key.

Pro, Max, Team, and Enterprise seats run on a rolling five-hour window plus a weekly window. On Team and Enterprise, that allowance is shared with Claude chat and Cowork, so your own Claude chat and Cowork use draws on the same seat allowance as your agent runs. With Claude Opus 5.5 on 2026-09-22, Anthropic raised the five-hour limits on Pro, Max, Team, and seat-based Enterprise and gave subscribers a rate-limit reset they can save and use when they choose.

Three choices draw on usage credits or a separate limit rather than the plan’s normal allowance: Fast mode (usage credits only on subscription plans), Claude Fable 5.1 (may bill to usage credits depending on plan and seat tier, with its own weekly limit in /usage), and managed Code Review or Ultrareview runs. Log these separately, because they are routing choices, not signs that the plan is too small.

For the month-long record, read the local logs with ccusage (npm ccusage 20.0.26, checked 2026-09-26):

Terminal window
# Terminal: one month of daily usage from local logs, split by agent
npx ccusage@latest daily --since 2026-09-01 --by-agent --json > usage-2026-09.json
# Claude Code five-hour blocks for the last three days, including the active one
npx ccusage@latest blocks --recent

daily --by-agent already includes Codex; use codex daily only for a Codex-only view, not as an additional total.

The usage views tell you how much you used. They do not tell you which limit message stopped real work. Keep a small log in the repository, one row per interruption, and fill it in when the message appears:

docs/ai/plan-fit-log.csv
date,tool,plan,task_class,parallel_sessions,model,effort,limit_hit,minutes_blocked,workaround,task_accepted,pr
2026-09-09,claude-code,Max 5x,multi-file change,3,claude-opus-5-5,high,five-hour,95,waited,yes,#412
2026-09-17,codex,Plus,review,1,gpt-6-astra,medium,weekly,0,used saved reset,yes,#431

Fill task_class from a fixed list, for example small edit, multi-file change, review, and long-running task, so the rows can be grouped. parallel_sessions and effort matter because they are the two levers that change the draw without a plan change.

Run this once per quarter, and again after any vendor change to plan terms. For one developer it takes an hour. For a team it takes an afternoon, most of it spent reading the numbers together.

  1. Fix the task set. Pick five to ten representative tasks per class from last month’s merged pull requests, including any parallel or long-running work you really do. Keep this set for every future review so the numbers stay comparable.

  2. Collect one month of evidence. Export ccusage JSON, the interruption log, and the list of agent pull requests with their merge status (gh pr list --state merged --limit 500 --search "merged:>=2026-09-01 merged:<2026-10-01" --json number,title,headRefName,labels,mergedAt > prs-2026-09.json). Without --limit, gh returns only 30 pull requests. Identify agent pull requests by their label or branch prefix, for example label:agent.

  3. Separate plan limits from workflow defects. Mark each interruption as capacity, routing (Fast mode, Fable, maximum effort, too many parallel agents), or workflow (vague task, missing context, reruns after failed gates). Only capacity rows justify buying more.

  4. Exclude ineligible plans. Drop any tier that fails your data, identity, retention, or admin requirements before you look at price. A personal plan that fails the governance check is not an option, however cheap.

  5. Check live terms. Read the current price, included usage, overage behaviour, model access, and data terms for the remaining tiers on the vendor page and in your account. Write down the date. The plan price comparison is the starting point, not the final word.

  6. Change one variable for two weeks. Move one tier up or down, or change routing (effort level, model, parallel sessions), not both. Rerun the fixed task set and keep logging.

  7. Record the decision. Commit the decision record below with the evidence files, an owner, a downgrade or escalation trigger, and the next review date.

Which plan change does the evidence support?

Section titled “Which plan change does the evidence support?”

Read the month’s record against this table. Change the plan only when the record matches a row that says so.

Pattern in the recordLikely causeAction
No limit hits, and less than half of the weekly window used in most weeksOversized planTrial one tier down for two weeks with the same task set, for example Max 20x to Max 5x, or a Team Premium seat back to Team Standard.
Five-hour limits hit on most working days, weekly window rarely exhaustedBursty use, often from parallel agentsLower parallel sessions to what you can review, then re-measure. If hits continue at reviewable concurrency, move one tier up.
Weekly limit exhausted before the week ends, on capacity rowsReal shortageMove one tier up (Pro to Max 5x, Plus to Pro, Team Standard seat to Team Premium) and rerun the task set.
Most interruptions are routing rowsFast mode, Fable, or maximum effort used by defaultFix the route first on the model routing page. No plan change.
Many reruns and failed gates before acceptanceWorkflow defectFix task scope, context, and gates. A larger plan pays for the same failures faster.
Heavy, uneven use across a teamSeats sized for the heaviest userSize seats per person: Team Premium for measured heavy users, Team Standard for the rest, or usage-billed Enterprise where the contract fits.
Governed code on a personal planEligibility failureMove to an organization plan before any other change, and see team accounts.

For usage-billed plans, compare against a vendor baseline only as a sanity check. Anthropic’s costs page (code.claude.com, read 2026-09-26) reports an enterprise average of “around $13 per developer per active day and $150-250 per developer per month” for Claude Code. That is Anthropic’s own figure; your month of logs replaces it.

Keep one record per person or team in the repository, next to the log it was built from. A reviewer checks the numbers against the linked files, not the prose.

docs/ai/plan-fit-2026-Q4.yaml
reviewed: 2026-09-26
owner: tech lead, platform team
scope: 6 developers, Claude Code primary, Codex for review
evidence:
usage: usage-2026-09.json
interruptions: docs/ai/plan-fit-log.csv
task_set: evals/plan-fit/tasks.md
terms_checked: 2026-09-26, claude.com/pricing and the admin console
eligible_plans: [Team Standard, Team Premium, Enterprise] # Pro and Max excluded: personal plans, our policy requires org SSO
decision: 4 Team Standard seats, 2 Team Premium seats
reason: two developers hit the weekly window on capacity rows in 3 of 4 weeks at 2 parallel sessions
cost_per_accepted_task: { before: "fill from invoice", after: "fill after trial" }
downgrade_trigger: a Premium seat uses under half its weekly window for 4 straight weeks
escalation_trigger: more than 2 capacity interruptions per developer per week at reviewable concurrency
next_review: 2026-12-15

The thresholds in the triggers are a starting policy, not an industry standard. Set yours from your own first month.

The third prompt asks for missing terms on purpose: a comparison that fills a gap from the model’s memory is the most common way a stale price gets into a decision.

How do you prove a plan change worked without reading every diff?

Section titled “How do you prove a plan change worked without reading every diff?”

A plan change is an experiment, and the evidence already exists in your tools.

  • The task set is the eval. Rerun the same fixed tasks after the change. Their gates (tests, type check, lint) decide pass or fail, exactly as in the model routing eval.
  • CI reports quality. Gate pass rate and the share of agent pull requests merged without rework, using the definitions in lifecycle metrics, must hold or improve after a downgrade.
  • The log reports capacity. Capacity interruptions per developer per week must fall after an upgrade. If they do not, the problem was routing or workflow.
  • The invoice reports cost. Cost per accepted task comes from the invoice and the merged pull requests, not from /usage or ccusage estimates. The cost per accepted change method defines the full formula when finance asks.
  • One person signs off. A developer signs off their own plan. The tech lead owns team seat tiers and the decision record through CODEOWNERS, and finance sees the record before any annual commitment.
SymptomCauseRecovery
The top tier on every seat, and nobody can say whyPlan chosen by brand or by the worst afternoonLog one month, run the review, and trial one tier down on the lightest users first.
An upgrade did not reduce interruptionsThe interruptions were routing or workflow rowsRevert the upgrade at the next billing date, fix the route or the task scope, and re-measure.
Limits hit only on days with many parallel agentsConcurrency above review capacityCap parallel sessions at what gets reviewed that day, and run the rest in isolated worktrees in sequence.
“You’ve hit your session limit” and switching model does not helpFive-hour and weekly windows are shared across all modelsWait for the reset the message shows, or spend a saved reset. Only a model-specific message (“You’ve hit your Opus limit”) is fixed by switching model family with /model.
A Codex limit banner at the five-hour or weekly windowThe ChatGPT plan allowance for that window is used upOpen /usage (Codex CLI 0.156.0 and later) to spend a saved usage-limit reset, or use credits. Log it as a capacity row only if the work stopped at reviewable concurrency.
A Team seat runs out although the agent use is lightThe seat allowance is shared with Claude chat and CoworkInclude chat and Cowork use in the log before you upgrade the seat.
The cheaper plan looks fine, but merged work dropsQuality fell and nobody measured itTreat quality as the gate: restore the previous tier, then try a routing change instead.
An annual commitment made on anecdotesNo trial and no exit conditionTrial monthly with the fixed task set first, and write the downgrade trigger into the record before you sign.
A comparison uses last quarter’s pricesTerms copied from an article or from memoryRe-read the vendor page and your account on the review day, and date every term in the record.
Company code on a personal planGovernance was checked after priceApply identity, retention, and audit requirements first, then compare only the eligible plans.
  • A fixed task set and a one-month measurement period are written down.
  • Usage exports, the interruption log, and merged pull requests are committed or linked.
  • Every interruption is classified as capacity, routing, or workflow.
  • Ineligible plans were excluded before any price comparison.
  • Live plan terms were read on a named date, in the vendor page and the account.
  • Cost per accepted task uses the invoice, not a usage-view estimate.
  • The decision record names an owner, a downgrade trigger, an escalation trigger, and a next review date.
  • Any plan change was validated against the same task set.

Before this page, choose a primary AI engineering harness (Q1), because the plan follows the tool. After it, route each task class to a model by evidence (Q3), which is often the cheaper fix for interruptions.