The economics of agent-built software: cost per accepted change
Cost per accepted change is the full cost of delivering software with coding agents in one period — licenses, metered usage, CI, human review, rework and incidents — divided by the changes that merged and were not reverted. It replaces seat price and time-saved estimates with one unit that finance and engineering can both audit.
The AI line in your budget doubled in two quarters. Engineering says the teams are faster, finance says prove it, and the only numbers on the table are a seat count and a time-saved survey. Neither says what a shipped change costs today compared with last year.
This page is for the executive who signs the budget and the CTO who has to defend it. It is the site’s one ROI method: plan prices per tool and model prices live on their own pages and feed into it.
What you get from the cost-per-accepted-change method
Section titled “What you get from the cost-per-accepted-change method”- A metric you can adopt as written (six cost terms, one denominator, a revert window, an owner per number), the cost curve from Level 2 to Level 5, and a test for when time saved becomes cash.
- A worked example over three quarters, a spreadsheet template for Google Sheets or Excel, and two agent prompts that fill it from your exports.
What goes into cost per accepted change?
Section titled “What goes into cost per accepted change?”The formula has six cost terms and one denominator:
cost per accepted change = (licenses + usage + CI and verification + review + rework + incidents) ÷ accepted changes| Term | What it includes | Where the number comes from | The usual omission |
|---|---|---|---|
| Licenses | Seats and plan fees for every coding tool, including review bots | Invoices | Seats bought for people who stopped using them |
| Usage | Metered tokens, usage credits, overage and API spend from CI runs | Vendor usage exports (see the tabs below) | Agent runs in CI billed to a separate API key |
| CI and verification | CI minutes, ephemeral environments, eval runs, and amortized platform work on the test harness | CI bill, cloud bill, platform team time | Treating harness work as free because salaried people did it |
| Review | Human hours spent reviewing changes (agent- and human-authored), times a loaded hourly cost | Sampled review time, or PR timeline data | Counting it at all |
| Rework | Hours spent fixing changes after merge that did not cause an incident | Follow-up PRs linked to the original | Bugs fixed silently inside the next feature |
| Incidents | Hours of response and repair for incidents traced to a change | Incident postmortems | Incidents without a linked change |
Engineer time spent writing code is deliberately in none of the six terms: it is the capacity agents free, and it belongs on the value side (see why time saved is not cash saved). Review, rework and incident hours are in the numerator because those are the hours that agent output consumes.
What counts as an accepted change?
Section titled “What counts as an accepted change?”Adopt this definition as written, then freeze it for the period:
- Accepted change: a pull request merged to the default branch in the period and not reverted within 30 days of merge.
- Excluded: automated dependency bumps and generated lockfile updates, which inflate the count without carrying intent.
- Included: changes written by people, by agents and by both, because the cost terms cover all of them. Tag agent-authored PRs so you can split the number later.
- Revert window: 30 days. Compute the quarter’s figure 30 days after the quarter closes, not on the last day.
Counting in a GitHub repository takes two commands in a terminal:
# Merged PRs in the quarter (repeat per repository; replace main with the# default branch; exclude bot authors such as Dependabot and Renovate)gh pr list --state merged --base main --limit 1000 \ --search "merged:2026-07-01..2026-09-30 -author:app/dependabot -author:app/renovate" \ --json number --jq length
# Reverts of those PRs, counted after the 30-day window closesgit log main --since=2026-07-01 --until=2026-10-30 --oneline --grep='^Revert "' | wc -lGitHub search returns at most 1,000 results; above that, split the range by month. The revert count is approximate: it misses fix-forwards without the word “Revert” and catches reverts of earlier PRs. Match each revert to its PR, as the first prompt below does, before the number goes to finance. Review 20 follow-up PRs each quarter to see how many fix-forwards the count misses in your repositories.
Why is seat price the wrong unit for agent costs?
Section titled “Why is seat price the wrong unit for agent costs?”A seat price was a fair proxy when every developer used autocomplete about equally. Agents break it in two ways.
First, the bill moves from seats to metered usage. Anthropic’s own cost guide for Claude Code reports an enterprise average of “around $13 per developer per active day and $150-250 per developer per month, with costs remaining below $30 per active day for 90% of users” (Anthropic, Claude Code docs, checked 2026-09-26), and its Enterprise plan is a seat plus usage at API rates. GitHub replaced request-based billing for Copilot with usage-based billing on 2026-06-01, where one GitHub AI Credit costs $0.01 (GitHub docs, checked 2026-09-26). The seat is now the smallest line on the invoice for a heavy user, not the whole of it.
Second, the largest costs never appear on a vendor invoice. Faros AI’s telemetry across 22,000 developers (AI Engineering Report 2026, April 2026) found task throughput per developer up 33.7% alongside incidents per pull request up 242.7% and median time in review up 441.5%. These are vendor figures from a self-selected customer base, but they point at costs a seat price cannot see. DX data agrees: median PR size grew from 44 to 72 lines between July 2025 and June 2026 (Justin Reock, DX newsletter, 2026-06-17).
DORA summarizes its 2025 research in one sentence: “AI improves throughput, but often at the cost of stability if your foundation isn’t solid” (Nathen Harvey, DORA, 2026-01-07). Cost per accepted change captures both halves in one number.
How does the cost curve move from Level 2 to Level 5?
Section titled “How does the cost curve move from Level 2 to Level 5?”The autonomy ladder describes who writes the code and who reads it. Each rung moves the dominant cost to a different line of the formula: from seats to human review at Level 3, then to metered usage and verification capacity at Level 4. The level is a property of the loop that produced a change, not of the whole company; one map of the ladder, the lifecycle and the stations explains how to count a team that runs loops at several levels.
| Level | Who reads the code | The cost line that dominates | What caps accepted changes | What to fund next |
|---|---|---|---|---|
| 2 · paired | A human, every line, while writing it | Licenses; review is folded into authoring | Human writing time | A baseline: measure the six terms before anything changes |
| 3 · reviewer | A human, every diff | Review hours, with usage starting to show | Reviewer hours; Dan Shapiro’s summary of this rung (“The Five Levels”, 2026-01-23) is “Your life is diffs.” | Smaller PRs, review triage by risk, review agents |
| 4 · spec writer | Tests, mostly | Metered usage plus CI and verification | The strength of the test oracle and CI capacity | Stronger oracles, an evidence bundle per change, ephemeral environments |
| 5 · factory | Nobody reads the diff | Usage and verification infrastructure; human cost moves to intent and oracle upkeep | Verification capacity and specification quality | Evals, observability, fast rollback |
The only cost figure for a factory-style run in our evidence is a single user’s: Tim Sehn of DoltHub reported that sixty minutes of an orchestrated multi-agent run “cost me about $100 in Claude tokens”, about ten times a normal Claude Code session per unit of time (DoltHub, 2026-01-15). One person, one setup, but at the top of the ladder usage is a real line, worth paying only when the verification around it holds.
Why is time saved not cash saved?
Section titled “Why is time saved not cash saved?”Time saved is a claim about capacity. Cash saved is a claim about the budget. The first becomes the second only through one of three routes, and each one has to be observable:
| Route | What you must be able to show | Where it appears |
|---|---|---|
| Redeployed capacity | The freed hours shipped roadmap items that were funded and would otherwise have waited | Lead time on committed roadmap items, not total PR count |
| Avoided cost | A contractor not renewed, a hire deferred, an outsourced backlog brought in-house | A line that left the budget |
| Revenue timing | A shorter lead time moved a launch or a contract date | A dated commercial milestone |
If none of the three is visible, report the hours as capacity, not money. Self-reported savings are weak evidence too. METR’s 2025 randomized trial found experienced open-source developers took 19% longer to complete tasks with AI tools (METR, 2025-07-10). Its 2026 follow-up produced point estimates in AI’s favour that are not statistically significant, and METR says selection effects make them hard to interpret (METR, February 2026 update and its published analysis code). DORA’s own ROI work describes a J-curve, with a dip before returns (DORA ROI report, 2026; secondary: seen only through search extracts and InfoQ coverage). Budget for the dip rather than assuming the first quarter pays for itself.
Worked example: one team, three quarters
Section titled “Worked example: one team, three quarters”The numbers below are illustrative assumptions for a ten-engineer team, not benchmarks; take real plan prices from the pricing analysis. The usage line in quarter B sits inside Anthropic’s published $150–250 per developer per month range; quarter C is above it because review agents and CI agent runs are added. Loaded engineer cost is set at $100 an hour.
| Line | Quarter A · Level 2 baseline | Quarter B · Level 3, agents on metered usage | Quarter C · Level 3–4, verification funded |
|---|---|---|---|
| Licenses | $1,200 (10 seats × $40 × 3 months) | $3,000 (10 × $100 × 3) | $3,000 |
| Usage | $0 (inside the seat) | $6,000 (10 × $200 × 3) | $9,000 (adds review-agent and CI agent runs) |
| CI and verification | $1,500 | $4,500 | $12,000 (includes amortized harness work) |
| Review | $30,000 (300 PRs × 1 h) | $45,000 (600 PRs × 0.75 h) | $38,500 (420 low-risk PRs × 0.25 h + 280 × 1 h) |
| Rework | $6,000 (60 h) | $18,000 (180 h) | $10,000 (100 h) |
| Incidents | $2,400 (2 × 12 h) | $4,800 (4 × 12 h) | $2,400 (2 × 12 h) |
| Total | $41,100 | $81,300 | $74,900 |
| Merged / reverted | 300 / 9 | 600 / 30 | 700 / 14 |
| Accepted changes | 291 | 570 | 686 |
| Cost per accepted change | $141 | $143 | $109 |
Three things in this table are the point of the method:
- Quarter B doubled output and did not lower unit cost. Spend rose from $41,100 to $81,300 and accepted changes rose from 291 to 570, so each change cost about the same. Review hours absorbed the gain: the Level 3 ceiling in numbers.
- Quarter C spent more on verification and less on people. CI and verification went from $4,500 to $12,000, and review, rework and incidents together fell from $67,800 to $50,900. Low-risk changes were reviewed against an evidence bundle in 15 minutes. The unit cost fell by about a quarter.
- None of this proves value yet. Quarter C delivered 395 more accepted changes than quarter A. They are worth money only through one of the three routes in the previous section. The value question belongs in the business case; this page gives it an honest cost.
Quarter A’s low unit cost is partly an artefact of the method: it excludes the authoring hours that ten engineers spent writing every line, which is the capacity agents free in quarters B and C. Record those hours on the sheet’s capacity row so the trade stays visible.
Spreadsheet template for cost per accepted change
Section titled “Spreadsheet template for cost per accepted change”Copy the block below and paste it into cell A1 of an empty Google Sheets or Excel sheet. The columns are tab-separated, so each line lands in columns A to C and the formulas calculate. The values are quarter C from the example. Replace the numbers in column B, and keep one sheet per team per quarter. The formulas contain no decimals, so they work unchanged in locales that use a decimal comma.
Metric Value Source or ownerPeriod 2026-Q3 Freeze definitions before the period startsLoaded cost per engineer hour (USD) 100 FinanceLicenses and seats (USD) 3000 Invoices — financeMetered usage (USD) 9000 Vendor usage exports — platform teamCI and verification (USD) 12000 CI and cloud bills, amortized harness workReview hours 385 Sampled review time — engineering leadReview cost (USD) =B7*B3Rework hours 100 Linked follow-up PRs — engineering leadRework cost (USD) =B9*B3Incident hours 24 Postmortems — on-call ownerIncident cost (USD) =B11*B3Total cost (USD) =B4+B5+B6+B8+B10+B12PRs merged (bots excluded) 700 gh pr listReverted within 30 days 14 git log --grepAccepted changes =B14-B15Cost per accepted change (USD) =B13/B16Review share of total cost =B8/B13 Above half: fund review triageUsage share of total cost =B5/B13 Rising with flat unit cost: check verificationAuthoring hours (capacity, not cost) Engineering lead — value side, not in the totalAdd two guardrail rows beside the result, change failure rate and escaped defects, and do not let a lower unit cost average them away. Their definitions live in metrics frameworks for agentic engineering.
Where each number comes from in Claude Code, Codex and Cursor
Section titled “Where each number comes from in Claude Code, Codex and Cursor”Licenses come from invoices for every tool. Usage is where the tools differ, so collect it per tool:
- Per developer:
/usageshows session cost and, on subscription plans, attribution to skills, subagents, plugins and MCP servers. The figure is computed at list price; set themodelPricingmanaged setting (Claude Code v2.1.242 or later) so it matches contracted rates. - Per organization: Team and Enterprise admins export a spend report as CSV from org analytics; on Enterprise, the Enterprise Analytics API returns per-user usage and cost; the spend report covers usage-credit spend only. Console (API) organizations use workspace spend limits and the Claude Code Analytics API.
- Any setup: OpenTelemetry export streams per-user token and cost metrics to your own stack. Cost governance has the configuration.
- CI runs: cap each headless run with
claude -p --max-budget-usd 5 "…"; the flag works only with--print. - Watch for: usage inside a Team or Enterprise seat allowance is not metered in dollars. Record allowance consumption separately or the usage line reads low.
- Per developer:
/usage(Codex CLI 0.156.0 or later) shows account usage;/statusshows estimated thread credits or cost for eligible workspaces (0.148.0 or later). - CI runs:
codex exec --jsonprints each run’s events as JSONL; store them with the PR number. When CI authenticates with an API key, that spend lands on the API project’s bill, not on the ChatGPT plan: add both to the usage line. - Watch for: allowances run on five-hour and weekly windows. A team that buys credits to get past a limit has a usage line even on a flat plan.
- Per team: Cursor’s plans, billing modes and admin reporting change often, so they are not restated here (not verifiable from cursor.com on 2026-09-26). Take usage from your plan’s admin reporting and record each team’s plan and billing mode.
- Review bots: confirm on cursor.com how Bugbot is billed. Anything billed per use belongs in the usage line, even inside a plan.
Review agents are a usage line too. Anthropic’s managed Code Review “averages $15-25” per review, and Ultrareview typically costs $5 to $25 per run (Anthropic, Claude Code docs, checked 2026-09-26): a four-figure line at a few hundred PRs a quarter.
How to check the number before finance sees it
Section titled “How to check the number before finance sees it”Run these steps each quarter, in order:
- Freeze the definitions before the period starts: revert window, bot exclusions, loaded hourly cost and repositories. A mid-period change restarts the comparison.
- Reconcile cost lines with the ledger. Licenses, usage and CI must match what finance paid within a few percent. A gap usually means an API key billed outside the tools’ own reports.
- Sample the hours. Review and rework hours come from a sample of 20 to 30 PRs, timed or reconstructed from PR timelines. Record the sample size beside the number.
- Check the guardrails. Change failure rate and escaped defects must not have worsened. A falling unit cost with a rising failure rate is a loan against the next quarter.
- Sign off in two places. The finance partner signs the cost lines; the engineering lead signs the denominator and the hour samples. The CTO owns the definition and any change to it.
The accepted-change count is only as trustworthy as the checks that decide acceptance: weak CI inflates today’s denominator and next quarter’s rework and incident lines. That is why verifying evidence instead of reading every diff is the prerequisite for this page, not an optional extra.
What goes wrong when you measure cost per accepted change?
Section titled “What goes wrong when you measure cost per accepted change?”Teams split changes to raise the denominator. Recovery: report median PR size beside the unit cost, and compare quarters by the share of committed roadmap items delivered, not by count alone.
Allowances hide the usage line. Seats absorb usage until someone buys credits, so early quarters look cheap. Recovery: record allowance consumption per team from the admin reports each month, even when the invoice does not change.
Review hours are guessed, not sampled. Recovery: time 20 PRs per quarter end to end, and show the sample size on the sheet.
Incidents are not traced to changes. Recovery: add a “causing change” field to the postmortem template, with “unknown” as an allowed answer, and report the unknown share.
Teams are ranked against each other. Different codebases and risk levels make cross-team unit costs misleading. Recovery: compare each team with its own previous quarters.
Hours saved are monetized in the board deck. A survey times salary becomes “savings” finance cannot find in the budget. Recovery: move it to a capacity line, and name the redeployment, avoided cost or revenue date that would turn it into money.
Where to go next with agent economics
Section titled “Where to go next with agent economics”The CTO and executive tracks both reach this page after the evidence steps; on the CTO track, the operating model comes just before it. Read how to verify evidence instead of diffs first if you have not. From here, executives turn the baseline into a funded decision; CTOs set the autonomy policy it pays for.
Frequently asked questions
What is cost per accepted change?
It is the full cost of delivering software with coding agents in one period (licenses and seats, metered usage, CI and verification, review hours, rework hours and incident hours) divided by the number of changes that merged and were not reverted within 30 days.
Why is time saved not the same as money saved?
Hours freed from writing code become money only when they are redeployed to work that has value, when they avoid a real cost such as a contractor or a deferred hire, or when a shorter lead time earns revenue sooner. Otherwise the saving stays on a timesheet and never reaches the budget.
Does moving to more autonomous agents make each change cheaper?
Only when verification capacity grows with it. More generation without stronger tests, evidence and review triage turns into review hours, rework and incidents, which leaves the cost per accepted change flat or higher even while output rises.