Skip to content

The economics of agent-built software: cost per accepted change

Cost per accepted change is the full cost of delivering software with coding agents in one period — licenses, metered usage, CI, human review, rework and incidents — divided by the changes that merged and were not reverted. It replaces seat price and time-saved estimates with one unit that finance and engineering can both audit.

The AI line in your budget doubled in two quarters. Engineering says the teams are faster, finance says prove it, and the only numbers on the table are a seat count and a time-saved survey. Neither says what a shipped change costs today compared with last year.

This page is for the executive who signs the budget and the CTO who has to defend it. It is the site’s one ROI method: plan prices per tool and model prices live on their own pages and feed into it.

What you get from the cost-per-accepted-change method

Section titled “What you get from the cost-per-accepted-change method”
  • A metric you can adopt as written (six cost terms, one denominator, a revert window, an owner per number), the cost curve from Level 2 to Level 5, and a test for when time saved becomes cash.
  • A worked example over three quarters, a spreadsheet template for Google Sheets or Excel, and two agent prompts that fill it from your exports.

The formula has six cost terms and one denominator:

cost per accepted change =
(licenses + usage + CI and verification + review + rework + incidents)
÷ accepted changes
TermWhat it includesWhere the number comes fromThe usual omission
LicensesSeats and plan fees for every coding tool, including review botsInvoicesSeats bought for people who stopped using them
UsageMetered tokens, usage credits, overage and API spend from CI runsVendor usage exports (see the tabs below)Agent runs in CI billed to a separate API key
CI and verificationCI minutes, ephemeral environments, eval runs, and amortized platform work on the test harnessCI bill, cloud bill, platform team timeTreating harness work as free because salaried people did it
ReviewHuman hours spent reviewing changes (agent- and human-authored), times a loaded hourly costSampled review time, or PR timeline dataCounting it at all
ReworkHours spent fixing changes after merge that did not cause an incidentFollow-up PRs linked to the originalBugs fixed silently inside the next feature
IncidentsHours of response and repair for incidents traced to a changeIncident postmortemsIncidents without a linked change

Engineer time spent writing code is deliberately in none of the six terms: it is the capacity agents free, and it belongs on the value side (see why time saved is not cash saved). Review, rework and incident hours are in the numerator because those are the hours that agent output consumes.

Adopt this definition as written, then freeze it for the period:

  • Accepted change: a pull request merged to the default branch in the period and not reverted within 30 days of merge.
  • Excluded: automated dependency bumps and generated lockfile updates, which inflate the count without carrying intent.
  • Included: changes written by people, by agents and by both, because the cost terms cover all of them. Tag agent-authored PRs so you can split the number later.
  • Revert window: 30 days. Compute the quarter’s figure 30 days after the quarter closes, not on the last day.

Counting in a GitHub repository takes two commands in a terminal:

Terminal window
# Merged PRs in the quarter (repeat per repository; replace main with the
# default branch; exclude bot authors such as Dependabot and Renovate)
gh pr list --state merged --base main --limit 1000 \
--search "merged:2026-07-01..2026-09-30 -author:app/dependabot -author:app/renovate" \
--json number --jq length
# Reverts of those PRs, counted after the 30-day window closes
git log main --since=2026-07-01 --until=2026-10-30 --oneline --grep='^Revert "' | wc -l

GitHub search returns at most 1,000 results; above that, split the range by month. The revert count is approximate: it misses fix-forwards without the word “Revert” and catches reverts of earlier PRs. Match each revert to its PR, as the first prompt below does, before the number goes to finance. Review 20 follow-up PRs each quarter to see how many fix-forwards the count misses in your repositories.

Why is seat price the wrong unit for agent costs?

Section titled “Why is seat price the wrong unit for agent costs?”

A seat price was a fair proxy when every developer used autocomplete about equally. Agents break it in two ways.

First, the bill moves from seats to metered usage. Anthropic’s own cost guide for Claude Code reports an enterprise average of “around $13 per developer per active day and $150-250 per developer per month, with costs remaining below $30 per active day for 90% of users” (Anthropic, Claude Code docs, checked 2026-09-26), and its Enterprise plan is a seat plus usage at API rates. GitHub replaced request-based billing for Copilot with usage-based billing on 2026-06-01, where one GitHub AI Credit costs $0.01 (GitHub docs, checked 2026-09-26). The seat is now the smallest line on the invoice for a heavy user, not the whole of it.

Second, the largest costs never appear on a vendor invoice. Faros AI’s telemetry across 22,000 developers (AI Engineering Report 2026, April 2026) found task throughput per developer up 33.7% alongside incidents per pull request up 242.7% and median time in review up 441.5%. These are vendor figures from a self-selected customer base, but they point at costs a seat price cannot see. DX data agrees: median PR size grew from 44 to 72 lines between July 2025 and June 2026 (Justin Reock, DX newsletter, 2026-06-17).

DORA summarizes its 2025 research in one sentence: “AI improves throughput, but often at the cost of stability if your foundation isn’t solid” (Nathen Harvey, DORA, 2026-01-07). Cost per accepted change captures both halves in one number.

How does the cost curve move from Level 2 to Level 5?

Section titled “How does the cost curve move from Level 2 to Level 5?”

The autonomy ladder describes who writes the code and who reads it. Each rung moves the dominant cost to a different line of the formula: from seats to human review at Level 3, then to metered usage and verification capacity at Level 4. The level is a property of the loop that produced a change, not of the whole company; one map of the ladder, the lifecycle and the stations explains how to count a team that runs loops at several levels.

LevelWho reads the codeThe cost line that dominatesWhat caps accepted changesWhat to fund next
2 · pairedA human, every line, while writing itLicenses; review is folded into authoringHuman writing timeA baseline: measure the six terms before anything changes
3 · reviewerA human, every diffReview hours, with usage starting to showReviewer hours; Dan Shapiro’s summary of this rung (“The Five Levels”, 2026-01-23) is “Your life is diffs.”Smaller PRs, review triage by risk, review agents
4 · spec writerTests, mostlyMetered usage plus CI and verificationThe strength of the test oracle and CI capacityStronger oracles, an evidence bundle per change, ephemeral environments
5 · factoryNobody reads the diffUsage and verification infrastructure; human cost moves to intent and oracle upkeepVerification capacity and specification qualityEvals, observability, fast rollback

The only cost figure for a factory-style run in our evidence is a single user’s: Tim Sehn of DoltHub reported that sixty minutes of an orchestrated multi-agent run “cost me about $100 in Claude tokens”, about ten times a normal Claude Code session per unit of time (DoltHub, 2026-01-15). One person, one setup, but at the top of the ladder usage is a real line, worth paying only when the verification around it holds.

Time saved is a claim about capacity. Cash saved is a claim about the budget. The first becomes the second only through one of three routes, and each one has to be observable:

RouteWhat you must be able to showWhere it appears
Redeployed capacityThe freed hours shipped roadmap items that were funded and would otherwise have waitedLead time on committed roadmap items, not total PR count
Avoided costA contractor not renewed, a hire deferred, an outsourced backlog brought in-houseA line that left the budget
Revenue timingA shorter lead time moved a launch or a contract dateA dated commercial milestone

If none of the three is visible, report the hours as capacity, not money. Self-reported savings are weak evidence too. METR’s 2025 randomized trial found experienced open-source developers took 19% longer to complete tasks with AI tools (METR, 2025-07-10). Its 2026 follow-up produced point estimates in AI’s favour that are not statistically significant, and METR says selection effects make them hard to interpret (METR, February 2026 update and its published analysis code). DORA’s own ROI work describes a J-curve, with a dip before returns (DORA ROI report, 2026; secondary: seen only through search extracts and InfoQ coverage). Budget for the dip rather than assuming the first quarter pays for itself.

The numbers below are illustrative assumptions for a ten-engineer team, not benchmarks; take real plan prices from the pricing analysis. The usage line in quarter B sits inside Anthropic’s published $150–250 per developer per month range; quarter C is above it because review agents and CI agent runs are added. Loaded engineer cost is set at $100 an hour.

LineQuarter A · Level 2 baselineQuarter B · Level 3, agents on metered usageQuarter C · Level 3–4, verification funded
Licenses$1,200 (10 seats × $40 × 3 months)$3,000 (10 × $100 × 3)$3,000
Usage$0 (inside the seat)$6,000 (10 × $200 × 3)$9,000 (adds review-agent and CI agent runs)
CI and verification$1,500$4,500$12,000 (includes amortized harness work)
Review$30,000 (300 PRs × 1 h)$45,000 (600 PRs × 0.75 h)$38,500 (420 low-risk PRs × 0.25 h + 280 × 1 h)
Rework$6,000 (60 h)$18,000 (180 h)$10,000 (100 h)
Incidents$2,400 (2 × 12 h)$4,800 (4 × 12 h)$2,400 (2 × 12 h)
Total$41,100$81,300$74,900
Merged / reverted300 / 9600 / 30700 / 14
Accepted changes291570686
Cost per accepted change$141$143$109

Three things in this table are the point of the method:

  1. Quarter B doubled output and did not lower unit cost. Spend rose from $41,100 to $81,300 and accepted changes rose from 291 to 570, so each change cost about the same. Review hours absorbed the gain: the Level 3 ceiling in numbers.
  2. Quarter C spent more on verification and less on people. CI and verification went from $4,500 to $12,000, and review, rework and incidents together fell from $67,800 to $50,900. Low-risk changes were reviewed against an evidence bundle in 15 minutes. The unit cost fell by about a quarter.
  3. None of this proves value yet. Quarter C delivered 395 more accepted changes than quarter A. They are worth money only through one of the three routes in the previous section. The value question belongs in the business case; this page gives it an honest cost.

Quarter A’s low unit cost is partly an artefact of the method: it excludes the authoring hours that ten engineers spent writing every line, which is the capacity agents free in quarters B and C. Record those hours on the sheet’s capacity row so the trade stays visible.

Spreadsheet template for cost per accepted change

Section titled “Spreadsheet template for cost per accepted change”

Copy the block below and paste it into cell A1 of an empty Google Sheets or Excel sheet. The columns are tab-separated, so each line lands in columns A to C and the formulas calculate. The values are quarter C from the example. Replace the numbers in column B, and keep one sheet per team per quarter. The formulas contain no decimals, so they work unchanged in locales that use a decimal comma.

Metric Value Source or owner
Period 2026-Q3 Freeze definitions before the period starts
Loaded cost per engineer hour (USD) 100 Finance
Licenses and seats (USD) 3000 Invoices — finance
Metered usage (USD) 9000 Vendor usage exports — platform team
CI and verification (USD) 12000 CI and cloud bills, amortized harness work
Review hours 385 Sampled review time — engineering lead
Review cost (USD) =B7*B3
Rework hours 100 Linked follow-up PRs — engineering lead
Rework cost (USD) =B9*B3
Incident hours 24 Postmortems — on-call owner
Incident cost (USD) =B11*B3
Total cost (USD) =B4+B5+B6+B8+B10+B12
PRs merged (bots excluded) 700 gh pr list
Reverted within 30 days 14 git log --grep
Accepted changes =B14-B15
Cost per accepted change (USD) =B13/B16
Review share of total cost =B8/B13 Above half: fund review triage
Usage share of total cost =B5/B13 Rising with flat unit cost: check verification
Authoring hours (capacity, not cost) Engineering lead — value side, not in the total

Add two guardrail rows beside the result, change failure rate and escaped defects, and do not let a lower unit cost average them away. Their definitions live in metrics frameworks for agentic engineering.

Where each number comes from in Claude Code, Codex and Cursor

Section titled “Where each number comes from in Claude Code, Codex and Cursor”

Licenses come from invoices for every tool. Usage is where the tools differ, so collect it per tool:

  • Per developer: /usage shows session cost and, on subscription plans, attribution to skills, subagents, plugins and MCP servers. The figure is computed at list price; set the modelPricing managed setting (Claude Code v2.1.242 or later) so it matches contracted rates.
  • Per organization: Team and Enterprise admins export a spend report as CSV from org analytics; on Enterprise, the Enterprise Analytics API returns per-user usage and cost; the spend report covers usage-credit spend only. Console (API) organizations use workspace spend limits and the Claude Code Analytics API.
  • Any setup: OpenTelemetry export streams per-user token and cost metrics to your own stack. Cost governance has the configuration.
  • CI runs: cap each headless run with claude -p --max-budget-usd 5 "…"; the flag works only with --print.
  • Watch for: usage inside a Team or Enterprise seat allowance is not metered in dollars. Record allowance consumption separately or the usage line reads low.

Review agents are a usage line too. Anthropic’s managed Code Review “averages $15-25” per review, and Ultrareview typically costs $5 to $25 per run (Anthropic, Claude Code docs, checked 2026-09-26): a four-figure line at a few hundred PRs a quarter.

How to check the number before finance sees it

Section titled “How to check the number before finance sees it”

Run these steps each quarter, in order:

  1. Freeze the definitions before the period starts: revert window, bot exclusions, loaded hourly cost and repositories. A mid-period change restarts the comparison.
  2. Reconcile cost lines with the ledger. Licenses, usage and CI must match what finance paid within a few percent. A gap usually means an API key billed outside the tools’ own reports.
  3. Sample the hours. Review and rework hours come from a sample of 20 to 30 PRs, timed or reconstructed from PR timelines. Record the sample size beside the number.
  4. Check the guardrails. Change failure rate and escaped defects must not have worsened. A falling unit cost with a rising failure rate is a loan against the next quarter.
  5. Sign off in two places. The finance partner signs the cost lines; the engineering lead signs the denominator and the hour samples. The CTO owns the definition and any change to it.

The accepted-change count is only as trustworthy as the checks that decide acceptance: weak CI inflates today’s denominator and next quarter’s rework and incident lines. That is why verifying evidence instead of reading every diff is the prerequisite for this page, not an optional extra.

What goes wrong when you measure cost per accepted change?

Section titled “What goes wrong when you measure cost per accepted change?”

Teams split changes to raise the denominator. Recovery: report median PR size beside the unit cost, and compare quarters by the share of committed roadmap items delivered, not by count alone.

Allowances hide the usage line. Seats absorb usage until someone buys credits, so early quarters look cheap. Recovery: record allowance consumption per team from the admin reports each month, even when the invoice does not change.

Review hours are guessed, not sampled. Recovery: time 20 PRs per quarter end to end, and show the sample size on the sheet.

Incidents are not traced to changes. Recovery: add a “causing change” field to the postmortem template, with “unknown” as an allowed answer, and report the unknown share.

Teams are ranked against each other. Different codebases and risk levels make cross-team unit costs misleading. Recovery: compare each team with its own previous quarters.

Hours saved are monetized in the board deck. A survey times salary becomes “savings” finance cannot find in the budget. Recovery: move it to a capacity line, and name the redeployment, avoided cost or revenue date that would turn it into money.

The CTO and executive tracks both reach this page after the evidence steps; on the CTO track, the operating model comes just before it. Read how to verify evidence instead of diffs first if you have not. From here, executives turn the baseline into a funded decision; CTOs set the autonomy policy it pays for.

Frequently asked questions

What is cost per accepted change?

It is the full cost of delivering software with coding agents in one period (licenses and seats, metered usage, CI and verification, review hours, rework hours and incident hours) divided by the number of changes that merged and were not reverted within 30 days.

Why is time saved not the same as money saved?

Hours freed from writing code become money only when they are redeployed to work that has value, when they avoid a real cost such as a contractor or a deferred hire, or when a shorter lead time earns revenue sooner. Otherwise the saving stays on a timesheet and never reaches the budget.

Does moving to more autonomous agents make each change cheaper?

Only when verification capacity grows with it. More generation without stronger tests, evidence and review triage turns into review hours, rework and incidents, which leaves the cost per accepted change flat or higher even while output rises.