Economics: cost per accepted change
Turn a reproduced pattern into a cost figure a CFO accepts.
The AI coding case studies that survive scrutiny are few and dated: Stripe’s Minions (over 1,300 agent-written pull requests merged weekly, February 2026), Microsoft’s study of its Claude Code and GitHub Copilot CLI rollout (about 24% more merged pull requests, July 2026), and Anthropic’s own teams. Each scaled on verification, not on the agent alone.
Your board forwarded a deck that promises “10x engineering output”, and the CFO wants to know whether the company should bet on it. Every example in the deck is a logo, a round number and no method. This page is for the CTO or executive who has to decide which stories to believe, what they prove, and what to copy first, and for the tech lead who then runs and signs off the first reproduction.
A case study earns a place in a plan when it names what was measured, over whom, when, and who published it. The table holds the ones that meet that bar as of 2026-09-26. Every one of them is vendor-internal or vendor-adjacent, so treat each as a fact about that company, not a forecast for yours.
| Case | What was measured | Result | Publisher, date | Caveat |
|---|---|---|---|---|
| Stripe Minions | Pull requests merged per week with no human-written code | “Over 1,300” per week, “human-reviewed” | Stripe, Alistair Gray, Minions, Part 2, 2026-02-19 | One company; a count of pull requests, not of value delivered |
| Microsoft’s CLI-agent rollout | Merged pull requests of adopters against a counterfactual | “roughly 24% more pull requests”, lift persisting across four months | Murphy-Hill, Butler and Savelieva, arXiv 2607.01418, 2026-07-01 | The authors’ own words: “a merged PR is not the same as the value it delivers” |
| Anthropic, merged code | Share of lines merged to production attributable to Claude | “more than 80%” as of May 2026, up from “low single digits” before February 2025 | Anthropic, When AI builds itself (the >80% sentence is dated “As of May 2026”; the page carries a 2026-09-18 update on LLM-judged session success, which is not a quality measure; read 2026-09-26) | The vendor measuring its own product; lines, not work |
| Anthropic, internal survey | Self-reported share of work and productivity gain | Claude used in 60% of work, a “50% productivity boost” | Anthropic, How AI is transforming work at Anthropic, 2025-12-02 | 132 engineers and researchers surveyed; self-report, not telemetry |
| Anthropic teams’ workflows | Concrete workflows, with a few timings | A 20-minute saving during an outage; up to 100 ad variations per batch | Anthropic, How Anthropic teams use Claude Code, 2025-07-24 | Anecdotes; useful for the workflow, not for a forecast |
| Google, share of new code | Share of new code that is AI-generated | 75% | Sundar Pichai at Google Cloud Next, as reported by Fast Company, 2026-04-24 | Secondary: no Google-owned transcript was reached; cite it as press-reported |
The industry-wide figure to set beside these is DX’s: across data from over 400 companies in Q2 2026, “on average, 51.9% of code is now AI-authored” (DX, 2026-06-17). That is a share of code, self-reported, and says nothing about whether the code was worth writing.
Stripe’s Minions posts give a detailed public account of agents shipping production code at scale, and the detail is all about verification. Alistair Gray’s two posts (2026-02-09 and 2026-02-19) describe four ingredients:
Every pull request is still human-reviewed. What changed is what the reviewer receives: a branch that has already passed a test suite of millions of tests, not a raw diff. That is the pattern this site calls evidence, not diffs.
The Microsoft paper is the largest adoption study in the table: “tens of thousands of engineers at Microsoft over its early-2026 rollout” of Claude Code and GitHub Copilot CLI. Adopters “merged roughly 24% more pull requests than they would have otherwise”, and “the lift persists across our four-month window”.
Two findings matter more for a rollout plan than the headline. First, “first use spread primarily through social networks”: colleagues, not mandates, drove adoption, which is why team adoption starts with champions. Second, the authors measured merged pull requests as an output proxy and said so. Your plan should inherit that honesty: count accepted changes, and price them with cost per accepted change.
Anthropic publishes two kinds of evidence about itself, and they answer different questions.
The workflow stories (2025-07-24) show what people typed. When Kubernetes clusters stopped scheduling pods, the Data Infrastructure team “fed it dashboard screenshots”, Claude Code guided them through Google Cloud’s UI to pod IP address exhaustion and gave “the exact commands to create a new IP pool”, “saving them 20 minutes of valuable time during a system outage.” Growth Marketing built a workflow that “processes CSV files with hundreds of ads, identifies underperformers, and generates new variations”, plus a Figma plugin that generates “up to 100 ad variations” per batch. Data scientists without TypeScript fluency built “entire React applications for visualizing RL model performance”.
The aggregate numbers show scale and carry their own warnings. The >80% merged-lines figure is Anthropic’s conservative measure; the page itself notes that leadership’s “90% or more” estimate includes scripts and experimental code. Before that figure existed, Redwood Research argued that “the productivity boost at a given fraction of code generated isn’t that high because AI allows people to cheaply generate lots of very low value code” (Ryan Greenblatt, 2025-10-22). Quote the share of code with that critique beside it, or not at all.
A plan built only on the success stories fails the first informed question. Put these beside them:
The success stories and the counter-evidence agree on one thing. DORA puts it plainly: “Without robust control systems, like strong automated testing, mature version control practices, and fast feedback loops, an increase in change volume leads to instability.” Stripe had the control system first. The state of agentic engineering in 2026 covers the full evidence base.
| Pattern | Where the evidence shows it | What to copy |
|---|---|---|
| The oracle came before the autonomy | Stripe’s three million tests; DORA’s “strong automated testing” | Measure how strong your tests are before widening what agents may merge |
| The loop has a hard stop | Stripe’s “at most two rounds of CI” | Cap retries and spend; hand back to a human with a report |
| A human signs off at merge, on evidence | Every Minions pull request is human-reviewed | Ship an evidence bundle with each agent pull request |
| Context and tools are wired once | Stripe’s Toolshed of nearly 500 MCP tools | Stand up shared MCP servers and context files for the whole team |
| The claim names its metric and its caveat | Microsoft’s merged-PR proxy, stated as a proxy | Report accepted changes and their cost, never lines or ”% AI-written” |
Run every vendor story, conference talk or analyst quote through these seven questions. A story that fails three or more is marketing; keep it out of the business case.
Copy the pattern, not the number. Pick one repetitive, well-tested task type from your backlog (a dependency bump, a flaky-test fix, a deprecated-API migration) and run the Stripe shape on it: a written task, an agent run with a hard stop, the test suite as the oracle, and a human review of the evidence. The run takes an afternoon; the measurement design belongs to a pilot that proves something.
The task file is the same in every tool. Save it as .agent/tasks/deprecated-api.md:
Replace every call to legacyFetch() in src/ with httpClient.get(), keepingbehaviour identical. Acceptance criteria:1. `npm test` passes.2. No file outside src/ and test/ changes.3. No existing test is deleted or weakened; add a test for any call site that had none.Process: run `npm test` after your change. If it fails, fix and run itonce more. If it still fails after the second run, stop and writeREPORT.md with the failing tests and your best diagnosis instead ofcontinuing. Always finish by writing REPORT.md: files changed, testsadded, test results, and anything you were unsure about.Run it headless from the repository root in a terminal. --max-budget-usd caps spend (it works only with --print), and --permission-prompts none denies anything that would prompt, so --allowedTools lets through only the test command.
claude -p "$(cat .agent/tasks/deprecated-api.md)" \ --permission-mode acceptEdits \ --permission-prompts none \ --allowedTools "Bash(npm test *)" \ --max-budget-usd 5 \ --output-format json > run.jsonTo run the same task on every pull request, see scripting Claude Code.
Run it non-interactively in a terminal. The :workspace permission profile, which OpenAI prefers, keeps writes inside the repository, and -o saves the agent’s final message for the reviewer. Permission profiles are beta and need Codex CLI 0.138.0 or later.
codex exec -c default_permissions=":workspace" \ -o last-message.md \ "$(cat .agent/tasks/deprecated-api.md)"On an older CLI, use the legacy sandbox instead: replace the -c option with --sandbox workspace-write. Use one or the other, not both: OpenAI says the profile and legacy sandbox systems “do not compose”.
Codex has no --full-auto flag (checked in codex exec --help, v0.157.1). For scheduled and cloud runs, see Codex automation.
Open the repository in Cursor, switch the agent to Plan Mode, and paste the task file. Approve the plan, then start the build in a worktree (Cursor’s Worktrees feature lets Agent work in isolated Git checkouts), so the change stays out of your working copy. Nothing in Cursor enforces the two-run cap here: it lives in the prompt alone, as the verification section below explains. To run the same task type unattended, use Cloud Agents; see Cursor automation.
The reviewer checks evidence, not every line of the diff. For each run:
Tests are the oracle. The run counts only if npm test passes and the reviewer prompt finds no weakened test. If the suite is thin, the pattern is not ready for this task type; strengthen the tests first.
Scope is checked mechanically. Run this in the terminal after the run. It lists every changed or new file outside src/ and test/ (the task file under .agent/ and the run’s own output files excepted), and any output means the run fails:
{ git diff --name-only HEAD; git ls-files --others --exclude-standard; } \ | grep -vE '^(src|test)/|^\.agent/|^(REPORT\.md|run\.json|last-message\.md)$' \ && echo 'SCOPE VIOLATION' || echo 'scope ok'The stop rule is counted. A run that hit its second failure produces REPORT.md and no pull request. Count these; they are your failure rate. Be clear about what enforces it: the two-run cap lives only in the prompt, so an agent can ignore it. For a hard limit, wrap the run, for example timeout 20m claude -p …, or run it in a CI job that counts test runs and kills the job at the third. Of the commands on this page, --max-budget-usd in the Claude Code tab is the only hard spend cap; the Codex and Cursor runs have none unless you add one.
A named person signs off. The tech lead who owns the repository merges or rejects, and records the verdict.
Quality is tracked after merge. Log reverts and incidents traced to agent changes for 30 days, then compute cost per accepted change.
After 10 to 20 runs of one task type, you have an internal case study with a baseline, a sample, a verification method and a quality number, which is more than most vendor decks carry. Write it up with this skeleton and file it beside the business case:
Internal case study: <task type>, <team>, <dates>Question: can agents complete <task type> to our merge standard?Baseline: median hours per task before the pilot (n = ?, source)Runs: n = ?, merged = ?, stopped by the retry cap = ?, rejected at review = ?Verification: test suite (size, mutation score if known), scope check, reviewerQuality after 30 days: reverts = ?, incidents = ?Cost: agent spend + review hours, per accepted changeCaveats: what this does not show (other task types, other teams)Decision: expand / repeat / stop, and who decidedEconomics: cost per accepted change
Turn a reproduced pattern into a cost figure a CFO accepts.
Design a pilot that proves something
Baselines, cohorts and a decision rule written before the pilot starts.
Build the business case
A decision memo that uses your evidence, not a vendor multiplier.
Report to the board
A quarterly one-page template, including what not to claim.
Stripe's Minions (over 1,300 agent-written pull requests merged each week, February 2026), Microsoft's study of its early-2026 Claude Code and GitHub Copilot CLI rollout (roughly 24% more merged pull requests, July 2026), and Anthropic's published accounts of its own teams.
Each scaled on verification: an existing test suite as the oracle, a hard limit on how often the agent may retry, and a human who reviews before merge. None of them removed the check; they made the check automatic.
Only as a dated, attributed data point about that company. Its number depended on its own test suite and review process, so measure your own baseline in a pilot before forecasting anything.