Skip to content

AI coding success stories: the case studies that hold up

The AI coding case studies that survive scrutiny are few and dated: Stripe’s Minions (over 1,300 agent-written pull requests merged weekly, February 2026), Microsoft’s study of its Claude Code and GitHub Copilot CLI rollout (about 24% more merged pull requests, July 2026), and Anthropic’s own teams. Each scaled on verification, not on the agent alone.

Your board forwarded a deck that promises “10x engineering output”, and the CFO wants to know whether the company should bet on it. Every example in the deck is a logo, a round number and no method. This page is for the CTO or executive who has to decide which stories to believe, what they prove, and what to copy first, and for the tech lead who then runs and signs off the first reproduction.

What can you take from AI coding case studies into your plan?

Section titled “What can you take from AI coding case studies into your plan?”
  • The verifiable case studies in one table, each with its publisher, date, what it measured and its caveat, ready to paste into a planning memo
  • The counter-evidence at full strength, so the plan survives the first skeptic who has read METR or Faros
  • A seven-question checklist and a copy-paste prompt to vet any case study a vendor sends you
  • One pattern from the evidence (the bounded agent loop) that a team can reproduce on its own backlog this week, in Claude Code, Codex or Cursor
  • A template for writing up your own internal case study so that it counts as evidence

Which AI coding case studies can you trust?

Section titled “Which AI coding case studies can you trust?”

A case study earns a place in a plan when it names what was measured, over whom, when, and who published it. The table holds the ones that meet that bar as of 2026-09-26. Every one of them is vendor-internal or vendor-adjacent, so treat each as a fact about that company, not a forecast for yours.

CaseWhat was measuredResultPublisher, dateCaveat
Stripe MinionsPull requests merged per week with no human-written code“Over 1,300” per week, “human-reviewed”Stripe, Alistair Gray, Minions, Part 2, 2026-02-19One company; a count of pull requests, not of value delivered
Microsoft’s CLI-agent rolloutMerged pull requests of adopters against a counterfactual“roughly 24% more pull requests”, lift persisting across four monthsMurphy-Hill, Butler and Savelieva, arXiv 2607.01418, 2026-07-01The authors’ own words: “a merged PR is not the same as the value it delivers”
Anthropic, merged codeShare of lines merged to production attributable to Claude“more than 80%” as of May 2026, up from “low single digits” before February 2025Anthropic, When AI builds itself (the >80% sentence is dated “As of May 2026”; the page carries a 2026-09-18 update on LLM-judged session success, which is not a quality measure; read 2026-09-26)The vendor measuring its own product; lines, not work
Anthropic, internal surveySelf-reported share of work and productivity gainClaude used in 60% of work, a “50% productivity boost”Anthropic, How AI is transforming work at Anthropic, 2025-12-02132 engineers and researchers surveyed; self-report, not telemetry
Anthropic teams’ workflowsConcrete workflows, with a few timingsA 20-minute saving during an outage; up to 100 ad variations per batchAnthropic, How Anthropic teams use Claude Code, 2025-07-24Anecdotes; useful for the workflow, not for a forecast
Google, share of new codeShare of new code that is AI-generated75%Sundar Pichai at Google Cloud Next, as reported by Fast Company, 2026-04-24Secondary: no Google-owned transcript was reached; cite it as press-reported

The industry-wide figure to set beside these is DX’s: across data from over 400 companies in Q2 2026, “on average, 51.9% of code is now AI-authored” (DX, 2026-06-17). That is a share of code, self-reported, and says nothing about whether the code was worth writing.

How did Stripe get agents to ship over 1,300 pull requests a week?

Section titled “How did Stripe get agents to ship over 1,300 pull requests a week?”

Stripe’s Minions posts give a detailed public account of agents shipping production code at scale, and the detail is all about verification. Alistair Gray’s two posts (2026-02-09 and 2026-02-19) describe four ingredients:

  1. An oracle that existed before the agents. The loop runs against “Stripe’s enormous preexisting battery of tests—over three million of them”. The agents inherited a verification system; they did not bring one.
  2. A hard stop on retries. A minion gets “at most two rounds of CI”. If tests fail after the first push, it fixes them and pushes once more, and then “we send the branch back to its human operator for manual scrutiny.”
  3. Deterministic structure around the agent. Blueprints are “workflows defined in code that direct a minion run”, a state machine that mixes deterministic code nodes with agent nodes.
  4. Tools wired once for everyone. Toolshed “currently contains nearly 500 MCP tools for internal systems and SaaS platforms” (Part 2). The agent itself is “a fork of Block’s coding agent goose”.

Every pull request is still human-reviewed. What changed is what the reviewer receives: a branch that has already passed a test suite of millions of tests, not a raw diff. That is the pattern this site calls evidence, not diffs.

What did Microsoft’s rollout study measure?

Section titled “What did Microsoft’s rollout study measure?”

The Microsoft paper is the largest adoption study in the table: “tens of thousands of engineers at Microsoft over its early-2026 rollout” of Claude Code and GitHub Copilot CLI. Adopters “merged roughly 24% more pull requests than they would have otherwise”, and “the lift persists across our four-month window”.

Two findings matter more for a rollout plan than the headline. First, “first use spread primarily through social networks”: colleagues, not mandates, drove adoption, which is why team adoption starts with champions. Second, the authors measured merged pull requests as an output proxy and said so. Your plan should inherit that honesty: count accepted changes, and price them with cost per accepted change.

Anthropic publishes two kinds of evidence about itself, and they answer different questions.

The workflow stories (2025-07-24) show what people typed. When Kubernetes clusters stopped scheduling pods, the Data Infrastructure team “fed it dashboard screenshots”, Claude Code guided them through Google Cloud’s UI to pod IP address exhaustion and gave “the exact commands to create a new IP pool”, “saving them 20 minutes of valuable time during a system outage.” Growth Marketing built a workflow that “processes CSV files with hundreds of ads, identifies underperformers, and generates new variations”, plus a Figma plugin that generates “up to 100 ad variations” per batch. Data scientists without TypeScript fluency built “entire React applications for visualizing RL model performance”.

The aggregate numbers show scale and carry their own warnings. The >80% merged-lines figure is Anthropic’s conservative measure; the page itself notes that leadership’s “90% or more” estimate includes scripts and experimental code. Before that figure existed, Redwood Research argued that “the productivity boost at a given fraction of code generated isn’t that high because AI allows people to cheaply generate lots of very low value code” (Ryan Greenblatt, 2025-10-22). Quote the share of code with that critique beside it, or not at all.

A plan built only on the success stories fails the first informed question. Put these beside them:

  • METR’s randomised trial of 16 experienced open-source developers on 246 issues found that “the use of AI causes tasks to take 19% longer”, with a confidence interval of +2% to +39% (METR, 2025-07-10). The February 2026 follow-up leans toward AI but is not statistically significant, and METR calls its new data “an unreliable signal” (METR, 2026-02-24). Its point estimates are about 18% less time for returning developers (interval -38% to +9%) and 4% less for new recruits (-15% to +9%); both intervals cross zero. (METR’s post labels these figures “speedup” of -18% and -4%, but its analysis code reports the change in completion time, so a negative value means less time.)
  • Faros AI’s telemetry of 22,000 developers (April 2026) shows throughput and quality moving together: epics per developer +66.2%, but bugs per developer +54% and incidents per pull request +242.7% (Faros AI).
  • DORA’s 2025 report found “a positive relationship between AI adoption on both software delivery throughput and product performance”, and in the same breath “a negative relationship with software delivery stability” (Google Cloud, 2025-09-23).
  • DX found median pull request size “growing from 44 lines to 72 lines” between July 2025 and June 2026 (DX, 2026-06-17), which is more for every reviewer to read.

The success stories and the counter-evidence agree on one thing. DORA puts it plainly: “Without robust control systems, like strong automated testing, mature version control practices, and fast feedback loops, an increase in change volume leads to instability.” Stripe had the control system first. The state of agentic engineering in 2026 covers the full evidence base.

What do the rollouts that worked have in common?

Section titled “What do the rollouts that worked have in common?”
PatternWhere the evidence shows itWhat to copy
The oracle came before the autonomyStripe’s three million tests; DORA’s “strong automated testing”Measure how strong your tests are before widening what agents may merge
The loop has a hard stopStripe’s “at most two rounds of CI”Cap retries and spend; hand back to a human with a report
A human signs off at merge, on evidenceEvery Minions pull request is human-reviewedShip an evidence bundle with each agent pull request
Context and tools are wired onceStripe’s Toolshed of nearly 500 MCP toolsStand up shared MCP servers and context files for the whole team
The claim names its metric and its caveatMicrosoft’s merged-PR proxy, stated as a proxyReport accepted changes and their cost, never lines or ”% AI-written”

Vet a case study before it reaches your plan

Section titled “Vet a case study before it reaches your plan”

Run every vendor story, conference talk or analyst quote through these seven questions. A story that fails three or more is marketing; keep it out of the business case.

  1. Who published it, and when? No date or no named author means no citation.
  2. What exactly was measured? Lines, pull requests, self-reported hours and delivered outcomes are four different things.
  3. Against what baseline? “24% more than they would have otherwise” has a counterfactual. “3x faster” usually does not.
  4. Over how many people and how long? One engineer’s month is an anecdote; tens of thousands over four months is a study.
  5. What verified the output? If the story never mentions tests, CI or review, the quality side was not measured.
  6. What did quality do? Look for change-failure rate, incidents, reverts or bugs. Their absence is a finding.
  7. Does the source state its own caveat? The credible ones do: Microsoft on merged PRs, METR on its 2026 data, Anthropic on its 90% estimate.

Reproduce the bounded-loop pattern on one task type

Section titled “Reproduce the bounded-loop pattern on one task type”

Copy the pattern, not the number. Pick one repetitive, well-tested task type from your backlog (a dependency bump, a flaky-test fix, a deprecated-API migration) and run the Stripe shape on it: a written task, an agent run with a hard stop, the test suite as the oracle, and a human review of the evidence. The run takes an afternoon; the measurement design belongs to a pilot that proves something.

The task file is the same in every tool. Save it as .agent/tasks/deprecated-api.md:

Replace every call to legacyFetch() in src/ with httpClient.get(), keeping
behaviour identical. Acceptance criteria:
1. `npm test` passes.
2. No file outside src/ and test/ changes.
3. No existing test is deleted or weakened; add a test for any call site
that had none.
Process: run `npm test` after your change. If it fails, fix and run it
once more. If it still fails after the second run, stop and write
REPORT.md with the failing tests and your best diagnosis instead of
continuing. Always finish by writing REPORT.md: files changed, tests
added, test results, and anything you were unsure about.

Run it headless from the repository root in a terminal. --max-budget-usd caps spend (it works only with --print), and --permission-prompts none denies anything that would prompt, so --allowedTools lets through only the test command.

Terminal window
claude -p "$(cat .agent/tasks/deprecated-api.md)" \
--permission-mode acceptEdits \
--permission-prompts none \
--allowedTools "Bash(npm test *)" \
--max-budget-usd 5 \
--output-format json > run.json

To run the same task on every pull request, see scripting Claude Code.

How do you verify the reproduced pattern without reading every line?

Section titled “How do you verify the reproduced pattern without reading every line?”

The reviewer checks evidence, not every line of the diff. For each run:

  • Tests are the oracle. The run counts only if npm test passes and the reviewer prompt finds no weakened test. If the suite is thin, the pattern is not ready for this task type; strengthen the tests first.

  • Scope is checked mechanically. Run this in the terminal after the run. It lists every changed or new file outside src/ and test/ (the task file under .agent/ and the run’s own output files excepted), and any output means the run fails:

    Terminal window
    { git diff --name-only HEAD; git ls-files --others --exclude-standard; } \
    | grep -vE '^(src|test)/|^\.agent/|^(REPORT\.md|run\.json|last-message\.md)$' \
    && echo 'SCOPE VIOLATION' || echo 'scope ok'
  • The stop rule is counted. A run that hit its second failure produces REPORT.md and no pull request. Count these; they are your failure rate. Be clear about what enforces it: the two-run cap lives only in the prompt, so an agent can ignore it. For a hard limit, wrap the run, for example timeout 20m claude -p …, or run it in a CI job that counts test runs and kills the job at the third. Of the commands on this page, --max-budget-usd in the Claude Code tab is the only hard spend cap; the Codex and Cursor runs have none unless you add one.

  • A named person signs off. The tech lead who owns the repository merges or rejects, and records the verdict.

  • Quality is tracked after merge. Log reverts and incidents traced to agent changes for 30 days, then compute cost per accepted change.

After 10 to 20 runs of one task type, you have an internal case study with a baseline, a sample, a verification method and a quality number, which is more than most vendor decks carry. Write it up with this skeleton and file it beside the business case:

Internal case study: <task type>, <team>, <dates>
Question: can agents complete <task type> to our merge standard?
Baseline: median hours per task before the pilot (n = ?, source)
Runs: n = ?, merged = ?, stopped by the retry cap = ?, rejected at review = ?
Verification: test suite (size, mutation score if known), scope check, reviewer
Quality after 30 days: reverts = ?, incidents = ?
Cost: agent spend + review hours, per accepted change
Caveats: what this does not show (other task types, other teams)
Decision: expand / repeat / stop, and who decided
  • You borrow a number without the oracle behind it. Stripe’s volume rests on three million tests. Recovery: before forecasting, measure your own suite’s strength on the task type and plan autonomy per loop, as the adoption roadmap does.
  • You count output instead of accepted change. Lines and pull requests rise while bugs and incidents rise faster, as Faros measured. Recovery: switch the headline metric to accepted, non-reverted changes and report quality beside throughput, using the metrics frameworks page.
  • Self-reported gains stand in for measurement. METR’s 2025 developers believed they were 20% faster while measuring slower. Recovery: take a baseline from delivery data before the pilot starts, not from a survey afterwards.
  • You copy the tool, not the harness. Buying seats reproduces none of Stripe’s blueprints, retry cap or tooling. Recovery: fund the shared harness (tests, context files, MCP servers, review gates) as its own line in the business case.
  • Survivorship hides the failures. Companies publish what worked. Recovery: keep your stopped and rejected runs in the internal case study; they are the evidence that the stop rule works.
  • Review becomes the bottleneck. More agent pull requests mean more to review, and DX measured them growing larger. Recovery: set pull request size budgets and route by risk, as in the review queue.

Design a pilot that proves something

Baselines, cohorts and a decision rule written before the pilot starts.

Design the pilot →

Frequently asked questions

Which AI coding case studies are publicly documented and verifiable?

Stripe's Minions (over 1,300 agent-written pull requests merged each week, February 2026), Microsoft's study of its early-2026 Claude Code and GitHub Copilot CLI rollout (roughly 24% more merged pull requests, July 2026), and Anthropic's published accounts of its own teams.

What do the successful AI coding rollouts have in common?

Each scaled on verification: an existing test suite as the oracle, a hard limit on how often the agent may retry, and a human who reviews before merge. None of them removed the check; they made the check automatic.

Can I use a vendor's case-study number in my business case?

Only as a dated, attributed data point about that company. Its number depended on its own test suite and review process, so measure your own baseline in a pilot before forecasting anything.