Skip to content

The State of Agentic Engineering, September 2026

Agentic engineering is the discipline of directing fallible coding agents without lowering the quality bar. As of September 2026, DX measures 51.9% of code as AI-authored across 400+ companies and Sonar’s developers report 42%. Anthropic reports over 80% of its merged lines, and Faros and DORA point to verification capacity, not generation, as the limit.

This report is for CTOs and executives setting strategy, tech leads deciding how far to trust agents, and developers placing their own workflow. A CTO asks what share of the codebase agents write. The lead says 40%. The developer running the agents knows half of those lines are tests the agent wrote for its own change. Three people, three rungs of one ladder, and nobody can name the rung.

What you’ll walk away with from the 2026 evidence

Section titled “What you’ll walk away with from the 2026 evidence”
  • An evidence table where every number has a publisher, a date and a unit
  • Shapiro’s six-rung autonomy ladder and a five-question test that places a repository on it
  • The six stations of a software factory, with the matching feature in Claude Code, Codex and Cursor
  • METR, Faros, DORA and SlopCodeBench, read correctly
  • Three copy-paste prompts and a metric-pair template for leadership

The share rises everywhere it is measured.

MeasureThenNowSource
Anthropic merged lines authored by Claude“low single digits”, Feb 2025“more than 80%”Anthropic, “as of May 2026”
AI-authored share, 400+ companies27.4%, Q1 202651.9% averageDX, 17 Jun 2026
Committed code reported as AI-generated or assistednot measured42%Sonar, 8 Jan 2026
Stripe PRs per week with no human-written code“over a thousand”, 9 Feb 2026“over 1,300”Stripe, 19 Feb 2026
Code merged per Anthropic engineer per day2024 baseline8×Anthropic, Q2 2026
Merged PRs of CLI-agent adopters at Microsoftsame engineers without the tools (modelled)“roughly 24% more”Murphy-Hill et al., arXiv, 1 Jul 2026

At the individual end the share reaches 100% and stays there. Boris Cherny, who leads Claude Code, told Fortune on 29 January 2026: “For me personally, it has been 100% for two+ months now, I don’t even make small edits by hand.”

Dario Amodei told the Council on Foreign Relations on 10 March 2025 that AI would write “essentially all of the code” within twelve months. Eighteen months on, Anthropic reports “more than 80%” of its own merged lines, and the industry’s best series is DX’s 51.9%. Fast Company reports Sundar Pichai putting Google’s new code at 75% in April 2026, a secondary report that no Google primary source confirms.

What do the percentage-of-code numbers actually measure?

Section titled “What do the percentage-of-code numbers actually measure?”

Every one of those percentages counts lines, and lines are the cheapest thing an agent makes. Ryan Greenblatt of Redwood Research wrote on 22 October 2025: “The productivity boost at a given fraction of code generated isn’t that high because AI allows people to cheaply generate lots of very low value code.”

Anthropic draws the same line. Its leadership has estimated 90% or more “including scripts and experimental code”, while the published figure measures “the share of lines merged to production that can be attributed to Claude.” DX adds the cost: median pull request size grew “from 44 lines to 72 lines per pull request between July 2025 and June 2026”. A rising AI share and a near-doubled diff are one event, and the reviewer pays for it.

Andrej Karpathy named the discipline in his Sequoia Ascent summary of 30 April 2026:

“I call it agentic engineering because it is an engineering discipline. You have agents, which are spiky entities. They are fallible and stochastic, but extremely powerful. How do you coordinate them to go faster without sacrificing your quality bar?

Vibe coding raises the floor. Agentic engineering is about extrapolating the ceiling.”

The constraint is explicit: “You are still responsible for your software, just as before.” What changed is the unit of work, “from typing lines of code to delegating larger ‘macro actions’.” The definition at the top of this page condenses that passage; it is not a quotation.

Which level of the autonomy ladder is your team at?

Section titled “Which level of the autonomy ladder is your team at?”

Dan Shapiro published the model on 23 January 2026 as “The Five Levels: from Spicy Autocomplete to the Dark Factory”, after the NHTSA driving-automation scale. It is zero-indexed, so six rungs sit under a title that says five. Simon Willison’s write-up of 28 January identifies the Level 5 team as StrongDM’s AI division.

LevelShapiro’s nameWritesReadsWhat it feels likeGuide
L0Spicy autocompleteHumanHuman“not a character hits the disk without your approval”L1–2
L1The coding internHumanHuman“you offload specific, discrete tasks to your AI intern”L1–2
L2The junior developerBothHuman“feels like you are done. But you are not done”L1–2
L3The developerAIReviewer“Your life is diffs.”L3
L4The engineering teamAITests“leave for 12 hours, and check to see if the tests pass”L4
L5The dark software factoryAINobody“It’s a black box that turns specs into software.”L5

Shapiro calls Level 2 “where 90% of ‘AI-native’ developers are living right now”, says “almost everyone tops out here” of Level 3, places himself at Level 4, and puts “a handful of people” at Level 5, where “humans are neither needed nor welcome.”

Answer each question with evidence, not impressions. The first one you cannot answer marks the ceiling.

  1. L0 to L1. Does code reach disk that nobody typed?
  2. L1 to L2. Do you hand over whole tasks with acceptance criteria, or only completions inside a function you are writing?
  3. L2 to L3. Does the agent run unattended long enough that you meet its work as a diff? That is a job change, not a speed change.
  4. L3 to L4. Is there a stop condition a machine can evaluate, written as commands that exit 0? Level 4 is defined by leaving.
  5. L4 to L5. Does anything merge that no human read, and can you name the oracle that made it safe?

Assign a level per loop, not per company. A dependency-upgrade loop with a deterministic validator can run at Level 4 while feature work in the same repository stays at Level 2. The one-map page lays the ladder over the delivery lifecycle.

What does a production software factory look like?

Section titled “What does a production software factory look like?”

Stripe’s is the most detailed public account. Its Minions are one-shot, end-to-end agents built on “a fork of Block’s coding agent goose”, drawing on a Toolshed of “nearly 500 MCP tools”. By 19 February 2026, “over 1,300 Stripe pull requests… merged each week are completely minion-produced, human-reviewed, but containing no human-written code.” Three design decisions are worth copying:

  • The loop has a hard bound. Stripe allows “at most two rounds of CI”, then “we send the branch back to its human operator for manual scrutiny.”
  • The workflow is code. Blueprints are “a state machine that intermixes deterministic code nodes and free-flowing agent nodes.”
  • The oracle predates the agents. The minions run against “Stripe’s enormous preexisting battery of tests — over three million of them.”

Microsoft’s study of its early-2026 rollout of Claude Code and GitHub Copilot CLI covers “tens of thousands of engineers”, and its authors qualify the 24% themselves: “a merged PR is not the same as the value it delivers.”

Factories cost real money: an hour of Steve Yegge’s Gas Town orchestrator cost DoltHub’s Tim Sehn “about $100 in Claude tokens” (15 January 2026), and a week cost $3,000. BCG Platinion wrote in March 2026 that “the defining shift is not the absence of humans; it is the relocation of human effort.”

Alexey Grigorev’s July 2026 taxonomy separates context engineering (what the agent knows before it starts), loop engineering (when it stops) and graph engineering (who does what once there is more than one agent). Add what may run, what proves the result and what ships, and a factory has six stations, each with a documented feature in every tool.

StationDecidesClaude CodeCodexCursor
IntentWhat it knows firstCLAUDE.mdAGENTS.mdRules
HarnessWhat runs per editHooksHooksHooks
LoopWhen it stops/goal/goal/goal
GraphWho does whatSubagentsSubagentsSubagents
VerificationWhat proves it/code-review/reviewBugbot
ReleaseHow it shipsGitHub Actionsopenai/codex-action@v1GitHub Actions

Claude Code cells were re-checked on 26 September 2026. Codex features were confirmed the same day with codex-cli 0.157.1 (codex features list shows goals, hooks and multi_agent stable and on); Codex documentation links and Cursor cells were last verified on 28 August 2026.

In Claude Code, spawning a subagent fails once 20 are running in a session (v2.1.217+; CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS changes the limit, and ultracode sessions are exempt). Cursor began rolling out /goal on 19 August 2026 (per its changelog, read 28 August 2026), the same day subagents gained their own virtual machines. /loop is built into Claude Code and has been a bundled skill in Cursor since 3.5 (20 May 2026); Codex 0.157.1 has no /loop (checked 26 September 2026) and covers scheduled cadence with Automations (the goal-directed loop is /goal).

What does the evidence against autonomy say?

Section titled “What does the evidence against autonomy say?”

Four bodies of evidence argue for caution; quote each one whole.

METR measured a slowdown, then lost its instrument. The 2025 randomised trial of 16 experienced open-source developers across 246 issues found tasks took 19% longer (interval +2% to +39%), while the developers believed afterwards they had been sped up by 20%. The February 2026 follow-up reports changes in completion time of -18% (interval -38% to +9%) for returning developers and -4% (-15% to +9%) for new recruits. Negative means less time, so both point estimates favour AI, but both intervals cross zero and METR says the new data “gives us an unreliable signal of the current productivity effect of AI tools.” Neither “18% slower” nor “18% faster” is a finding.

Faros measured the whiplash. Two years of telemetry from 22,000 developers and 4,000+ teams, published April 2026. Throughput rose: epics per developer +66.2%, tasks +33.7%, merge rate +16.2%, while deployments per week fell 11.7%. Quality fell in the same window: bugs per developer +54%, incidents per pull request +242.7%, churn +861%, pull requests merging unreviewed +31.3%. Median time in review rose 441.5% and time to first review 156.6%; those are two different metrics.

DORA named the mechanism. Its 2025 report found “a positive relationship between AI adoption on both software delivery throughput and product performance” and that “AI adoption does continue to have a negative relationship with software delivery stability.” The reason: “Without robust control systems, like strong automated testing, mature version control practices, and fast feedback loops, an increase in change volume leads to instability.”

SlopCodeBench measured decay over time. Across 36 problems, 196 checkpoints and 15 agents extending their own earlier solutions, “no agent fully solves any problem end-to-end, and the best agent passes 14.8% of checkpoints”.

How is agent output verified without reading every line?

Section titled “How is agent output verified without reading every line?”

None of that argues for staying at Level 2. It shows the climb: autonomy is earned per loop against a verification oracle, never declared for a codebase. Faros shows generation scaling while verification did not, which is the Level 3 ceiling in telemetry. Climbing means adding a proof, not removing a reader.

LayerWhat it provesWho signs off
Stop condition as commands that exit 0The loop’s stated goal holdsThe spec author, before the run starts
Deterministic validators (types, lint, tests, schema checks)Known classes of error are absentCI, on every push
Review agent (/code-review, /review, Bugbot)A second model found no blocking issueThe human assignee reads its findings, not the diff
Retry bound and escalationThe agent cannot loop forever or weaken a test to passThe human operator, after the bound (Stripe: two CI rounds)
Production signals (incidents per PR, rollback rate)The oracle is catching what mattersThe tech lead, weekly

Naresh B A puts the rule in one line: “A model can argue that its work is complete. A deterministic validator can prove that a required field is missing.” Linear’s product rule: “issues can only be assigned to humans, and only delegated to agents.” The full method is reading evidence instead of diffs.

Six responsibilities survive every rung below the dark factory: taste, architecture, product direction, stop-condition design, permissions and the decision not to automate. Karpathy says “you are in charge of taste, engineering, design, and whether the system makes sense.” Understanding is the one that erodes quietly. In Anthropic’s randomised trial of 52 mostly junior engineers learning a new library (29 January 2026), “the AI group averaged 50% on the quiz, compared to 67% in the hand-coding group.” Oversight needs the skill that over-delegation wears down. The human’s job covers each responsibility in depth.

Copy-paste prompts and templates for placing your team on the ladder

Section titled “Copy-paste prompts and templates for placing your team on the ladder”

Run these prompts against a repository you ship from; each works unchanged in Claude Code, Codex and Cursor.

For leadership, the adoptable artifact is a reporting rule: never show a share-of-code or throughput number without its quality pair.

Report thisAlways next toWhy the pair
AI-authored share of merged linesIncidents per pull requestAnthropic and Redwood: lines are not work
Pull requests merged per engineerMedian time in review and time to first reviewFaros: both review metrics moved with throughput
Median pull request sizeShare of PRs merged without reviewDX and Faros: bigger diffs, thinner review
Autonomy level per loopThe named oracle and its retry boundStripe: a two-round CI bound plus 3M+ tests is what lets humans review instead of rewrite

DORA’s AI Capabilities Model (23 September 2025) names seven foundations to fund first. Metrics frameworks turns the pairs into a dashboard.

When the 2026 agentic engineering numbers get misused

Section titled “When the 2026 agentic engineering numbers get misused”

Five common misreadings, each with its fix.

  • Reading “percentage of lines” as “percentage of work”. Anthropic’s footnote and Redwood’s critique both refuse that step. Pair every share figure with a work figure.
  • Quoting half of Faros. The +66.2% epics and the +242.7% incidents come from one dataset. Quote throughput and quality together, or neither.
  • Reading METR’s sign backwards. METR reports a change in completion time, so -18% means less time, not a slowdown. Quote the interval and METR’s “unreliable signal” caveat every time.
  • Citing a vendor’s internal share as an industry rate. Anthropic, Stripe, Google and Cherny all report on themselves. For an industry number, use DX’s 51.9% or Sonar’s 42%.
  • Publishing feature claims that have expired. Cursor had no /goal before it began rolling out on 19 August 2026. Re-check the vendor page, or run the CLI’s --help, on the day you publish, and date every negative claim.
  1. Shapiro, “The Five Levels: from Spicy Autocomplete to the Dark Factory”, 23 Jan 2026. https://www.danshapiro.com/blog/2026/01/the-five-levels-from-spicy-autocomplete-to-the-software-factory/
  2. Simon Willison, “The Five Levels: from Spicy Autocomplete to the Dark Factory”, 28 Jan 2026. https://simonwillison.net/2026/Jan/28/the-five-levels/
  3. Karpathy, “Sequoia Ascent 2026”, 30 Apr 2026. https://karpathy.bearblog.dev/sequoia-ascent-2026/
  4. Anthropic, “When AI builds itself”, undated (“as of May 2026”). https://www.anthropic.com/institute/recursive-self-improvement
  5. Fortune, “Top engineers at Anthropic, OpenAI say AI now writes 100% of their code”, 29 Jan 2026. https://fortune.com/2026/01/29/100-percent-of-code-at-anthropic-and-openai-is-now-ai-written-boris-cherny-roon/
  6. DX, “AI-authored code has nearly doubled”, 17 Jun 2026. https://newsletter.getdx.com/p/ai-authored-code-has-nearly-doubled
  7. Sonar, “State of Code Developer Survey”, 8 Jan 2026. https://www.sonarsource.com/blog/state-of-code-developer-survey-report-the-current-reality-of-ai-coding/
  8. Murphy-Hill, Butler and Savelieva, “Adoption and Impact of Command-Line AI Coding Agents”, arXiv:2607.01418, 1 Jul 2026. https://arxiv.org/abs/2607.01418
  9. Stripe, “Minions”, parts 1 and 2, 9 and 19 Feb 2026. https://stripe.dev/blog/minions-stripes-one-shot-end-to-end-coding-agents-part-2
  10. Ryan Greenblatt, Redwood Research, “Is 90% of code at Anthropic being written by AIs?”, 22 Oct 2025. https://blog.redwoodresearch.org/p/is-90-of-code-at-anthropic-being
  11. METR, early-2025 developer study, 10 Jul 2025, and uplift update, 24 Feb 2026, with the analysis code that defines the sign. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/, https://metr.org/blog/2026-02-24-uplift-update/ and https://github.com/METR/Measuring-Late-2025-AI-on-OSS-Devs
  12. Faros AI, “The Acceleration Whiplash”, Apr 2026. https://www.faros.ai/blog/ai-acceleration-whiplash-takeaways
  13. Google Cloud, “Announcing the 2025 DORA Report”, 23 Sep 2025, and “Introducing DORA’s inaugural AI Capabilities Model”, 23 Sep 2025. https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report and https://cloud.google.com/blog/products/ai-machine-learning/introducing-doras-inaugural-ai-capabilities-model
  14. Orlanski et al., “SlopCodeBench”, 25 Mar 2026. https://arxiv.org/abs/2603.24755
  15. Anthropic research (Judy Hanwen Shen and Alex Tamkin), randomised trial on AI assistance and coding skills, 29 Jan 2026. https://www.anthropic.com/research/AI-assistance-coding-skills
  16. BCG Platinion, “The Agentic Software Factory”, 26 Mar 2026. https://www.bcgplatinion.com/insights/the-agentic-software-factory
  17. Grigorev, “AI-Native Development”, 22 Jul 2026. https://aishippingblog.com/p/ai-native-development-specifications
  18. Naresh B A, “Graph Engineering for AI Coding Agents”, 30 Jul 2025. https://dev.to/naresh_007/graph-engineering-for-ai-coding-agents-beyond-prompt-loops-48h4
  19. Linear, “Our approach to building the Agent Interaction SDK”, 1 Aug 2025. https://linear.app/now/our-approach-to-building-the-agent-interaction-sdk
  20. DoltHub (Tim Sehn), “A Day in Gas Town”, 15 Jan 2026, and “A Week in Gas Town”, 24 Mar 2026. https://www.dolthub.com/blog/2026-01-15-a-day-in-gas-town/ and https://www.dolthub.com/blog/2026-03-24-a-week-in-gas-town/
  21. Council on Foreign Relations, “CEO Speaker Series With Dario Amodei”, 10 Mar 2025. https://www.cfr.org/event/ceo-speaker-series-dario-amodei-anthropic
  22. Fast Company (secondary), “Google CEO says 75% of the company’s code is AI-generated”, 24 Apr 2026. https://www.fastcompany.com/91531519/google-ceo-says-75-of-the-companys-code-is-ai-generated

Frequently asked questions

What percentage of code is written by AI in 2026?

The broadest industry measure is DX's Q2 2026 panel of over 400 companies: 51.9% of code AI-authored on average. Sonar's January 2026 survey of over 1,100 developers reports 42% of committed code as AI-generated or assisted. Anthropic's "more than 80%" is vendor-internal and counts lines merged to production, not share of the work.

What is agentic engineering?

Andrej Karpathy named the discipline on 30 April 2026: "I call it agentic engineering because it is an engineering discipline. You have agents, which are spiky entities. They are fallible and stochastic, but extremely powerful. How do you coordinate them to go faster without sacrificing your quality bar?"

What are the levels of AI coding autonomy?

Dan Shapiro's scale, published 23 January 2026, runs from L0 to L5: spicy autocomplete, the coding intern, the junior developer, the developer (your life is diffs), the engineering team (you write a spec and check whether the tests pass), and the dark software factory, where nobody reads the code.

Does AI make developers faster?

No measurement settles it. METR's 2025 randomised trial found tasks took 19% longer. Its February 2026 follow-up has point estimates in AI's favour (about 18% and 4% less time) but intervals that cross zero, and METR calls that data an unreliable signal. Microsoft measured roughly 24% more merged pull requests; Faros measured quality falling alongside throughput.