Skip to content

Executive track: the ten-minute brief and the five decisions you own

The executive track is the AI Developer Toolkit reading path for business leaders who fund software delivery: a ten-minute brief on coding agents in 2026, the dated evidence both ways, and the five decisions leadership owns.

Your CTO wants a coding agent for every engineer, a vendor deck promises a multiple, a board member forwards a study saying developers got slower, and finance asks about headcount. You decide without reading code or taking anyone’s word.

Coding agents such as Claude Code, Codex, and Cursor take a task, change files across a codebase, run the tests, and open a change for a person to approve; the decisions below hold for any tool, and the tool comparison covers the choice. Agents already write much of the code (Anthropic: “more than 80% of the code we merge”, May 2026, vendor-internal; DX: 51.9% across 400+ companies, self-reported, 17 June 2026), but share of code measures volume, not value.

FindingPublisher, date
Microsoft adopters of Claude Code and GitHub Copilot CLI merged about 24% more pull requests; the authors note a merged PR is not delivered valueMicrosoft researchers, arXiv, 1 July 2026
Task throughput per developer +33.7%, but incidents per pull request +242.7% and median time in review +441.5%Faros AI telemetry, 22,000 developers, April 2026
AI adoption relates positively to throughput and negatively to delivery stabilityDORA 2025 report, 23 September 2025
Experienced developers (16, early-2025 tools) took 19% longer with AI; the effect is unproven, and METR calls its 2026 follow-up data “an unreliable signal”METR, 10 July 2025 and 24 February 2026

DORA names the mechanism: without strong automated testing and fast feedback loops, “an increase in change volume leads to instability.” Sources: the state of agentic engineering.

How mature is your engineering organization?

Section titled “How mature is your engineering organization?”

Read Dan Shapiro’s autonomy scale (23 January 2026, L0 to L5) as a maturity model. The free C-Level Scorecard gives your organization a level from 1 to 4 that maps onto the ladder below.

LevelWhat proves it worksWhere most teams areWhat moving up costs you
L0–L2: autocomplete to junior developerPeople reading every lineL2 is “where 90% of ‘AI-native’ developers are living right now”Licenses and training
L3: the developer, agents write most codeA reviewer reading every change“almost everyone tops out here”Review capacity
L4: the engineering teamAutomated tests and stop conditionsShapiro places himself hereTests and evidence on every change
L5: the dark software factoryNobody reads the code“a handful of people”Full evidence pipeline; no human reads code

Moving from L3 to L4 is an investment in tests and evidence, not seats; see the autonomy ladder and what stays human.

  1. Why now: What changed, according to whom, and when: the dated evidence first.

    Done when: You can separate sourced claims from vendor marketing.

  2. Why now: One picture of how far along your engineering organization is.

    Done when: You can ask your CTO where each team sits on the map.

  3. Why now: Why quality can be proven without people reading every line of code.

    Done when: You know what evidence to ask for before approving more autonomy.

  4. Why now: The economics of agent-built software: cost per accepted change.

    Done when: A cost baseline your finance team agrees with.

  5. Why now: Build the case with ranges and sources, not multipliers.

    Done when: A business case with stated assumptions and a stop condition.

  6. Why now: How teams, roles and headcount change shape.

    Done when: An org-design hypothesis you can test in one division.

  7. Why now: Set the pace: from pilots to an agentic delivery system, with gates.

    Done when: A roadmap with a go or no-go gate per phase.

  8. Why now: Report progress, risk and spend so the board can steer.

    Done when: The first quarterly board update delivered.

DecisionAsk your CTO forDecided when
Economics: what does an accepted change cost?Seats, tokens, and review hours per change merged and not revertedFinance agrees with the cost baseline
Business case: is it worth funding at scale?Ranges with sources, stated assumptions, and a stop conditionA funded case with a kill criterion
Org design: how do teams, roles, and hiring change?Where review time goes; the plan for junior hiringA hypothesis under test in one division
Roadmap: how fast do you move?Phases with a go or no-go gate on stability metricsA roadmap with signed-off gates
Board reporting: what does the board see?Output always paired with stability, spend, and riskThe first quarterly update delivered

How do you know the program works without reading code?

Section titled “How do you know the program works without reading code?”

Use this set quarterly: the CTO owns the definitions, finance signs the cost line, and a named engineering owner approves each autonomy increase.

MetricDefinitionRule
OutputChanges merged per engineer per weekNever reported alone
StabilityChange failure rate and incidents per merged changeWorse two months running: pause autonomy increases
Review loadMedian hours a change waits for and spends in reviewRising: fund verification before planning headcount
Unit costAgent spend plus review hours per accepted changeCompared with the Economics baseline
Evidence% of merged changes with passing checks and a named approver100% before any team moves from L3 to L4
  • Share of code becomes the target. Report unit cost and stability.
  • One study is read as the verdict. Run your own pilot with a baseline and a stop condition.
  • Autonomy is granted by memo. Tie each increase to one team’s evidence (autonomy governance).

Frequently asked questions

What is the executive track?

A reading path for business leaders who fund software delivery. It opens with the free C-Level Scorecard, gives a ten-minute brief with a publisher and a date on every number, and then takes one step per decision leadership owns: economics, business case, org design, roadmap, and board reporting.

Do coding agents make engineering teams faster?

Output rises and stability falls unless verification keeps up. Microsoft researchers measured roughly 24% more merged pull requests among adopters (July 2026); Faros AI measured task throughput per developer up 33.7% alongside incidents per pull request up 242.7% (April 2026); DORA's 2025 report found the same split between throughput and stability.

What should an executive ask the CTO about AI coding agents?

Five questions: where each team sits on the autonomy ladder and what evidence places it there, what an accepted change costs, what happened to incidents and review time when throughput rose, what an agent may do without human approval, and which number would make you pause.