Executive track: the ten-minute brief and the five decisions you own
The executive track is the AI Developer Toolkit reading path for business leaders who fund software delivery: a ten-minute brief on coding agents in 2026, the dated evidence both ways, and the five decisions leadership owns.
Your CTO wants a coding agent for every engineer, a vendor deck promises a multiple, a board member forwards a study saying developers got slower, and finance asks about headcount. You decide without reading code or taking anyone’s word.
What do coding agents do in 2026?
Section titled “What do coding agents do in 2026?”Coding agents such as Claude Code, Codex, and Cursor take a task, change files across a codebase, run the tests, and open a change for a person to approve; the decisions below hold for any tool, and the tool comparison covers the choice. Agents already write much of the code (Anthropic: “more than 80% of the code we merge”, May 2026, vendor-internal; DX: 51.9% across 400+ companies, self-reported, 17 June 2026), but share of code measures volume, not value.
What does the evidence say, both ways?
Section titled “What does the evidence say, both ways?”| Finding | Publisher, date |
|---|---|
| Microsoft adopters of Claude Code and GitHub Copilot CLI merged about 24% more pull requests; the authors note a merged PR is not delivered value | Microsoft researchers, arXiv, 1 July 2026 |
| Task throughput per developer +33.7%, but incidents per pull request +242.7% and median time in review +441.5% | Faros AI telemetry, 22,000 developers, April 2026 |
| AI adoption relates positively to throughput and negatively to delivery stability | DORA 2025 report, 23 September 2025 |
| Experienced developers (16, early-2025 tools) took 19% longer with AI; the effect is unproven, and METR calls its 2026 follow-up data “an unreliable signal” | METR, 10 July 2025 and 24 February 2026 |
DORA names the mechanism: without strong automated testing and fast feedback loops, “an increase in change volume leads to instability.” Sources: the state of agentic engineering.
How mature is your engineering organization?
Section titled “How mature is your engineering organization?”Read Dan Shapiro’s autonomy scale (23 January 2026, L0 to L5) as a maturity model. The free C-Level Scorecard gives your organization a level from 1 to 4 that maps onto the ladder below.
| Level | What proves it works | Where most teams are | What moving up costs you |
|---|---|---|---|
| L0–L2: autocomplete to junior developer | People reading every line | L2 is “where 90% of ‘AI-native’ developers are living right now” | Licenses and training |
| L3: the developer, agents write most code | A reviewer reading every change | “almost everyone tops out here” | Review capacity |
| L4: the engineering team | Automated tests and stop conditions | Shapiro places himself here | Tests and evidence on every change |
| L5: the dark software factory | Nobody reads the code | “a handful of people” | Full evidence pipeline; no human reads code |
Moving from L3 to L4 is an investment in tests and evidence, not seats; see the autonomy ladder and what stays human.
The executive track, step by step
Section titled “The executive track, step by step”-
Why now: What changed, according to whom, and when: the dated evidence first.
Done when: You can separate sourced claims from vendor marketing.
-
Why now: One picture of how far along your engineering organization is.
Done when: You can ask your CTO where each team sits on the map.
-
Why now: Why quality can be proven without people reading every line of code.
Done when: You know what evidence to ask for before approving more autonomy.
-
Why now: The economics of agent-built software: cost per accepted change.
Done when: A cost baseline your finance team agrees with.
-
Why now: Build the case with ranges and sources, not multipliers.
Done when: A business case with stated assumptions and a stop condition.
-
Why now: How teams, roles and headcount change shape.
Done when: An org-design hypothesis you can test in one division.
-
Why now: Set the pace: from pilots to an agentic delivery system, with gates.
Done when: A roadmap with a go or no-go gate per phase.
-
Why now: Report progress, risk and spend so the board can steer.
Done when: The first quarterly board update delivered.
Further reading
Which five decisions does leadership own?
Section titled “Which five decisions does leadership own?”| Decision | Ask your CTO for | Decided when |
|---|---|---|
| Economics: what does an accepted change cost? | Seats, tokens, and review hours per change merged and not reverted | Finance agrees with the cost baseline |
| Business case: is it worth funding at scale? | Ranges with sources, stated assumptions, and a stop condition | A funded case with a kill criterion |
| Org design: how do teams, roles, and hiring change? | Where review time goes; the plan for junior hiring | A hypothesis under test in one division |
| Roadmap: how fast do you move? | Phases with a go or no-go gate on stability metrics | A roadmap with signed-off gates |
| Board reporting: what does the board see? | Output always paired with stability, spend, and risk | The first quarterly update delivered |
How do you know the program works without reading code?
Section titled “How do you know the program works without reading code?”Use this set quarterly: the CTO owns the definitions, finance signs the cost line, and a named engineering owner approves each autonomy increase.
| Metric | Definition | Rule |
|---|---|---|
| Output | Changes merged per engineer per week | Never reported alone |
| Stability | Change failure rate and incidents per merged change | Worse two months running: pause autonomy increases |
| Review load | Median hours a change waits for and spends in review | Rising: fund verification before planning headcount |
| Unit cost | Agent spend plus review hours per accepted change | Compared with the Economics baseline |
| Evidence | % of merged changes with passing checks and a named approver | 100% before any team moves from L3 to L4 |
Where executive AI programs go wrong
Section titled “Where executive AI programs go wrong”- Share of code becomes the target. Report unit cost and stability.
- One study is read as the verdict. Run your own pilot with a baseline and a stop condition.
- Autonomy is granted by memo. Tie each increase to one team’s evidence (autonomy governance).
Beyond this track
Section titled “Beyond this track”Frequently asked questions
What is the executive track?
A reading path for business leaders who fund software delivery. It opens with the free C-Level Scorecard, gives a ten-minute brief with a publisher and a date on every number, and then takes one step per decision leadership owns: economics, business case, org design, roadmap, and board reporting.
Do coding agents make engineering teams faster?
Output rises and stability falls unless verification keeps up. Microsoft researchers measured roughly 24% more merged pull requests among adopters (July 2026); Faros AI measured task throughput per developer up 33.7% alongside incidents per pull request up 242.7% (April 2026); DORA's 2025 report found the same split between throughput and stability.
What should an executive ask the CTO about AI coding agents?
Five questions: where each team sits on the autonomy ladder and what evidence places it there, what an accepted change costs, what happened to incidents and review time when throughput rose, what an agent may do without human approval, and which number would make you pause.