Skip to content

The organization-wide roadmap: from pilots to an agentic delivery system

The organization-wide transformation roadmap is a 12-month, five-stage plan that takes an engineering organization from scattered AI pilots to an agentic delivery system. It funds a platform team first, pilots named delivery loops, builds verification before granting autonomy, and advances each loop one autonomy-ladder rung at a time, only when that loop’s gate evidence is met.

This page is for the executive who funds the change and the CTO or VP Engineering who runs it. The situation it solves: licenses went out to every engineer six months ago, usage dashboards look busy, and nobody can say whether delivery got faster, cheaper or riskier. Some teams let agents merge overnight, others ban them, and the board wants a plan with dates by the next quarter.

What you’ll walk away with from this transformation roadmap

Section titled “What you’ll walk away with from this transformation roadmap”
  • A five-stage, 12-month plan with an owner and an exit gate for every stage
  • A stop/go gate table written as ladder transitions per loop (L1 → L2 through L4 → L5), with the evidence each one requires
  • A loop register template your platform team can commit today
  • The org-level metric: the distribution of merged changes across ladder rungs
  • Per-tool control-plane settings for Claude Code, Codex and Cursor
  • Two copy-paste prompts that compile gate evidence from your own pull requests

Why a seat rollout is not a transformation

Section titled “Why a seat rollout is not a transformation”

Buying seats changes who types. It does not change how a change is specified, verified or released, so the delivery system stays the one you had. The 2025 DORA report (Google Cloud, 23 September 2025) states the risk directly: “AI doesn’t fix a team; it amplifies what’s already there. Strong teams use AI to become even better and more efficient. Struggling teams will find that AI only highlights and intensifies their existing problems.”

The telemetry agrees. Faros AI’s April 2026 report, built on two years of data from 22,000 developers, measured throughput and quality moving in opposite directions: epics completed per developer rose 66.2% while incidents per pull request rose 242.7% and median time in review rose 441.5%. DX measured median pull request size growing from 44 to 72 lines between July 2025 and June 2026 (DX, 17 June 2026). Generation scaled; verification did not.

The individual-speed evidence is weaker than vendor decks suggest. METR’s 2025 randomized trial found experienced developers took 19% longer with AI; its 2026 follow-up produced point estimates in AI’s favour whose intervals cross zero, and METR itself calls the new data an unreliable signal. Microsoft’s study of its early-2026 rollout of Claude Code and GitHub Copilot CLI (arXiv, July 2026) measured roughly 24% more merged pull requests among adopters, with the authors’ own caveat that “a merged PR is not the same as the value it delivers.”

So the roadmap below spends its first two stages on the parts that do not scale by themselves: a shared harness, measurement and verification. Autonomy comes after, per loop, on evidence. The dated sources behind these numbers are collected on the state of agentic engineering.

What a loop is, and how to measure the organization’s level

Section titled “What a loop is, and how to measure the organization’s level”

A loop is one recurring kind of change in one codebase, with one owner: dependency updates in payments-api, UI bug fixes in the web app, new endpoints in the partner API. The autonomy ladder rates loops, not people or companies, because the same team can safely run dependency updates at L4 while its billing changes stay at L3.

That makes the organization’s level a distribution, not a single number. Report it as the share of merged changes by rung, across all registered loops:

RungWho reads whatShare of merged changes (example, month 0)Example, month 12
L1–L2A human reads every line as it is written70%25%
L3The agent runs unattended; a human reads every diff30%45%
L4A human reads the spec, the tests and the evidence, not the diff0%28%
L5Nobody reads the code; the oracle and release controls decide0%2%

The percentages are an illustration of the shape, not a target or a benchmark. Your targets come from your own loop register. The one map page defines the rungs, the lifecycle stages and the factory stations in one grid, and reading evidence instead of code explains what replaces the diff at L4.

StageMonthsGoalAccountableExit gate (go)
0. Mandate and baseline0–1Name the sponsor, the loops and the current numbersExecutive sponsor, CTOBaseline for every pilot loop; budget and stop rule signed
1. Platform and first loops1–3Stand up the shared harness; pilot three to five loops from L1–L2 to L3Platform leadPilot loops pass the L2 → L3 gate; managed policy on every seat
2. Verification before autonomy3–6Protected tests, evidence bundles, agent review; first loops to L4CTO, loop ownersAt least one low-risk loop passes the L3 → L4 gate for 30 days
3. Scale by loop6–9Register every team’s loops; roll the harness out; start trust transferPlatform lead, tech leadsMost active repositories have registered loops; no gate skipped
4. Operate the system9–12Quarterly gate reviews, board reporting, L5 only where the oracle allowsCTO, executive sponsorCost per accepted change and change-fail rate at or better than baseline

Each stage has a calendar window, but a stage ends on its exit gate, never on the date. A stage that misses its gate extends, and the board report says so.

  1. Stage 0 — mandate and baseline (month 0–1). The executive sponsor writes a one-paragraph AI stance: what agents may do, what stays human, and that the roadmap will stop loops that fail their gates. DORA lists a “clear and communicated AI stance” as the first capability in its AI Capabilities Model (Google Cloud, 23 September 2025). The CTO then picks three to five pilot loops, records their baseline with the method from designing a pilot that proves something, and puts the full cost, including review and rework, into the business case. Exit when every pilot loop has at least five comparable historical changes on record.

  2. Stage 1 — platform and first loops (months 1–3). Fund a small platform team, two or three engineers to start, whose product is the harness: managed settings, shared rules and skills, CI gates, agent identities and telemetry. DORA’s 2025 report found that 90% of organizations have adopted at least one platform and “a direct correlation between a high quality internal platform and an organization’s ability to unlock the value of AI”. The pilot loops move from L1–L2 to L3 on that harness, repository by repository, using the per-repository adoption roadmap. The agent platform team page has the charter; the operating model says who owns which control.

  3. Stage 2 — verification before autonomy (months 3–6). Before any loop leaves L3, its tests must be an oracle the agent cannot quietly rewrite. Build oracle strength, protect the oracle, require an evidence bundle on every pull request, and add agent PR review. Then, and only then, move one low-risk loop to L4. Stripe’s published account of its agents shows the pattern: they run against “over three million” pre-existing tests and get “at most two rounds of CI” before the branch returns to a human (Stripe, 9 and 19 February 2026).

  4. Stage 3 — scale by loop (months 6–9). Every team registers its loops in the same register and inherits the platform harness. Tech leads run the trust transfer protocol (shadow review, then sampled review, then evidence-only) for loops approaching L4. Governance moves from memo to mechanism: one managed policy across tools, agent identities and secrets, cost governance and, for EU operations, the EU AI Act obligations. Leading the change covers the people side.

  5. Stage 4 — operate the system (months 9–12). The CTO runs a quarterly gate review for every loop (advance, hold or roll back one rung) and reports the rung distribution, cost per accepted change and incident trend with the board reporting template. L5 is considered only for low-risk loops whose oracle, progressive delivery and automated rollback have held for a full quarter. The roadmap for year two is written from the register, not from vendor launches.

Stop/go gates as ladder transitions per loop

Section titled “Stop/go gates as ladder transitions per loop”

A gate is the evidence a loop must show before it climbs one rung. Loops climb one rung at a time; a loop never skips from L2 to L4. The loop owner assembles the evidence, the named approver signs, and a stop signal sends the loop back one rung automatically.

TransitionGo: entry evidence for this loopStop: roll back one rung whenSigns
L1 → L2 (paired)Managed policy applied to every seat; shared rules file committed; secrets out of the agent’s reach; engineers trained on the loopA secret or production credential appears in an agent session or transcriptPlatform lead
L2 → L3 (agent unattended, human reads every diff)The agent runs typecheck, lint and tests before opening a pull request; first-pass CI rate at or above baseline; time to first review within the team’s service levelMedian time in review or pull request size rises for four weeks; change-fail rate above baselineLoop owner (tech lead)
L3 → L4 (human reads spec and evidence, not the diff)Acceptance criteria are executable; test files protected from agent edits; evidence bundle on every pull request for 30 days; sampled human review of 20 changes finds no defect the evidence missed; rollback rehearsedAny escaped defect traced to the loop; any agent edit to protected tests; evidence bundle missing on a merged changeCTO or delegate, plus the service owner
L4 → L5 (no human reads the code)Risk class low; oracle owned by someone other than the agent’s operator; progressive delivery with automated rollback; a full quarter at L4 with change-fail rate at or below baselineAny high-severity incident traced to an unread changeCTO, with security sign-off

The thresholds in the table, such as 30 days and 20 sampled changes, are starting values to adapt, not research findings. Write your own into the register before the loop starts, because a threshold chosen after the results arrive proves nothing. Metrics frameworks holds the canonical definitions of change-fail rate, lead time and the AI-specific measures.

The roadmap lives as one file in a repository the platform team owns, reviewed like code. Each entry is also a portfolio item in the sense of the AI tooling roadmap: a hypothesis with an owner, a budget and a stop gate.

# loop-register.yaml — one entry per loop; reviewed quarterly
- id: payments-api/dependency-updates
owner: tech-lead-payments # accountable for the gate evidence
risk_class: low # low | medium | high | critical
current_rung: L3
target_rung: L4 # never more than one rung above current
target_by: 2027-03-31
entry_evidence:
- acceptance checks run in CI on every pull request
- test files protected; agent edits need the label oracle-change
- evidence bundle attached to every merged PR for 30 days
- 20 sampled human reviews with no defect the evidence missed
stop_signals:
- any escaped defect traced to this loop
- change-fail rate above baseline for two consecutive weeks
rollback_rung: L3
metrics: [cost_per_accepted_change, lead_time, change_fail_rate, time_in_review]
baseline: { lead_time_days: 3.1, change_fail_rate: 0.04 }
next_review: 2026-12-15

baseline holds your own measured values; the ones above are placeholders. Replace tech-lead-payments with a named person, not a team alias, so every gate has one accountable signature.

DecisionExecutive sponsorCTO / VP EngPlatform leadLoop ownerSecurityFinance
AI stance and 12-month budgetAccountableResponsibleConsultedInformedConsultedConsulted
Loop register and risk classesInformedAccountableResponsibleResponsibleConsulted—
L2 → L3 gate—InformedConsultedAccountable——
L3 → L4 and L4 → L5 gatesInformedAccountableConsultedResponsibleConsulted—
Stop a loop or the programAccountableResponsibleConsultedConsultedConsultedConsulted

Anyone in the table can trigger a stop signal. Only the accountable role can restart a stopped loop.

The platform team’s first deliverable in stage 1 is a policy that every seat receives and no user can override. The mechanism differs per tool; the gates above do not. Enforcing one policy across every coding agent compares all three in one table.

Claude Code reads a managed-settings.json file deployed to the system-level settings location, through MDM, or as server-managed settings from the Claude admin console on Team and Enterprise plans. Model governance keys include availableModels, enforceAvailableModels and maxEffortLevel. A minimal stage-1 policy restricts models and keeps secrets out of reach:

{
"availableModels": ["opus", "sonnet"],
"permissions": {
"deny": ["Read(./.env)", "Read(./.env.*)", "Read(./secrets/**)"]
}
}

For unattended loops at L3 and above, run the agent in CI with anthropics/claude-code-action@v1, or headless with claude -p and a spending cap such as --max-budget-usd 5. The analytics dashboard (Team and Enterprise) supplies usage, but not outcome, data; join it to your pull request metrics yourself. Checked against Claude Code v2.1.283 on 2026-09-26.

The quarterly gate review should read evidence the agent gathered from the repository, not a status slide. Both prompts work in Claude Code, Codex or Cursor’s agent with the GitHub CLI (gh) authenticated.

Speed claims do not count unless quality held. Check these four signals at every quarterly review, and treat a gain in the first two with a loss in the last two as a failure, not a trade-off:

  • Cost per accepted change: licenses, usage, CI, review time, rework and incident cost, divided by merged changes that were not reverted. The economics of agent-built software defines it and gives a worked example.
  • Lead time from accepted spec to production, per loop.
  • Change-fail rate and escaped defects, per loop, against the stage-0 baseline.
  • Time in review and pull request size, the two measures that rose most visibly in the Faros and DX data.

“Percentage of code written by AI”, seat counts and prompts per day are activity, not outcomes. They can go in an appendix; they never decide a gate. The verification itself, not a human reading every diff, is what the gates rely on: executable acceptance criteria, protected tests, the evidence bundle and agent review, with a named human signing each rung.

What derails an organization-wide rollout?

Section titled “What derails an organization-wide rollout?”

Seats go out before the harness exists. Symptom: usage rises, lead time is flat, and every team configures the agent differently. Recovery: freeze new seat grants, stand up the platform team, apply the managed policy, and register loops before granting more access.

The pilot proves nothing. Symptom: the pilot used typo fixes and volunteers, and the result is a satisfaction survey. Recovery: rerun it on production-relevant loops with a baseline and a decision rule written in advance, as pilot design describes.

Autonomy arrives before verification. Symptom: throughput up, incidents per change and review time up, the Faros pattern. Recovery: roll the affected loops back one rung, and do not re-apply the L3 → L4 gate until the tests are protected and the evidence bundle is mandatory.

The platform team becomes the bottleneck. Symptom: every new loop waits weeks for a platform review. Recovery: publish the gate evidence as a self-service checklist; the platform team owns the harness, not every loop’s approval.

Gates pass on the calendar. Symptom: the stage 2 date arrives and loops advance because “the plan says month six”. Recovery: the board report shows stage status by gate, and a stage that misses its gate is reported as extended, not complete.

Review becomes the new queue. Symptom: agents open more and larger pull requests than reviewers can read. Recovery: cap pull request size per loop, route by risk class, and move qualifying loops through trust transfer rather than adding reviewers.

Unmanaged agents run outside the register. Symptom: a team runs an overnight agent with a personal token. Recovery: revoke the token, give the loop a service identity, and register it at L3 until it earns its gate.

Where to go next from the transformation roadmap

Section titled “Where to go next from the transformation roadmap”
Edit page

Last updated:

Cite this page — https://developertoolkit.ai/en/strategy/transformation-roadmap/, developertoolkit.ai