Skip to content

Adoption roadmap for one repository: ladder transitions per loop

The per-repository adoption roadmap moves one codebase’s delivery loops up the autonomy ladder one rung at a time. Each loop, such as dependency bumps, climbs from L2 to L3, L4 and occasionally L5 only when its gate evidence is visible in Git, CI and the pull requests. It is the repository slice of the organization-wide transformation roadmap.

This page is for the tech lead who owns one repository. Your organization’s plan says “pilot loops reach L3 this quarter”, eight engineers use agents in eight different ways, and the review queue is already longer than it was in spring. You need to know what to change in this repository, in which order, and what proves each step worked.

  • A loops.yaml register that names the repository’s loops, their current rung and the next gate.
  • A baseline per loop, computed from merged pull requests, so every later claim has a comparison.
  • The repository changes for each transition: instructions file, isolated checkouts, protected tests, the evidence bundle, auto-merge rules.
  • A go and stop rule for every transition, with who signs.
  • Three prompts: baseline the loops, run an L3 task, and check a loop’s gate before you promote it.

Why does a repository climb by loop and not by wave?

Section titled “Why does a repository climb by loop and not by wave?”

A repository is not at one level. Dependency bumps with contract tests can safely run at L4 while payment-provider work beside them stays at L2, and both answers are correct at once. The one map defines a loop as a repeatable class of change with its own trigger, owner and oracle, and it rates each loop by what a human reads before merge.

Earlier versions of this roadmap used fixed waves 0–5 for the whole repository. Waves failed in a predictable way: a wave ended on the calendar, and every kind of change advanced together, so the riskiest loop set the pace for the safest. If your team still uses the wave names, translate them with this table and then stop using them.

Retired waveWhat it addedWhere it lives now
0. BaselineOne repository, current metricsBaseline and loop register, before any transition
1. Artifact firstintent.md → spec.md → plan.md, Plan modeEntry evidence for L2 → L3
2. Review firstAgent review, protected tests, branch protectionL2 → L3 (branch protection) and L3 → L4 (protected tests, agent review)
3. Safe parallelismOne worktree per task, scoped credentialsL2 → L3, isolated checkouts
4. Bounded automationHeadless checks, machine-readable outputL3 → L4, evidence bundle and CI gate
5. Closed loopA production signal writes the next intentL4 → L5, maintain stage

Before the first transition: choose the repository and record a baseline

Section titled “Before the first transition: choose the repository and record a baseline”

Pick a repository that has all of the following. If one is missing, fix it first, because every later gate depends on it.

  • Local typecheck, lint, test and build commands that pass on the default branch.
  • Branch protection with required status checks, and a named person who approves production releases.
  • No need to give the agent production credentials.
  • At least five merged changes per month in the loop you plan to move first, so a 30-day window holds enough evidence.

Then export the history the baseline is computed from. Run this in a terminal at the repository root with the GitHub CLI authenticated:

Terminal window
gh pr list --state merged --search "merged:>=$(date -d '90 days ago' +%F)" --limit 1000 \
--json number,title,labels,createdAt,mergedAt,additions,deletions,files,reviews,commits,statusCheckRollup,body \
> merged-prs.json

The date -d form is GNU; on macOS write $(date -v-90d +%F), or type the date 90 days before today by hand.

Record five numbers per loop: lead time from pull request opened to merged, first-pass CI success, review rounds, pull request size, and revert or change-fail rate. Metrics frameworks holds the canonical definitions. The goal is not maximum agent activity. It is shorter lead time with change-fail rate at or below baseline.

Write the loops into loops.yaml at the repository root and review it like code. The one map page has the full schema and a prompt that drafts the file from history; the excerpt below shows the fields this roadmap reads.

# loops.yaml (excerpt): one entry per repeatable class of change
- loop: dependency-bumps
owner: anna-kowalska # GitHub login of a named person, never a team alias
risk_class: low # low | medium | high | critical
level: L3
oracle: contract tests in tests/contract, owned in CODEOWNERS
agent_can_edit_oracle: false
baseline: { lead_time_days: 2.4, change_fail_rate: 0.03 }
next_transition:
to: L4
exit_criterion: evidence bundle on every merged PR for 30 days; 20 sampled reviews found no defect the evidence missed
- loop: checkout-feature-work
owner: piotr.nowak
risk_class: medium
level: L2
oracle: unit and integration tests in the same diff (agent-editable)
agent_can_edit_oracle: true
next_transition:
to: L3
exit_criterion: agent runs typecheck, lint and tests before opening a PR; first-pass CI at or above baseline for 30 days

Then give every loop a pull request label, so later gate checks can filter by loop. Create one label per loops.yaml entry, for example gh label create loop:dependency-bumps, and require it: add a “Loop:” line to the pull request template and a CI check that fails any pull request without exactly one loop: label.

The baseline values above are placeholders; use the numbers from your own export. A loop’s risk_class caps how far it may go: a critical-class loop, such as authentication or money movement, does not go past L3 in the first year, however good its evidence looks; a high-class loop stays at L3 unless the governance policy names it for L4; only low-class loops reach L5. Governance and autonomy assigns the classes.

Which transitions does each loop go through?

Section titled “Which transitions does each loop go through?”

A loop climbs one rung at a time and never skips from L2 to L4. The thresholds match the transformation roadmap’s gates, so your evidence feeds the organization’s quarterly review unchanged. They are starting values to adapt, not research findings: write your own into loops.yaml before the loop starts.

TransitionWhat you add to the repositoryGo: the loop’s evidenceStop: roll back one rung whenSigns
L1 → L2Shared instructions file; deny rules for secretsEvery engineer on the loop uses the committed file; no secret in any agent sessionA secret or production credential appears in a session or transcriptTech lead
L2 → L3Check commands in the instructions file; Plan mode; one isolated checkout per task; required status checksThe agent runs typecheck, lint and tests before it opens a PR; first-pass CI at or above baseline for 30 daysMedian time in review or PR size rises for four weeks; change-fail rate above baselineTech lead (loop owner)
L3 → L4Executable acceptance criteria; protected tests; evidence bundle and its CI check; agent PR review; rollback rehearsedEvidence bundle on every PR for 30 days; 20 sampled human reviews found no defect the evidence missedAn escaped defect traced to the loop; an agent edit to protected tests; a merged PR without a bundleService owner and CTO or delegate
L4 → L5Auto-merge by rule for the low class; progressive delivery with automatic rollback; production signals that write the next intentRisk class low; oracle owned by someone other than the agent’s operator; a full quarter at L4 with change-fail rate at or below baselineAny high-severity incident traced to an unread changeCTO, with security sign-off

L1 → L2 needs no section of its own here: it is one committed instructions file with deny rules for secrets, and shared agent rules covers it.

A loop that stays at L3 or L4 because its risk class or its oracle allows no more is a correct outcome, not a failure of the roadmap.

Move a loop from L2 to L3: the agent works unattended, you read every diff

Section titled “Move a loop from L2 to L3: the agent works unattended, you read every diff”

At L3 you first meet the agent’s work as a finished diff. The repository has to make that diff arrive small, isolated and already checked.

  1. Put the check commands in the instructions file. Add the exact typecheck, lint, test and build commands to CLAUDE.md for Claude Code, AGENTS.md for Codex, or a project rule for Cursor, with the instruction to run them before reporting done. Shared agent rules shows how to keep one core file for all three tools.

  2. Plan before editing. For every task above a one-file fix, the agent drafts a plan in read-only Plan mode and a human accepts it. Keep spec.md and plan.md in the branch when the task has them; the artifact chain describes the handoffs.

  3. Give each task its own checkout. One worktree per task keeps two agents from colliding on files, ports or local state. Give each run only the credentials the task needs, never production ones. Team parallelism covers the queue cap.

  4. Make CI the judge. Mark typecheck, lint and tests as required status checks on the default branch, so a red run blocks merge instead of producing a warning.

  5. Cap the size. Agree a maximum pull request size for the loop, for example 400 changed lines, and split anything larger. DORA names working in small batches as one of its AI capabilities because “AI can easily generate massive blocks of code, which are hard to review and test” (Nathen Harvey and Allison Park, Google Cloud, 10 December 2025).

The unattended run differs per tool; the gate does not.

Run the task in a new worktree, headless, with only the tools it needs. claude -p starts in Manual mode, and dontAsk denies anything the allow list does not name (checked against Claude Code v2.1.283 on 2026-09-26):

Terminal window
claude -w orders-142 -p "$(cat .github/prompts/l3-task.md)" \
--permission-mode dontAsk \
--allowedTools "Read" "Edit" "Bash(npm run typecheck)" "Bash(npm run lint)" "Bash(npm test)" \
--max-budget-usd 5 \
> run-orders-142.md

Under dontAsk the allow list matches these commands exactly: a piped or extended variant, such as npm test | tail -20 or npm test -- orders.test.ts, is denied. The prompt below therefore tells the agent to run each check without pipes or extra arguments.

The final report lands in run-orders-142.md, the same file name the Codex run writes. The shell redirect writes it in the checkout you ran the command from, while the diff sits in the worktree at .claude/worktrees/orders-142/, on a new branch worktree-orders-142. That file, with the exit codes the agent pastes into it, is the evidence the reviewer reads, next to the uncommitted diff in the worktree. Neither the Claude Code run nor the Codex run commits, so both tools hand the reviewer the same two artefacts.

For interactive work, start with claude -w orders-142 and switch to Plan mode with /plan before the agent edits anything.

Save this prompt as .github/prompts/l3-task.md so the commands above can read it.

Move a loop from L3 to L4: you read the spec and the evidence, not the diff

Section titled “Move a loop from L3 to L4: you read the spec and the evidence, not the diff”

L4 changes what the reviewer reads. The move is only safe when the checks that decide “done” sit outside the agent’s reach, so this is the transition to budget the most time for.

  1. Write acceptance criteria a machine can run. Every task in the loop starts from criteria that become tests before implementation. See acceptance criteria.

  2. Protect the oracle. Give the tests, the CI files and loops.yaml, which holds the loop’s own gate, an owner who is not the agent’s operator, and turn on “Require review from Code Owners” for the default branch:

    # .github/CODEOWNERS
    /tests/contract/ @anna-kowalska
    /.github/ @anna-kowalska
    /loops.yaml @anna-kowalska

    Protecting the oracle adds the session deny rules and sandbox profiles per tool, and oracle strength measures whether the tests would catch a wrong change at all.

  3. Require the evidence bundle. Add the pull request template and the CI check from the evidence bundle, so a pull request without spec delta, acceptance results and check output cannot merge.

  4. Add agent review. Run an agent PR review on every pull request in the loop, as a second signal beside the evidence.

  5. Transfer trust in stages. Keep reading every diff while the evidence is also produced (shadow review), then read a random sample, then read evidence only. Trust transfer has the protocol and the sampling workflow.

  6. Rehearse the rollback. Revert one merged change from the loop in staging and time it before you sign the gate.

Move a loop from L4 to L5: nobody reads the code

Section titled “Move a loop from L4 to L5: nobody reads the code”

L5 is rare and belongs to low-risk loops only. The repository adds three things. First, an auto-merge rule for the low class that fires only when the evidence bundle check and every required check pass. Second, progressive delivery with an automatic rollback on a service-level breach. Third, the closed loop from the old wave 5: a production signal opens a triage record and the next intent.md, and the normal merge and release gates still apply to the fix, as intent to production describes. Prove the path with a drill: a synthetic breach must reach triage with evidence attached and without any direct change to production.

A transition works when the loop’s speed improves and its quality holds, measured from the repository, not from impressions. DORA’s 2025 report found both sides of that trade: “Unlike last year, we observe a positive relationship between AI adoption on both software delivery throughput and product performance”, and “However, AI adoption does continue to have a negative relationship with software delivery stability” (Nathen Harvey and Derek DeBellis, Google Cloud, 23 September 2025). Check both at every promotion:

  • The distribution. Report the share of merged change produced at each rung over 30 days, as the one map defines it. Never average it into one number.
  • The paired stability measure. Change-fail or revert rate per loop, against the baseline. A gain in lead time with a loss here is a failure.
  • The audit. Sample 10 merged pull requests from the loop at random and check they carry the evidence the claimed rung requires. A failed audit demotes the loop.
  • The signature. The tech lead signs L1 → L2 and L2 → L3; the service owner and the CTO or a delegate sign L3 → L4; the CTO and security sign L4 → L5. The date and the sample go into loops.yaml.

None of these checks needs anyone to read every line the agent wrote. They read the evidence, the tests’ protection and the production numbers.

What goes wrong when a repository climbs the ladder?

Section titled “What goes wrong when a repository climbs the ladder?”

The pilot loop is trivial. Typo fixes pass every gate and prove nothing. Recovery: pick a production-relevant loop with real tests and a real reviewer, even if it moves more slowly.

The whole repository is promoted at once. Someone writes “we are L4 now” after one loop passes. Recovery: promotions name one loop; the distribution shows the rest where they are.

The agent edits its own oracle. A test “fix” makes CI green on an L4 loop. Recovery: roll the loop back to L3, add the paths to CODEOWNERS and the tool’s deny rules, and re-run the 30-day window from zero.

Review becomes the queue. Pull requests get larger and wait longer after L3. Recovery: enforce the size cap, stop adding parallel tasks, and use the review queue practices before moving any loop further.

Artifacts become paperwork. plan.md and the evidence bundle are filled in but nobody reads them. Recovery: delete every field that no reviewer or check consumes, and make the CI check reject empty sections.

The tool becomes the process. The gates only work in one tool’s interface. Recovery: keep instructions, loops.yaml, the evidence bundle and CODEOWNERS in the repository, so the team can run Claude Code, Codex or Cursor under the same gates.

Where to go next with the repository roadmap

Section titled “Where to go next with the repository roadmap”
Edit page

Last updated:

Cite this page — https://developertoolkit.ai/en/teams/adopt/adoption-roadmap/, developertoolkit.ai