Skip to content

One map: the autonomy ladder, the lifecycle and the factory stations

The one map combines three axes in a single grid. Maturity is the autonomy ladder, Level 0 to Level 5, measured per loop. Process is the six lifecycle stages, from plan to maintain. Capability is the six factory stations, from intent to release. Each cell names the evidence a loop must produce before it may run at that level.

Your adoption roadmap says “wave 3”, the CTO scorecard says “Level 2”, a vendor deck says “around L3”, and the compound engineering guide says “stage 2”. Then the board asks: what level are we at? Four numbers, four definitions, and none says which work may merge without a human reading it. This page reconciles them: one vocabulary for maturity, process and capability, and a map from every older scheme on this site back to it.

  • Developers: which evidence your loop has to produce before you stop reading its diffs, stage by stage.
  • Tech leads: a loop register you can copy, and a rule for promoting or demoting a loop on evidence rather than on confidence.
  • CTOs: a metric definition for team and organization level that cannot be averaged into a vanity number, and risk classes that sit beside levels instead of inside them.
  • Executives: one sentence to report upward (“share of merged change produced at Level 4 or above, with its failure rate”) and the list of outside models that are not the same thing.

Each axis answers a different question with a different unit, and confusing the units starts most maturity arguments.

AxisQuestion it answersValuesUnit it applies toCanonical page
MaturityWho reads what before a change merges?L0 By hand · L1 Assisted · L2 Paired · L3 Review manager · L4 Spec manager · L5 Dark factoryOne loopThe autonomy ladder
ProcessWhere in the life of a change are we?Plan · design · build · test · deploy · maintainOne stage of the lifecycleThe AI-native lifecycle
CapabilityWhat machinery exists to run that stage without a person?Intent · harness · loop · graph · verification · releaseOne station of the factoryLevel 5: the six stations

A loop is a repeatable class of change with its own trigger, oracle and stop condition: dependency bumps, flaky-test repair, feature work in the checkout service, schema migrations. A level belongs to a loop, not to a company: a dependency-bump loop with a deterministic check can run at Level 4 while the feature work beside it sits at Level 2, and both answers are correct at once.

The ladder’s six levels follow Dan Shapiro, “The Five Levels: from Spicy Autocomplete to the Dark Factory” (January 2026); the short names are this site’s. Shapiro places most AI-native developers at Level 2 and says almost everyone tops out at Level 3, which is why most cells in the grid below sit in those two columns for most teams.

The tool trees (Claude Code, Codex, Cursor) are organized by nine task nodes. They are navigation, not a fourth axis, and each one maps onto the process and capability axes:

Task node in the tool treesLifecycle stageStations it mostly documents
Quick startSetup, before the lifecycleIntent, harness
PlanPlan, designIntent
BuildBuildHarness, loop, graph
TestTestVerification
ShipDeployVerification, release
OperateMaintainLoop, release
AutomateBuild to maintainLoop, graph, release
TeamAcross stages (shared rules and people)Intent
ReferenceNoneNone

The grid: what evidence does each level require at each stage?

Section titled “The grid: what evidence does each level require at each stage?”

Rows are lifecycle stages; each row names the stations that produce its evidence. Columns are ladder levels. A cell is the evidence that must exist, and be produced by a machine wherever the level is L4 or above, before a loop may claim that level at that stage. L0 and L1 are left out: at those levels the evidence is the human who typed or accepted each line.

Stage (stations)L2 PairedL3 Review managerL4 Spec managerL5 Dark factory
Plan (intent)Intent lives in the prompt and the developer’s headA written task with acceptance criteria in the issueA committed intent.md or spec with machine-checkable acceptance criteria, accepted by a named owner before the runIntents arrive from a queue (tracker, alert, incident), each carrying its own oracle; the owner signs the intent, not the code
Design (intent, graph)Decided live in the sessionA plan reviewed in plan mode before implementation startsCommitted spec.md and plan.md; architecture constraints encoded as rules or fitness functions the agent cannot editDecomposition runs as a versioned workflow; a design change outside the agreed envelope escalates to a human
Build (harness, loop, graph)The human watches every edit and answers permission promptsThe agent runs unattended in a worktree or sandbox; the evidence is the diff plus a local test runAn explicit permission mode and sandbox for the run type, a stop condition a machine evaluates, and a retry boundEvent-triggered runs with no approval prompts; the harness configuration is versioned and reviewed like code
Test (verification)Tests the human writes or readsCI is green and a human reads the diff, tests includedAn oracle outside the agent’s reach (protected tests, acceptance tests from the spec) and an evidence bundle attached to the pull requestOracle strength is measured, not assumed; CI rounds are bounded; the factory itself has evals
Deploy (verification, release)The human merges what they wrote with the agentA human approves the merge after reading every diffThe merge decision rests on the evidence bundle and the risk class; code is read only for the classes that require itLow-risk classes auto-merge by rule; progressive delivery with automatic rollback on a service-level breach
Maintain (loop, release)Humans watch productionThe agent triages on request; a human drives the fixScheduled or event-triggered runs open pull requests from alerts, with runtime signals in the evidenceA breached control band writes the next intent, and every incident adds a case to the oracle

Read the grid by row, not by column. A loop is at the lowest level any of its stages supports. A loop with an L4 test row and an L3 deploy row is an L3 loop, because a human still reads every diff before merge.

The canonical pages for the cells are evidence instead of diffs (what a human reads at L4), the evidence bundle (what the pull request must carry), oracle strength (how strong the test row is) and progressive delivery (the L5 deploy row).

How do you score a team or an organization?

Section titled “How do you score a team or an organization?”

A team does not have a level. It has a distribution: the share of its merged change produced by loops at each level. Use this definition as it stands.

FieldDefinition
UnitOne merged pull request (or merged change set), tagged with the loop that produced it
Level of a changeThe level the loop held on the merge date, established by the evidence test in the next section, not by self-report
WindowA rolling 30 days for teams; a quarter for the organization
Team levelShare of merged changes at L0–L2, L3, L4 and L5 in the window
Organization levelThe same distribution across all teams, plus the number of loops at L4 or above
Paired stability measureChange failure rate (or revert rate) per level, in the same window
ForbiddenAveraging levels into one number (“Level 3.4”), or reporting the best loop as the team’s level

The paired stability measure is not optional. DORA’s 2025 report (Nathen Harvey and Derek DeBellis, Google Cloud, 23 September 2025) found “a positive relationship between AI adoption on both software delivery throughput and product performance”, and in the same breath: “However, AI adoption does continue to have a negative relationship with software delivery stability.” A distribution that moves right while failures rise is not progress.

An illustrative team, not a measured one:

LoopLevelMerged PRs, 30 daysShare
Dependency bumpsL46030%
Flaky-test repairL42010%
Checkout feature workL38442%
Schema migrations (critical class, always read)L3126%
Payment-provider integrationL22412%

The report line is: “Checkout: 40% of merged change at L4, 48% at L3, 12% at L2, with change failure rate per level beside it.” The team is not “Level 3”. It runs two L4 loops, and the next promotion decision is about one named loop, not the team.

Run the distribution from your merged pull requests

Section titled “Run the distribution from your merged pull requests”

Export the last 30 days of merged pull requests with the GitHub CLI, in your terminal, from the repository root:

Terminal window
# GNU date (Linux); on macOS use: date -v-30d +%F
SINCE=$(date -d '30 days ago' +%F)
gh pr list --state merged --search "merged:>=$SINCE" --limit 1000 \
--json number,title,labels,author,headRefName,files,commits,reviews,mergedBy,mergedAt,body > merged-prs.json
# Slim the export: the raw file is one JSON line, often several MB, which a
# file-reading tool cannot page through.
jq '[.[] | {number,title,labels:[.labels[].name],author:.author.login,branch:.headRefName,
files:[.files[].path],commits:[.commits[]|{headline:.messageHeadline,body:.messageBody,
authors:[.authors[]|{login,name}]}],reviews:[.reviews[]|{author:.author.login,state}],
mergedBy:.mergedBy.login,body}]' merged-prs.json > prs-slim.json

Save the prompt below as prompts/level-distribution.txt, then run it read-only in your tool. --tools "Read" removes every other Claude Code tool; in Codex, the :read-only permission profile (beta) replaces the legacy --sandbox flag, and the two must not be combined:

Terminal window
claude -p "$(cat prompts/level-distribution.txt)" --tools "Read" > level-distribution.md

A level describes a loop. A risk class describes a single change. They answer different questions, and mixing them produces the worst failure on this page: “we are Level 4, so the agent may merge the auth change”.

Risk classTypical changesMay merge on evidence alone?Who signs
LowDocs, copy, internal tooling, patch-level dependency bumps with green contract testsYes, in L4 loops; by rule in L5 loopsThe loop owner, through the rule
MediumFeature logic behind a flag, internal API changesYes, when a human reads the evidence bundleA reviewer of the evidence
HighPublic API contracts, data model changes, performance-critical paths, any change to tests or CI configurationNo: evidence plus a human reading the sensitive pathsThe code owner
CriticalAuthentication, money movement, schema migrations on production data, secrets and permissions, the agent’s own harnessNo: a human reads the code, and deployment needs a separate approvalA named owner and a second approver

The class caps what a level may do with a given change. An L5 loop that opens a critical-class pull request escalates it exactly as an L2 team would. Changes to tests and CI are high class on purpose: they are the oracle, and an agent that edits its own oracle has turned the check into a suggestion. The policy that assigns classes lives in governance and autonomy; the list of changes where someone still reads code lives in evidence instead of diffs.

How do you prove a loop is at the level you claim?

Section titled “How do you prove a loop is at the level you claim?”

A level is a claim about evidence, so it can be checked without reading the code. Run this audit before promoting a loop and once a quarter after.

  1. Sample. Pick 10 merged pull requests from the loop at random from the last window.
  2. Check the evidence for the claimed level. For L3, each has a human approval after a diff review. For L4, each carries an evidence bundle, its tests and CI files are untouched (or code-owner approved), and the run recorded a stop condition. For L5, each merged by rule with no human approval, and the rollback path has fired at least once in a drill.
  3. Check the four ledger questions from software factories: what oracle decides done, can the agent fake it, how long until a wrong answer surfaces, and what the blast radius is.
  4. Compare stability. The loop’s change failure rate or revert rate over the window must be no worse than at its previous level.
  5. Record the verdict in the loop register with the date and the sample, and have the tech lead sign the promotion.

A failed check demotes the loop; it does not start a debate. The tech lead signs promotions, and the CTO owns the risk-class policy.

The artifact that makes this repeatable is a loop register, one file per repository:

# loops.yaml: one entry per repeatable class of change
- loop: dependency-bumps
owner: platform-team
trigger: Renovate pull request opened
stages: [build, test, deploy]
level: L4
oracle: required CI jobs defined outside the paths the agent may write
agent_can_edit_oracle: false
time_to_surface_wrong_answer: minutes in CI, hours in canary
blast_radius: one service, reversible by revert
merge_on_evidence: [low]
escalate_to_code_reading: [high, critical]
last_audit: 2026-09-26, 10 of 10 PRs carried an evidence bundle
next_transition:
to: L5
exit_criterion: auto-merge for the low class with canary rollback, 60 days without a loop-caused revert
- loop: checkout-feature-work
owner: checkout-team
trigger: ticket moved to Ready with acceptance criteria
stages: [plan, design, build, test, deploy]
level: L3
oracle: unit and integration tests in the same repository (agent-editable)
agent_can_edit_oracle: true
merge_on_evidence: []
next_transition:
to: L4
exit_criterion: acceptance tests generated from the spec and protected by CODEOWNERS; evidence bundle required by CI

Which older schemes does the one map replace?

Section titled “Which older schemes does the one map replace?”

This site used to run several numbering schemes side by side. They are retired, and each maps back to one axis of the map. If you meet one in an older bookmark or an internal deck, translate it with this table.

Retired schemeWhat it describedWhat replaces itWhere to read it now
Waves 0–5A repository rollout in fixed wavesLadder transitions, one loop at a timeAdoption roadmap
Weeks 1–8, “Month 2+”A developer’s learning calendarTrack steps with exit criteriaDeveloper track
Weeks 1–12, “Month 4+”, “Phased Enterprise Rollout”An organization rollout calendarStop/go gates stated as ladder transitions per loopTransformation roadmap
Adoption Curve, Phase 0, org-size modelsWho adopts first, and in what orderLadder transitions plus a designed pilotPilot design
Phases 1–4A tool migration checklistUnnumbered migration stepsMigration guide
“Tier 2 / Tier 3”Parallel agents, then overnight runsL3 (parallel agents, a human reads every diff) and L4–L5 (unattended runs)Team parallelism
“Reactive → Strategic Leader”Leadership maturity labelsThe Level 1–4 bands the CTO scorecard scores: Assisted, Paired, Review manager, Spec managerCTO scorecard guide
Every’s stage ladder (stages 0 to 5), used as a second ladderAn individual’s adoption pathThe ladder, per loop; Every’s stages remain Every’s own model (see below)Compound engineering
“Risk tiers 0–3”How dangerous a change isRisk classes: low, medium, high, criticalGovernance and autonomy

The CTO scorecard band is a self-assessment of practices (tooling, review, governance), not the level distribution of merged change. Never report it as “our level”; quote it separately from the distribution.

The pattern behind the table: a date is not evidence, and a measure of danger is not a measure of maturity.

Averaging levels into one number. “The org is at Level 3.2” hides the two loops that could run dark and the critical loop that should never have left L3. Recovery: report the distribution and the count of loops at L4 or above, never a mean.

Scoring the team by its showcase loop. One polished dependency-bump loop at L4 becomes “we are an L4 team”. Recovery: the team level is the share of merged change, so a loop that produces 5% of merges moves the distribution by 5%.

Reading a risk class as permission. A loop promoted to L4 starts merging schema migrations on evidence alone. Recovery: classes cap levels; add the critical paths to CODEOWNERS and make CI fail when an evidence-only merge touches them.

Claiming a level the stations cannot support. A loop is called L4 while its oracle sits in the same diff the agent writes. Recovery: run the audit above; a loop whose agent can edit its oracle is L3 until the oracle moves out of reach.

Moving right while getting worse. The share at L4 rises and so does the revert rate. Recovery: demote the loop that caused the reverts, fix its oracle, and re-audit before promoting it again.

Frequently asked questions

What is the one map?

A single grid with three axes. Maturity is the autonomy ladder, Level 0 to Level 5, measured per loop. Process is the six lifecycle stages, plan, design, build, test, deploy and maintain. Capability is the six factory stations, intent, harness, loop, graph, verification and release. Each cell names the evidence a loop must produce before it runs at that level.

What level is my team or organization at?

Not one level. A team's level is the distribution of its merged changes across the levels of the loops that produced them, for example 40% at Level 4, 48% at Level 3 and 12% at Level 2 over 30 days. An organization is the same distribution across all its teams, reported beside a stability measure for each level.

Is a risk class the same as a maturity level?

No. A risk class (low, medium, high, critical) is a property of one change, such as a schema migration or a copy edit. A level is a property of a loop. A Level 5 loop still escalates a critical change to a human who reads the code, and a Level 2 team still has low-risk changes.

Are JetBrains AIDEs, Every's stage ladder and the DORA AI Capabilities Model the same as the ladder?

No. They are separate models with their own definitions, and this site does not map them onto the ladder. The AIDEs L1–L5 labels collide with the ladder's level numbers but mean something else; the DORA model describes seven organizational capabilities, not levels.