Skip to content

The Autonomy Ladder: Which Level Is Your Workflow At?

The autonomy ladder is Dan Shapiro’s six-level scale, Level 0 to Level 5, for how much of a team’s code coding agents write and how much of it humans still read. Each level describes one loop. From Level 3 upward, a loop climbs only when a machine-checkable signal, such as commands that exit 0, replaces a human reading the diff.

Your team uses Claude Code, Codex or Cursor daily, yet every change still waits for someone to read the diff. Developers, tech leads, CTOs and executives each ask “where are we?” and mean something different. This page gives all four one scale, a five-question test, and the guide for each rung.

What does each level of the ladder look like?

Section titled “What does each level of the ladder look like?”
LevelSite name (Shapiro’s name)Who writesWho readsIn Shapiro’s words
L0By hand (Spicy autocomplete)HumanHuman, every character“not a character hits the disk without your approval”
L1Assisted (The coding intern)Human, mostlyHuman, all of it“you offload specific, discrete tasks to your AI intern”
L2Paired (The junior developer)Human and agent, pairedHuman, every line“feels like you are done. But you are not done”
L3Review manager (The developer)Agent, most of itHuman, as full-time reviewer“Your life is diffs.”
L4Spec manager (The engineering team)Agent, from a specTests, mostly“leave for 12 hours, and check to see if the tests pass”
L5Dark factory (The dark software factory)AgentNobody“It’s a black box that turns specs into software.”

Source: Dan Shapiro, “The Five Levels: from Spicy Autocomplete to the Dark Factory”, 23 January 2026; level names as quoted by Simon Willison, 28 January 2026. The short names are this site’s.

Shapiro puts 90% of “AI-native” developers at Level 2, says “almost everyone tops out” at Level 3, and places “a handful of people” at Level 5. See the evidence behind the ladder, sourced and dated.

A loop is a repeatable class of change with its own trigger, check and stop condition: dependency bumps, flaky-test repair, feature work in one service. Answer per loop, against the repository you ship from. The first question you cannot answer yes with evidence is that loop’s ceiling.

  1. L0 to L1. Did code reach disk this week that you did not type?
  2. L1 to L2. Do you hand the agent whole tasks with acceptance criteria, not completions inside a function you are writing?
  3. L2 to L3. Does the agent run long enough, and often enough, that you first meet its work as a diff?
  4. L3 to L4. Is there a stop condition a machine can evaluate, such as commands that exit 0, so you can leave the run and check the result?
  5. L4 to L5. Does anything merge that no human read, and can you name the oracle that made that safe?

Evidence is an artifact: a CI job, a test command, a merged pull request. The one map lists the evidence each level requires at each lifecycle stage and defines a team’s level as the distribution of its merged changes across loop levels.

How do you move a loop from Level 3 to Level 4?

Section titled “How do you move a loop from Level 3 to Level 4?”

Level 3 is where review becomes the bottleneck; Level 4 is defined by leaving. The move needs a stop condition that runs without you: commands the agent must make pass before it stops.

Use /goal interactively, or run the loop headless in CI with claude -p "Bump dependencies until npm test exits 0" --output-format json --allowedTools "Edit Bash(npm install *) Bash(npm update *) Bash(npm test)" --disallowedTools "Edit(**/tests/**) Edit(**/*.test.*)" so the job can parse the result. Headless runs start in Manual permission mode and refuse any tool call that would need approval (edits, shell commands) unless it is on the allow list; read-only tools such as Read and Grep still run. So grant exactly what the loop needs: the install commands that change dependencies and the test command that checks them. Deny rules are evaluated before allow rules, so the agent can edit source but not the test files that are its stop condition (Claude Code v2.1.283). See goal-directed runs with /goal.

A stop condition the agent can edit, such as a test it may rewrite, is not a stop condition. Path deny rules do not close every route: the agent still edits package.json, including the script npm test runs, and npm install runs the install scripts of every new dependency. So after the agent exits, the CI job re-runs the stop-condition commands itself, fails if the diff touches a test file or the test script, and runs on a disposable runner with no deploy secrets in its environment. Reading evidence instead of code covers what a human reads instead of the diff, and which changes (auth, money, schema migrations) still get read line by line. The tech lead who owns the loop signs off its promotion, and records the stop-condition commands and the last month of rejected diffs they would have caught in the loop register.

Where does placing a loop on the ladder go wrong?

Section titled “Where does placing a loop on the ladder go wrong?”
  • One level for the whole team. Averaging hides the loops that are ready to move. Recovery: keep a register with one row per loop, as in the one map.
  • Claiming a level from confidence. A loop “at Level 4” whose stop condition is “looks fine to me” is at Level 3. Recovery: name the command; if you cannot, demote the loop.
  • Skipping Level 3. Merging unread before you know how the agent fails produces incidents. Recovery: stay at review until your checks would have caught last month’s rejected diffs.
  • Mixing in other maturity models. JetBrains’ AIDEs L1–L5 and Every’s five stages define levels differently. Recovery: do not map them onto the ladder.

Follow your role track (developer, tech lead, CTO, executive) or go to your loop’s rung.

Frequently asked questions

What is the autonomy ladder?

The autonomy ladder is Dan Shapiro's six-level scale, Level 0 to Level 5, for how much of a team's code coding agents write and how much of it humans still read. It runs from autocomplete that needs approval for every character, through delegated tasks, paired sessions and reviewing every diff, to writing specs and checking tests, and finally a dark factory where nobody reads the code.

Which level are most developers at?

Level 2 and Level 3. Shapiro calls Level 2 "where 90% of 'AI-native' developers are living right now" and says of Level 3 that "almost everyone tops out here". His warning is that Level 2, and every level after it, feels like you are done when you are not.

Does a team have one level?

No. A level belongs to a loop, a repeatable class of change with its own trigger, check and stop condition. A dependency-bump loop can run at Level 4 while feature work in the same repository sits at Level 2.

What changes at Level 4?

What the human reads. Below Level 4 a human reads every diff. At Level 4 the developer writes the spec and a stop condition a machine can evaluate, the agent runs unattended, and the human reads the evidence: whether the tests passed and what they covered.

Is Level 5 real?

Shapiro puts only "a handful of people" at Level 5, in teams of fewer than five, and Simon Willison's write-up identifies one of them as StrongDM's AI division. It is rare. Shapiro's Level 5 has nobody reading the code; what makes that survivable is the strength of the verification around it (see Level 5: You Run the Software Factory).

Edit page

Last updated:

Cite this page — https://developertoolkit.ai/en/ladder/, developertoolkit.ai