Level 3: You Review the Diffs
Level 3 of the autonomy ladder is the stage where agents write most of the code in parallel sessions and a developer reviews every diff before it merges. It is a stage, not a destination: review bandwidth caps it. Teams leave it by moving one change class at a time from line-by-line reading to evidence and automated gates.
This page is for developers who already run two or more agents at once, and for tech leads whose review queue grows faster than the team. Five agents are running and four of them are right. You find out which four by reading, and reading is the one part of the pipeline that did not get faster. CTOs and executives can skip to the measures section: it defines the four numbers that show whether your teams have hit the ceiling.
What you get from this Level 3 guide
Section titled “What you get from this Level 3 guide”- A four-point test that tells you whether you are at Level 3 or still at Level 2.
- The parallel-agent setup in Claude Code, Codex and Cursor, with the commands to start it.
- A map of what each tool’s review automation does, and the one detail that decides whether it gates anything.
- Six steps that raise the review ceiling, including a CI check that turns Claude Code’s review into a blocking gate.
- Three copy-paste prompts: measure your review queue, make agents hand you evidence, and design the review gate.
- The exit criteria for moving a change class to reading evidence instead of code.
Are you at Level 3? The four-point test
Section titled “Are you at Level 3? The four-point test”Dan Shapiro’s ladder (The Five Levels: from Spicy Autocomplete to the Dark Factory, 23 January 2026) describes Level 3 in three words: “Your life is diffs.” He adds that “almost everyone tops out here.” You are at Level 3 when all four statements are true:
- The agent writes most of each change. You edit its output, but you rarely type the first draft.
- More than one agent works at a time, each in its own checkout, while you do something else.
- You read every diff before it merges. Reading, not the agent, is the step that decides.
- Your queue of unread work grows during the day. Agents finish faster than you review.
If statement 2 is false, you are at Level 2 and the Level 1–2 guide comes first. If statement 3 is false for some changes and nothing replaced your reading, you are not at Level 4. You are at Level 3 with the check removed, which is the most common way to fail at this rung.
How do you run parallel agents at Level 3?
Section titled “How do you run parallel agents at Level 3?”The mechanical unlock is filesystem isolation: one git worktree per agent, so parallel edits never touch the same files. The second unlock is one screen that shows which sessions need you.
Start each task in its own worktree, in the foreground or in the background (checked with Claude Code 2.1.285):
# Terminal: one isolated session per taskclaude --worktree feature-authclaude --worktree fix-pagination --bg # runs in the background, prints an id
# Terminal: agent view lists every background session and what it needsclaude agentsAgent view is a research preview. Inside a session, the bundled /batch skill splits one large change into 5 to 30 units, each run by a background subagent in its own worktree. See agent view in Claude Code.
Worktree support is on by default since Codex 0.156.0, but a session runs in a worktree only when you opt in (checked with codex-cli 0.157.1):
# Terminal: one session per task, each in a new managed Git worktreecodex --worktree
# Terminal: browse every agent session on the local app-server daemoncodex agentsIn the TUI, /worktree does the same for a running session and /agents opens the command center. See Codex worktrees.
Cursor documents worktrees as the isolation unit: “Worktrees let Agent work in isolated Git checkouts” (checked 2026-08-28; cursor.com was unreachable from the writing environment on 2026-09-26). You run and watch parallel agents from the Agents Window. Check the current setup steps in Cursor’s documentation before you script anything around them.
Worktrees isolate files, not ports, databases, caches or a shared dev server. Give each agent its own port block and its own local state, or two green runs will disagree about the same machine. The git worktrees guide covers ports and state; ephemeral environments cover the rest. For fleets you start from a script, see tmux for agent fleets and herdr, the agent multiplexer.
Why review bandwidth is the Level 3 ceiling
Section titled “Why review bandwidth is the Level 3 ceiling”Generation scales with the number of agents; reading scales with the number of reviewers. Two independent datasets measured what happens in between.
Faros AI’s AI Engineering Report 2026: The Acceleration Whiplash (April 2026) covers two years of telemetry from 22,000 developers and more than 4,000 teams. Throughput rose: epics completed per developer +66.2%, task throughput per developer +33.7%, pull request merge rate per developer +16.2%. Quality and review fell behind in the same window:
| Faros measure (April 2026) | Change | What it tells you |
|---|---|---|
| Median time in review | +441.5% | The queue: how long a pull request sits in review |
| Median time to first review | +156.6% | The response time: how long before anyone looks |
| Pull requests merged with no review | +31.3% | Review turning into approval |
| Incidents per pull request | +242.7% | What the skipped reading cost |
| Bugs per developer | +54% | The same, measured per person |
Faros’s data is vendor telemetry from its own customer base, so read it as a direction, not a forecast for your team. DX measured the load in a different unit: median pull request size grew “from 44 lines to 72 lines per pull request between July 2025 and June 2026” (DX, Justin Reock, 17 June 2026). More diffs, bigger diffs, same reader. The state of agentic engineering page has the full numbers with sources.
What does review automation cover in Claude Code, Codex and Cursor?
Section titled “What does review automation cover in Claude Code, Codex and Cursor?”All three tools ship an automated reviewer. None of them is a merge gate until you make it one.
| Claude Code | Codex | Cursor | |
|---|---|---|---|
| Local review | /code-review in a session (/review is an alias; --fix, --comment) | /review in the TUI; codex review --base main or codex exec review from a script | Not verified here (cursor.com unreachable on 2026-09-26) |
| Deep or security review | claude ultrareview [target] or /code-review ultra: a cloud fleet in which “every reported finding is independently reproduced and verified”, typically 5 to 10 minutes | @codex security review on the pull request (checked 2026-08-28) | Security Agents for the vulnerability pass (checked 2026-08-28) |
| Review on the pull request | Managed Code Review: research preview, Team and Enterprise, tuned by a root REVIEW.md | @codex review on GitHub and GitLab, automatic reviews, rules in AGENTS.md (checked 2026-08-28) | Bugbot “reviews pull requests and identifies bugs, security issues, and code quality problems” (checked 2026-08-28) |
| Can it block a merge? | Not by itself: the check run “always completes with a neutral conclusion” | Not documented (checked 2026-08-28) | PR Routing & Approval “can approve low-risk PRs when your criteria are met” (checked 2026-08-28) |
Claude Code rows were checked against Anthropic’s documentation and claude --help on 2026-09-26; the Codex CLI rows against codex --help 0.157.1. Costs for Ultrareview and Code Review are billed as usage credits; see automated code reviews in Claude Code for the current figures and plan limits. The per-tool guides for Codex review and Cursor Bugbot cover setup.
An automated reviewer that cannot block is a second opinion. It raises your ceiling only after you decide which of its findings fail the build.
How do you raise the review ceiling at Level 3?
Section titled “How do you raise the review ceiling at Level 3?”Six steps, in this order. Each one removes a category of work from the human queue before the next one starts.
-
Measure the queue. Record median time to first review, median time in review, and median pull request size per repository for the last 50 to 100 merged pull requests. Use the first prompt below. Without a baseline, you cannot tell whether any later step helped.
-
Make the deterministic gate block. Types, tests, lint, a schema check, and a grep for the pattern you are migrating away from all exit non-zero on failure. They cost nothing to run and cannot be talked round. Mark them as required checks, so no pull request reaches a human while one is red.
-
Let the automated reviewer fail the build on a short list. Pick the finding classes you trust, such as Important findings in Claude Code Review, and block on those only. Claude Code writes a machine-readable severity count into the check run, which a CI step can read with a token that has read access to checks:
Terminal window # CI step, after the "Claude Code Review" check run completesCHECK_RUN_ID=$(gh api "repos/$OWNER/$REPO/commits/$SHA/check-runs" \--jq '.check_runs[] | select(.name=="Claude Code Review") | .id')counts=$(gh api "repos/$OWNER/$REPO/check-runs/$CHECK_RUN_ID" \--jq '.output.text | split("bughunter-severity: ")[1] | split(" -->")[0] | fromjson')important=$(echo "$counts" | jq -e '.normal') || { echo 'No severity counts in check run'; exit 1; }if [ "$important" != "0" ]; thenecho "Claude Code Review reported $important Important findings"exit 1fiThe first call finds the ID of the
Claude Code Reviewcheck run on commit$SHA, as Anthropic’s Code Review page describes. Thenormalkey holds the count of Important findings.jq -emakes the step fail closed: if the check run carries no severity counts, the build stops instead of passing. -
Make agents hand you evidence, not a summary of the diff. Every agent pull request states the behaviour it changes, maps each acceptance criterion to a check that ran, and lists any test, fixture or CI file it touched. Use the second prompt below in your context file or task template.
-
Split the queue by risk class. Authentication and authorization, money, schema, data migrations, and any change that edits its own tests always get a human reader. Everything else can be triaged on evidence. The agent PR review triage turns this into a six-step protocol with verdicts.
-
Move one change class to evidence review. Choose the class with the strongest checks, such as dependency bumps or copy changes, and approve it on the evidence from step 4 while you keep reading everything else. The exit criteria below say when a class is ready.
Who signs off stays the same throughout: a named reviewer approves every merge. What changes is what that person reads. At Level 3 it is the diff; after each class moves, it is the evidence plus the diff only for the escalation classes.
Copy-paste prompts for the Level 3 review ceiling
Section titled “Copy-paste prompts for the Level 3 review ceiling”What should tech leads, CTOs and executives measure at Level 3?
Section titled “What should tech leads, CTOs and executives measure at Level 3?”Four measures, per repository, show whether a team has hit the ceiling. Adopt the definitions as written and set each threshold from your own last-quarter baseline, not from another company’s numbers.
| Measure | Definition | Ceiling signal |
|---|---|---|
| Time to first review | Median hours from pull request opened to first review submitted | Rises while agent count rises |
| Time in review | Median hours from opened to merged | Rises faster than merge rate |
| Unreviewed merge share | Merged pull requests with zero reviews, divided by all merged | Any rise that no written policy allows |
| Median pull request size | Median of additions plus deletions | Grows beyond what one reviewer reads in one sitting |
A rising unreviewed merge share is the number to watch. It means reviewers have stopped reading without a gate taking their place. The review queue guide for teams covers how to staff and route the queue.
What breaks when five agents share one reviewer?
Section titled “What breaks when five agents share one reviewer?”When is a change class ready to leave Level 3?
Section titled “When is a change class ready to leave Level 3?”Leave one class at a time, never the whole repository at once. A change class is ready for evidence review when all four hold:
- A deterministic gate for the class is a required check and has blocked at least one bad change.
- Agents cannot weaken the oracle unnoticed: any edit to tests, fixtures or CI in the same pull request sends it to a human.
- Every pull request in the class carries the evidence summary from the second prompt.
- Your reading over a run of recent pull requests in the class (for example, the last 10) found nothing the checks had missed.
Keep a short trust log of each class, when it moved, and what sent it back. Reading evidence instead of code explains what you read instead of the diff and the signs that you stopped reading too early. Reviewing an agent’s pull request is the day-to-day triage. When most of your merged changes are approved on evidence, your runs can last hours, and Level 4 begins.
Where to go next from Level 3
Section titled “Where to go next from Level 3”Frequently asked questions
What is Level 3 on the autonomy ladder?
Level 3 is the stage where agents write most of the code in parallel sessions, each in its own worktree, and a developer reviews every diff before it merges. Dan Shapiro sums the rung up as "Your life is diffs" and says almost everyone tops out there.
Why is review bandwidth the ceiling at Level 3?
Generation scales with the number of agents and reading does not. Faros AI's April 2026 report measured median time in review up 441.5% and 31.3% more pull requests merged with no review, while DX measured median pull request size growing from 44 to 72 lines between July 2025 and June 2026.
Do AI code reviewers remove the Level 3 ceiling?
Not by default. Claude Code's managed Code Review check run always completes with a neutral conclusion, so it never blocks a merge until you read its severity counts in your own CI. An automated reviewer raises the ceiling only once you decide which findings may fail the build.
How do you leave Level 3?
One change class at a time. When a class has a blocking deterministic gate, a protected test suite, an agent-written evidence summary, and a run of reviews in which your reading found nothing the checks missed, you approve that class on evidence and keep reading code only for the escalation classes.