Skip to content

Running the review queue when agents open the pull requests

A review queue for agent-opened pull requests stays healthy when pull requests arrive no faster than the team can review them. A tech lead keeps that balance with four controls measured by time in review: a work-in-progress (WIP) cap that pauses agent dispatch, a PR size budget, risk lanes that decide who reviews what, and load-balanced reviewer assignment.

Your six engineers each run two or three agents, and on Monday morning there are 34 open pull requests. Each one looks reasonable. None of them is urgent. By Thursday the oldest has been waiting four days, two people have quietly started approving without reading, and a senior engineer tells you in a one-to-one that they spend the whole day reviewing and never build anything. That is the review ceiling from Level 3 of the autonomy ladder, and this page is the operating manual for it.

This page is for the tech lead who owns one team’s flow of work and assumes agents already open pull requests; if not, start with the agent-ready backlog.

What a managed review queue gives your team

Section titled “What a managed review queue gives your team”
  • A one-line capacity check that tells you whether the queue will grow or drain this week.
  • A WIP cap derived from your own numbers, and a script that stops agent dispatch when the queue is full.
  • A PR size budget that the agents follow and CI enforces, with an escape hatch you control.
  • A risk-lane table that routes each pull request to a green pipeline, an automated reviewer, one human, or two.
  • Eight metric definitions you can compute this afternoon with gh and jq, and the thresholds that trigger a change.
  • A sustainable-pace policy: session caps, a review rota, and the signals that tell you to lower concurrency.

Generation scaled and reading did not. Faros AI’s AI Engineering Report 2026: The Acceleration Whiplash (April 2026, telemetry from 22,000 developers) measured both sides at once. The pull request merge rate per developer rose 16.2%. In the same window, median time in review rose 441.5%, and median time to first review rose 156.6%. Incidents per pull request rose 242.7%, and 31.3% more pull requests merged with no review at all. The pull requests also grew: Faros reports PR size up 51%. DX’s panel (June 2026) saw the median go “from 44 lines to 72 lines per pull request between July 2025 and June 2026”.

The unreviewed-merge number is the one to worry about: in an unmanaged queue, reviewers do not refuse the load, they stop reading and keep approving.

DORA names the counter-measure in its AI Capabilities Model. On “working in small batches”, Google Cloud’s DORA write-up (December 2025) says: “AI can easily generate massive blocks of code, which are hard to review and test. Enforcing the discipline of small batches counteracts this risk”. The rest of this page turns that sentence into rules you can enforce.

A queue is stable only when the arrival rate is below the service rate. Measure both for the last four weeks, not from memory:

  • Arrival rate (λ): non-draft pull requests opened per working day, agent and human together.
  • Review throughput (μ): pull requests your reviewers can finish per working day. Compute it as reviewer-hours spent on review per day divided by the median hours of review per pull request.

A worked example with round numbers: four reviewers each protect 90 minutes a day for review, which is 6 reviewer-hours. A median pull request takes 30 minutes of review, so μ is 12 pull requests a day. If agents and humans open 18 a day, the queue grows by six every day, forever. No amount of reviewer effort fixes that; only fewer or smaller arrivals do.

Queueing theory adds a warning: waiting time grows sharply, not linearly, as utilisation (λ ÷ μ) approaches 1. A team whose reviewers are fully loaded waits a long time even when the arithmetic works. Plan for review to be busy about three quarters of the time, and treat anything above that as a queue about to jam.

Little’s law links the three numbers you care about: items in the queue = throughput × time in the queue. Choose the time in review you want, multiply by your measured throughput, and that is your WIP cap. With μ = 12 per day and a target of one working day in review, the cap is 12 open non-draft pull requests awaiting review, team-wide.

The cap only works if it stops dispatch, not pull request creation: an agent that cannot open a pull request leaves a branch nobody sees, which is the same inventory hidden somewhere worse. So put the gate in front of whatever starts agent work: the issue-to-PR workflow, a scheduled routine or automation, or the command your engineers use to launch a background agent.

scripts/review-queue-gate.sh
#!/usr/bin/env bash
# Exit 1 when the review queue (open, non-draft, not yet approved) is at or above the WIP cap.
set -euo pipefail
CAP="${REVIEW_WIP_CAP:-12}"
open=$(gh pr list --state open --search "draft:false -review:approved" \
--limit 200 --json number --jq 'length')
if [ "$open" -ge "$CAP" ]; then
echo "Review queue full: $open PRs awaiting review (cap $CAP). Review before dispatching more." >&2
exit 1
fi
echo "Review queue $open/$CAP: dispatch allowed."

The script counts exactly what the cap defines: every open non-draft pull request without an approval, from agents and humans alike. Approved pull requests waiting to merge no longer need a reviewer, so they do not count. Keep the cap in one place, the REVIEW_WIP_CAP variable in your CI settings, so changing it is a one-line decision rather than a code change. Review the cap every two weeks against the metrics below; it is derived from throughput, so it moves when throughput moves.

Enforce a PR size budget the agents follow

Section titled “Enforce a PR size budget the agents follow”

Small pull requests are the cheapest capacity you can buy: they are faster to review, they pinpoint the fault, and they revert cleanly. Set the budget from your own history, for example at the size below which three quarters of your human-reviewed pull requests fall. If you have no history, start at 400 changed lines excluding lockfiles and snapshots, and tune it after a month. That number is a starting point, not a finding.

Tell the agents first, in the instruction file every session reads, so they plan the split before writing code:

CLAUDE.md or AGENTS.md (same text in Cursor project rules)
## Pull request budget
- One concern per pull request. Stay under 400 changed lines, excluding lockfiles and snapshots.
- If a task needs more, stop before coding and propose a split into stacked pull requests,
each independently green and reviewable on its own.
- Never change tests, CI config or lint rules in the same pull request as the code they check.
Open those separately with the label `oracle-change`.
- Open every pull request with `gh pr create --label agent` and put the evidence bundle
(acceptance criteria, commands run, results) in the description.

Then make CI the backstop, because an instruction is a request and a failing check is a rule:

.github/workflows/pr-size-budget.yml
name: pr-size-budget
on:
pull_request:
types: [opened, synchronize, reopened, labeled, unlabeled]
permissions:
contents: read
jobs:
size:
if: ${{ !contains(github.event.pull_request.labels.*.name, 'size-exception') }}
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
with:
fetch-depth: 0
persist-credentials: false
- name: Compare changed lines with the budget
env:
BASE: ${{ github.event.pull_request.base.sha }}
BUDGET: 400
run: |
changed=$(git diff --numstat "$BASE" HEAD -- . \
':(exclude,glob)**/package-lock.json' ':(exclude,glob)**/pnpm-lock.yaml' \
':(exclude,glob)**/*.lock' ':(exclude,glob)**/__snapshots__/**' \
| awk '{ a += $1; d += $2 } END { print a + d + 0 }')
echo "Changed lines: $changed (budget $BUDGET)"
if [ "$changed" -gt "$BUDGET" ]; then
echo "::error::Over the $BUDGET-line review budget. Split the change, or ask the tech lead for size-exception."
exit 1
fi

The size-exception label is the escape hatch for mechanical changes such as a rename across the codebase. Make it team policy that only the tech lead applies it. GitHub does not enforce that, because anyone with triage or write access can add a label, so the weekly metrics list every use of size-exception and who applied it (from the pull request’s timeline), which makes misuse visible.

Not every pull request needs the same reader. The evidence-based review model says what a human reads instead of the diff; the lanes below say which pull requests reach a human at all. Write the table down, commit it as docs/review-lanes.md next to CODEOWNERS, and let the paths decide the lane.

LaneWhat falls in itGate before mergeWho signs off
A. Green pipelineDependency patch bumps, docs, copy changes, generated clients, changes that match a pattern you have already approved many timesTypes, tests, lint, and the size budget pass; automated reviewer raises no blocking findingThe pipeline; the tech lead samples a few each week
B. Automated plus sampled humanOrdinary feature and bug-fix work inside one module with strong testsLane A gates plus an automated reviewer; a human reads the evidence bundle, and reads code only when the evidence is thinOne reviewer, assigned by load
C. Code ownerPublic APIs, shared libraries, performance-sensitive paths, anything crossing module boundariesLanes A and B plus a code owner’s approvalThe owning team via CODEOWNERS
D. Two humans, code readAuthentication and authorisation, money, schema and data migrations, and any oracle-change (tests, CI, lint, or type config)All gates plus two approvals, one from a code owner; a deep automated review firstTwo named reviewers; the tech lead for oracle changes

CODEOWNERS enforces the code-owner approval for lanes C and D without a new tool, provided the branch protection rule “Require review from Code Owners” is on:

.github/CODEOWNERS
/db/migrations/ @acme/data-owners
/src/auth/ @acme/security-reviewers
/src/billing/ @acme/security-reviewers
/.github/ @acme/tech-leads
/tests/ @acme/tech-leads
/eslint.config.js @acme/tech-leads

GitHub branch protection has a single required-approvals count for the whole branch, and that shapes the other lanes. Keep the count at 1, and satisfy lane A in one of two ways. Either bind an auto-approver to the lane A criteria (Cursor PR Routing & Approval, or a GitHub Action that approves only lane A paths; the Action needs the repository setting “Allow GitHub Actions to create and approve pull requests”), or give the automation that merges lane A (a GitHub App or bot account) bypass permission in the branch ruleset, and have it merge only pull requests whose paths all match lane A. Either way, the tech lead’s weekly sample is the check on it.

A count of 1 also means nothing enforces lane D’s second approval: CODEOWNERS requires one code-owner approval, not two. Give that second approval its own control. Either add a required CI check that fails when lane D paths changed and the pull request has fewer than two approvals (read them with gh pr view <number> --json reviews), or keep it as policy and audit it in the weekly metrics below, where the Unreviewed merges row also counts lane D merges with fewer than two approvals.

Lane A starts empty. A class of change earns its way into it after a run of merges with no reverts and no escaped defects, and falls back out after one. The trust-transfer protocol describes that staging in detail.

Balance reviewer load instead of assigning by habit

Section titled “Balance reviewer load instead of assigning by habit”

Left alone, review requests flow to the two people who answer fastest, and they become the bottleneck and then the burnout. GitHub’s team code review settings fix the mechanics. Under Team settings › Code review, select Enable auto assignment and choose the Load balance routing algorithm rather than round robin.

GitHub’s documentation says load balance “considers the number of outstanding reviews for each member”. It aims for each member to review an equal number of pull requests in any 30-day period. Round robin, by contrast, alternates “regardless of the number of outstanding reviews they currently have”.

Two more settings matter. Members whose GitHub status is Busy are not selected, so a focus block becomes a real state rather than a polite fiction. Never assign certain team members keeps someone off the rotation during onboarding or an incident week.

Pre-review and routing in Claude Code, Codex, and Cursor

Section titled “Pre-review and routing in Claude Code, Codex, and Cursor”

The queue controls above are forge-level and identical for all three tools. What differs is how each tool reviews a change before a human sees it and how it helps route the result.

  • Before the pull request opens: have the agent run the bundled /code-review skill on its own diff and fix what it finds. Add --fix to apply fixes in the same pass.
  • On the pull request: the managed Code Review service (research preview, Team and Enterprise plans) posts inline findings sorted into Important, Nit and Pre-existing, and averages $15–25 per review according to Anthropic. Its check run “always completes with a neutral conclusion so it never blocks merging” (Code Review docs, checked 2026-09-26), so it informs lane B and cannot be lane A’s gate on its own.
  • Lane D deep review: run claude ultrareview 482 from the shell, where 482 is the pull request number (/code-review ultra inside a session). It runs a multi-agent review in the cloud and independently reproduces each finding; Anthropic lists it at typically $5 to $25 after three free runs on Pro and Max. It is not available on Bedrock, Google Cloud, or Foundry, or to Zero Data Retention organisations.
  • Watching the fleet: claude agents opens agent view (research preview), which lets you “dispatch and manage many Claude Code sessions from one screen”. Use it to see in one place how many sessions your engineers are supervising.

An automated reviewer that cannot block trains people to scroll past it. Decide which findings it may fail the build on, a short list such as secrets, SQL built from strings, and disabled tests, and treat everything else as a comment. The agent PR review workflow covers how to configure that layer, and PR review automation for teams covers rolling it out across a team.

Throughput counts are easy to game and tell you nothing about the queue. Track these eight, weekly, split by lane and by agent versus human author:

MetricDefinitionAct when
Arrival rateNon-draft pull requests opened per working dayIt exceeds review throughput for two weeks running
Review throughputPull requests reaching merge or close after a review, per working dayIt falls while arrivals hold steady
Queue depthOpen non-draft pull requests with no approval yet, agent and human (what the gate script counts)It sits at the WIP cap for more than two days
Time to first reviewOpened (or marked ready) until the first submitted human review, median and 90th percentileThe median exceeds half a working day
Time in reviewOpened (or marked ready) until merge, median and 90th percentileThe 90th percentile exceeds your target by 50%
PR sizeChanged lines excluding lockfiles and snapshots, medianThe median rises two weeks in a row
Unreviewed mergesShare of merges with no submitted human review outside lane A, plus lane D merges with fewer than two approvalsAny merge outside lane A has no review, or any lane D merge has fewer than two approvals
Reviewer loadOpen review requests per person, and the spread between the busiest and the least busyOne person carries twice the team median

The thresholds are starting points for a conversation, not industry benchmarks. The first four weeks give you your baseline; after that, compare against it. For how these fit the DORA and SPACE families, see engineering metrics frameworks.

This command exports the timing, size and changed-path columns for the last 200 merged agent pull requests. Run it in a terminal inside the repository, with gh authenticated:

Terminal window
gh pr list --state merged --label agent --limit 200 \
--json number,author,createdAt,mergedAt,additions,deletions,reviews,files \
--jq '.[] | (.createdAt | fromdateiso8601) as $o | .author.login as $me
| [ .number, .createdAt, .mergedAt,
(.additions + .deletions),
(([.reviews[] | (.author.login // "") as $r
| select($r != $me and ($r | test("\\[bot\\]$|^copilot-|^chatgpt-codex") | not))
| .submittedAt] | min) as $f
| if $f then ((($f | fromdateiso8601) - $o) / 3600 | floor) else "none" end),
(((.mergedAt | fromdateiso8601) - $o) / 3600 | floor),
([.files[].path] | join(",")) ]
| @tsv' > review-queue.tsv

The columns are pull request number, created and merged timestamps, changed lines, hours to first review (none means merged unreviewed), hours in review, and the changed paths, comma-separated. gh fetches at most the first 100 files per pull request, so a larger one lists only part of its paths. The filter drops reviews by bots and by the pull request’s own author, so the first-review column measures human review; the automated reviewers above would otherwise almost always review first. GitHub App reviewers do not always carry the [bot] suffix in this output, so check once with gh pr view <number> --json reviews --jq '.reviews[].author.login' on a pull request your bots reviewed, and add the logins it shows to the pattern.

createdAt includes draft time, so the numbers overstate the queue if agents open drafts first. When that distortion matters, use the ready_for_review event from the timeline API instead. The size column also counts lockfiles, which the CI budget excludes.

Keep a sustainable pace when people supervise parallel agents

Section titled “Keep a sustainable pace when people supervise parallel agents”

The queue has a human cost that the metrics above see late. Anthropic’s study of its own engineers (December 2025) found some reporting that their work shifted “70%+ to being a code reviewer/reviser rather than a net-new code writer”, and it names a “paradox of supervision”: overseeing agents needs the very skills that over-delegation erodes. Faros AI’s earlier 2025 AI Productivity Paradox report (secondary source: search extracts, not the report itself) put PR review time up about 91% on teams with high AI adoption. Supervising several agents at once is tiring in a way no dashboard shows until someone stops reading.

Adopt these four rules as a team policy, and revisit them at each retrospective:

  1. Cap concurrent sessions per engineer at what they can review the same day. Two or three is a reasonable place to start. An agent that finishes at 5 p.m. with nobody to read its output has produced inventory, not progress.

  2. Protect review blocks and focus blocks separately. Schedule review in fixed blocks rather than as interrupts, and have people set their GitHub status to Busy during focus time so load-balanced assignment skips them.

  3. Rotate a daily queue owner. One person per day watches the queue depth, chases pull requests past the time-to-first-review target, and applies nothing but routing. The role rotates so nobody becomes the permanent reviewer, and so juniors see how the whole queue behaves.

  4. Lower concurrency before you lower the bar. When time to first review stays above target for two weeks, when unreviewed merges appear outside lane A, or when two people in the same week say they are only reviewing, reduce the WIP cap and the per-engineer session cap first. Relaxing the lanes to clear the queue is how teams end up looking like the Faros data: more unreviewed merges and more incidents per pull request.

What breaks when you manage the review queue

Section titled “What breaks when you manage the review queue”

The WIP cap moves inventory into branches. Engineers hold agent branches locally until the queue has room. Recover by counting open agent branches as well as pull requests in the weekly review, and by gating dispatch rather than pull request creation, as the script above does.

Agents meet the size budget with splits that cannot be reviewed alone. Four 390-line pull requests that only make sense together are one 1,560-line pull request with extra overhead. Require each stacked pull request to pass the suite on its own and state its single concern; reject the stack when it does not.

Lane A grows by habit. A class of change enters the green-pipeline lane because it was convenient in a busy week. Recover by keeping lane membership in the committed table, reviewed by the tech lead, and by dropping a class back to lane B after any revert or escaped defect.

The metrics punish drafts or reward tiny pull requests. createdAt counts draft time, and a team measured on time in review alone can split changes past the point of usefulness. Read time in review together with PR size, reverts, and escaped defects; never report one of them alone.

The busiest reviewer burns out anyway. Load balancing equalises request counts, not difficulty, and lane D reviews cost far more than lane B. Track who carries lane D and rotate it deliberately.

Automated review becomes background noise. If everyone scrolls past the bot, its useful findings are lost too. Cut its blocking list to the few findings that justify a failed build.

Edit page

Last updated:

Cite this page — https://developertoolkit.ai/en/teams/review-queue/, developertoolkit.ai