Skip to content

Reviewing an agent's pull request without reading every line

Reviewing an agent’s pull request without reading every line is a six-step triage: check the spec delta, confirm the evidence, scan the risk flags, inspect every oracle change, read a small sample of hotspots, then approve, return, or escalate. A review agent runs the mechanical checklist first, and a named human reads code only for the escalation classes.

It is Tuesday morning and nine agent pull requests are waiting for you. The first is 640 lines across 14 files, every check is green, and the description says “Implemented pagination as requested.” You can spend the morning reading it, or skim it and hope. This page gives you a third option: about 15 minutes per pull request, and a record you can defend.

  • A six-step triage protocol with a time budget and a stop condition for each step.
  • A review-agent checklist you commit to the repository: the former 50-point human checklist, rewritten as checks an agent runs and reports with evidence.
  • A human escalation list of the change classes where someone still reads the code, and the question that person answers.
  • Three verdicts (approve, return, escalate) with the criteria for each, so two reviewers reach the same decision.
  • Three copy-paste prompts and the wiring for Claude Code, Codex and Cursor.
  • Four measures that tell you whether the triage itself is working.

This page puts reading evidence instead of code into practice; read that page first if approving code you did not read is still new to you.

Why a 50-point checklist for humans stopped working

Section titled “Why a 50-point checklist for humans stopped working”

The old answer was a 50-point checklist applied by a person reading the diff. The checks were right; the reader was the problem.

Pull requests got bigger and slower to review at the same time. DX measured median pull request size growing from 44 to 72 lines between July 2025 and June 2026 (DX, Justin Reock, 17 June 2026). Faros AI’s Acceleration Whiplash report (April 2026, telemetry from 22,000 developers) found median time in review up 441.5% and 31.3% more pull requests merging with no review at all. A checklist that depends on reading every line ends one of two ways: the queue stalls, or the reading quietly stops.

The fix keeps every check and moves it: an agent runs the mechanical checks and reports evidence, and the human reads that report, the oracle changes and a short code sample, then decides. That is the move from Level 3, where you review diffs, to Level 4, where you write the specs and judge the evidence.

Run the steps in order. Each can end the review early, so the expensive steps run only on pull requests that survived the cheap ones.

StepThe questionWhat you look atBudgetStops the review when
1. Spec deltaDid the agent change what we asked for, and only that?The ticket and the plain-language list of behaviour changes2 minA behaviour change nobody asked for, or no spec at all → return
2. EvidenceDoes every acceptance criterion have a check that ran on this commit?The acceptance mapping, test output, runtime screenshots or traces3 minAny criterion marked UNVERIFIED, or results from an older commit → return
3. Risk flagsDoes the change touch an escalation class, or did the review agent raise a blocker?The review agent’s report and the list of touched paths1 minAn escalation class → escalate; a P0 or P1 finding → return
4. Oracle changesDid the change alter what “green” means?The diff of tests, fixtures, snapshots, CI, lint and type config3 minA check made looser without a reason in the spec → return
5. Sampled hotspotsDoes the code I did not read hide anything the evidence missed?Up to three hunks, chosen by rule before you look5 minA real problem → return, and add the missing check
6. DecideApprove, return or escalate?Your notes from steps 1–51 min—

The 15-minute budget is our starting policy, not a research finding. If a pull request cannot be triaged in that time, that is information: it is too big, its evidence is thin, or it belongs to an escalation class.

  1. Read the spec delta against the ticket. The pull request states, in sentences, which behaviour changed: “GET /orders returns 50 items per page and a next_cursor; the old page parameter returns 400.” Compare it with what the ticket asked for. Scope creep, a misread requirement and a silent contract change all show here without opening a file. If the agent cannot explain in a few sentences the architecture it changed, return the pull request.

  2. Confirm the evidence, line by line. Every acceptance criterion needs one line: criterion, the check that proves it, and its result, for example old page param returns 400 → orders.contract.test.ts:88 → pass. Check that the results are for the head commit, not an earlier push. For visual or operational changes, look for a screenshot or trace of each acceptance path, including the named error cases. The format of this manifest is defined once, on the evidence bundle.

  3. Scan the risk flags. The review agent’s report lists the escalation classes touched and its findings by severity. An escalation class routes the pull request to the named code reader; a P0 or P1 finding returns it to the authoring agent.

  4. Read every oracle change in full. This is the one part of the diff you always read. A change to a test, fixture, snapshot, CI step, lint rule or tsconfig changes what every other check means: an assertion loosened from an exact value to “truthy”, a skipped test, or a regenerated snapshot can make a wrong implementation pass. If the tests are new, ask whether they would fail with the feature removed; how strong your oracle is covers the mechanical answer, mutation testing.

  5. Read a sample of hotspots, chosen before you look at the code. Pick up to three hunks. Choose the first two by rule, in this order: a hunk the review agent flagged below its confidence cutoff; error handling around external calls, retries or timeouts; concurrency or state transitions; a new dependency or a new abstraction; a hunk the spec delta does not explain; the largest non-test hunk. Choose the third at random (a random changed file, then its largest hunk), so that nothing about the pull request decides what gets read. Read those hunks properly.

  6. Decide, and write the verdict on the pull request. Use the verdict table below, and record which hunks you sampled so the trust record shows what a human read.

Pick the random file with a command, not by eye, then read its largest hunk:

Terminal window
# Terminal, on the pull request branch
git diff --name-only main...HEAD -- . ':!*.test.*' ':!*.spec.*' ':!**/__snapshots__/**' | sort -R | head -n 1

Approve, return or escalate: which verdict fits?

Section titled “Approve, return or escalate: which verdict fits?”

Write the verdict and its reason as a pull request comment; “LGTM” is not a reason.

VerdictCriteria (all must hold)Who acts next
ApproveThe spec delta matches the ticket; every criterion has a passing check on the head commit; no oracle change loosens a check; no escalation class is touched; the sampled hunks raised no question the evidence did not answerYou merge, or auto-merge proceeds; progressive delivery watches production
ReturnAny of: an unasked behaviour change, an UNVERIFIED criterion, stale results, a loosened check, a P0 or P1 finding, a problem in a sampled hunk, or a pull request over the team’s size budgetThe authoring agent, with the failing items as its next prompt. Cap it at two rounds, then a human takes over
EscalateAny of: an escalation class is touched; the review agent and the evidence disagree; the agent could not run a check it cites; the change makes a design decision the spec did not settleThe named code reader for that class, through CODEOWNERS

The two-round cap on returns is borrowed from practice. At Stripe, the Minions agents are bounded to “at most two rounds of CI”, and they produce “over 1,300 Stripe pull requests” merged each week (Stripe engineering blog, Alistair Gray, Minions Parts 1 and 2, 9 and 19 February 2026). That is one company’s internal figure, not a benchmark. A pull request that is still failing after two returns has a spec problem, not a code problem. The bounded review-fix loop shows how to automate the return path with that stop condition.

When a sampled hunk finds a real problem, also write the check that would have caught it, so the next pull request with the same flaw fails before anyone samples it.

The checklist below is the former 50-point human checklist, kept item for item and rewritten so every item asks an agent for evidence (a command and its exit status, a file:line, a search result) rather than an opinion. Commit it as .github/review/agent-checklist.md, and make the tech lead its code owner, so the review standard changes only through a reviewed pull request.

.github/review/agent-checklist.md
# Review-agent checklist
Review the changes on this branch against main. Do not edit any file.
Run checks; never infer their results. Report only what could matter in
production. No style comments unless the style hides a bug.
Mark every item PASS, FAIL or N/A with one line of evidence: a command and its
exit status, a file:line, or a search result. For every FAIL give severity
(P0 blocks merge, P1 fix before merge, P2 fix soon, P3 optional), file:line,
why it matters in production, and the smallest safe fix.
## 1. Scope and intent
- Link the ticket or spec. If there is none, FAIL and stop.
- List every behaviour change as a sentence. Flag any the ticket did not ask for.
- List files the change did not need, and any refactor mixed into feature work.
- List what the change intentionally does not do, and every assumption made.
- Explain in three sentences the architecture this change touches. If you cannot, FAIL.
- Check names, copy, comments and examples against the product's domain terms.
## 2. Context fit
- For each new piece, name the existing local pattern it follows (file:line).
FAIL a new architecture where a local pattern exists.
- Search for an existing utility, type, hook, component or service this duplicates.
- For each new abstraction, show the current call sites it deduplicates.
- Check public interfaces stay backward-compatible unless the PR declares a break.
- Check naming, error handling, logging, telemetry and analytics against the
surrounding module and its existing helpers.
- Check module boundaries and the CODEOWNERS owner of every touched path.
## 3. Correctness
- Per changed function, list which of these a test covers: empty input,
null/undefined, limits, duplicates, timeouts. Name the missing ones.
- For each external call, state what happens on partial failure.
- Find un-awaited async work, and fire-and-forget work without lifecycle handling.
- Check state transitions are explicit and cannot skip a required step.
- Check date, currency, locale and timezone logic is deterministic.
- Check retries, idempotency and duplicate events wherever messages, webhooks or
payments are handled.
- Compare output shapes with the existing API contracts and types.
- Say whether tests use realistic data or only toy fixtures.
## 4. Tests (the oracle)
- List every changed test, fixture, snapshot, CI, lint and type-config file,
each marked stricter, looser or neutral.
- For each new test, name the bug it catches. FAIL a test that passes with the
feature removed.
- Confirm at least one failure-path test per changed behaviour.
- Flag mocks inside the logic under test instead of at system boundaries.
- Flag regenerated snapshots, sleeps, fixed waits and live network calls.
- Run the project's test, type-check and lint commands. Paste each command and
its exit status, and the commit SHA they ran on.
## 5. Security and privacy
- Validate every new input at the boundary.
- Check authorization separately from authentication on every new route or action.
- Trace every user-controlled value into SQL, shell, file paths, HTML and URLs.
- Check no secret is logged, returned, committed or shipped in the client bundle.
- List any widened permission, OAuth scope, CORS, CSP or webhook trust, with its reason.
- Check sensitive data is redacted in logs, analytics and error messages.
- Check uploads, redirects and callbacks are limited to expected origins and types.
- For every new dependency, confirm it exists on the registry under the intended
name and is maintained.
## 6. Data and migrations
- Check schema changes stay compatible during a rolling deploy.
- Check migrations are idempotent or ship a rollback, and preserve existing data.
- Check new queries use an index or a bounded scan where volume matters.
- Check jobs and webhooks tolerate duplicate delivery.
- Check deletes are soft, recoverable, or justified in the PR description.
## 7. Performance and operations
- Flag unbounded loops, N+1 queries and repeated network calls.
- Check expensive work is cached, batched, queued or paginated.
- Report client bundle growth from new imports, if the build reports sizes.
- Check errors carry enough context to debug without private data.
- Check new critical paths have monitoring, alerting or analytics.
- Check security-sensitive paths fail closed and UX paths fail gracefully.
## 8. Reviewability
- Report changed lines excluding generated files. FAIL above 400 lines.
- Flag formatting-only churn and generated code that was not simplified.
- Flag comments that restate the code instead of explaining a decision.
- Check screenshots or traces are attached for visual or operational changes.
- Name the single riskiest hunk and why, in two sentences.
## 9. Escalation classes
List every touched path in these classes: auth, permissions, tenancy, admin;
billing, payments, refunds, pricing; schema and migrations; public API, SDK and
webhook contracts; privacy, deletion, export, consent; rate limits, abuse
prevention, security headers; tests, CI, lint and type config.
## Output, in this order
1. Findings, P0 first.
2. Checklist results: PASS, FAIL or N/A with evidence, per item.
3. Escalation classes touched, with paths.
4. Suggested verdict: APPROVE, RETURN or ESCALATE, and the reason.

The 400-line budget in section 8 is a starting value; set it to the size budget your team agreed in the review queue. The suggested verdict is a suggestion: the review agent never approves or merges.

Section 9 of the checklist finds these classes. For each, a named person also reads the code, because the oracle is weak, feedback is slow, or the damage is hard to undo.

Escalation classWhy evidence is not enoughThe question the code reader answers
Authentication, authorization, permissions, tenancy, admin featuresTests prove the allowed paths; the failure is a path nobody testedCan any caller reach data or actions it should not?
Money: billing, payments, refunds, pricingRounding and idempotency errors surface weeks later, in customer accountsIs every amount computed once, rounded correctly, and safe to retry?
Schema and migrations that modify production dataOften irreversible, and run against data the fixtures do not resembleCan it run twice, and how do we roll it back?
Public API, SDK and webhook contractsThe break shows up in someone else’s code, outside this repository’s testsWhich existing client breaks, and was that declared?
Privacy and compliance: deletion, export, consent, retentionThe failure is a legal exposure, not a failing testDoes personal data go only where the policy says?
Abuse and incident controls: rate limits, abuse prevention, security headersThey matter under attack, which tests rarely simulateDoes it fail closed when someone pushes on it?
The oracle: tests, CI, lint and type configIt changes what “green” means for every other changeDoes green still mean what it meant yesterday?
Anything no check covers, or where evidence and review agent disagreeWith no oracle, the only evidence is the codeWhat would a test for this look like, and who writes it?

Five rows match the escalation table on reading evidence instead of code; the other three come from the old human checklist. Keep a class only if it maps to a small set of paths; an escalation list that covers half the repository sends you back to Level 3. Enforce the list with CODEOWNERS plus the branch protection rule “Require review from Code Owners”:

.github/CODEOWNERS
/src/auth/ @acme/security-reviewers
/src/billing/ @acme/payments-owners
/db/migrations/ @acme/data-owners
/tests/ @acme/tech-leads
/.github/ @acme/tech-leads

How do you run the review agent in Claude Code, Codex and Cursor?

Section titled “How do you run the review agent in Claude Code, Codex and Cursor?”

The checklist and the triage are the same in all three tools. What differs is where the review agent runs and whether it can post to the pull request.

Run the checklist headless from your terminal or CI. claude -p sessions start in Manual permission mode, so tools outside the allowlist are denied unless the branch’s own project settings allow them. Section 4 of the checklist runs the test, type-check and lint commands and records the commit SHA, so all of them must be on the allowlist; replace the npm spellings with your project’s own:

Terminal window
# Terminal or CI, from the repository root (Claude Code 2.1.283)
claude -p "$(cat .github/review/agent-checklist.md)" \
--allowedTools "Read,Grep,Glob,Bash(git diff *),Bash(git log *),Bash(git rev-parse *),Bash(npm test *),Bash(npm run typecheck *),Bash(npm run lint *)" \
--output-format json --max-budget-usd 2 > review.json

For a second pass on correctness, use the bundled /code-review in a session. For escalation-class pull requests, claude ultrareview 482 runs a cloud-hosted multi-agent review of pull request 482, with each finding independently reproduced. After three free runs on Pro and Max it bills usage credits (typically $5 to $25 per run); it is not available on Bedrock, Google Cloud or Foundry, or to Zero Data Retention organizations.

For step 4, the Anthropic plugin pr-review-toolkit focuses on tests and silent failures:

Terminal window
claude plugin install pr-review-toolkit@claude-plugins-official

Then, in a session on the branch: /pr-review-toolkit:review-pr tests errors. The managed Code Review service (research preview, Team and Enterprise) reads REVIEW.md, so copy the checklist there. Its check run always completes as neutral and never blocks merging: treat it as comments, not a gate.

Run the review agent in a fresh session, not the one that wrote the code, which carries its assumptions into the review. For how review bots from each vendor and from third parties compare on noise and configuration, see AI code review bots compared. To split the review into separate passes for correctness, security, tests and spec compliance before human sign-off, see governing layered pull-request review.

The third prompt is for round one and round two only. After the second return, a human reads the pull request or rewrites the spec.

The triage needs evidence of its own. Track four numbers per change class, monthly:

MeasureDefinitionWhat a bad trend tells you
Sample yieldShare of sampled hunks where the human found a real problem the evidence and the review agent missedRising: the evidence has a gap. Write the missing check, then watch the yield fall
Escape rate by verdictProduction defects traced to pull requests approved through triage, divided by all triage approvalsRising: a class is being approved too early; move it back to sampled or full reading
Return roundsAverage returns per pull request before approvalAbove two: specs are unclear, not code quality
Review-agent precisionFindings the human accepted, divided by findings raisedFalling: the checklist produces noise; tighten it before people stop reading the report

Sign-off stays explicit: the approver’s name is on every approved pull request, the CODEOWNERS owners sign off on escalation classes, and the tech lead changes the checklist and the escalation list only through reviewed pull requests. Canonical versions of these measures sit with the other team metrics in metrics frameworks for agentic engineering, and the staged rollout that moves a whole team to this protocol is helping the team stop reading every diff.

What breaks when you triage agent pull requests?

Section titled “What breaks when you triage agent pull requests?”

The review agent buries you in nits. Forty findings, most about naming, and the one real bug is 31st. Recovery: keep the “no style comments” rule, ask only for P0 to P2, and when precision falls, remove the checklist items that produce the noise.

The author and the reviewer share the same blind spot. The same model, in the same session, with the same context misreads the spec twice. Recovery: run the review in a fresh session, and for escalation classes use a different tool or a deeper pass (claude ultrareview, @codex review, Bugbot). The review agent’s opinion is not evidence; only checks that ran are.

The results are for a different commit. The agent pushed a fix after running the tests, so the green checks belong to the previous commit. Recovery: require the commit SHA next to every command in the report, and make CI, not the agent, the source of the final test results.

Sampling drifts into skimming, or stops. Under pressure, “three hunks” becomes “I glanced at it”. Recovery: record the sampled hunks in the verdict comment, pick the random file with the command, not by eye, and count pull requests with no recorded sample as unsampled.

Returned pull requests loop. The authoring agent fixes one finding and breaks another. Recovery: stop at two returns. Rewrite the spec or its acceptance criteria before a third attempt; see writing acceptance criteria an agent cannot misread.

The escalation list swallows everything. Every pull request touches “shared” code, so every pull request escalates, and you are back to reading diffs. Recovery: escalate by path through CODEOWNERS, not by judgment, and add a class only through a reviewed change to that file.

A plausible dependency does not exist. The code imports a package name that looks right and was never published, or was published by an attacker. Recovery: keep the dependency item in section 5 of the checklist, and add the registry check from dependency verification for agent changes to CI.