Reviewing an agent's pull request without reading every line
Reviewing an agent’s pull request without reading every line is a six-step triage: check the spec delta, confirm the evidence, scan the risk flags, inspect every oracle change, read a small sample of hotspots, then approve, return, or escalate. A review agent runs the mechanical checklist first, and a named human reads code only for the escalation classes.
It is Tuesday morning and nine agent pull requests are waiting for you. The first is 640 lines across 14 files, every check is green, and the description says “Implemented pagination as requested.” You can spend the morning reading it, or skim it and hope. This page gives you a third option: about 15 minutes per pull request, and a record you can defend.
What you get from agent PR triage
Section titled “What you get from agent PR triage”- A six-step triage protocol with a time budget and a stop condition for each step.
- A review-agent checklist you commit to the repository: the former 50-point human checklist, rewritten as checks an agent runs and reports with evidence.
- A human escalation list of the change classes where someone still reads the code, and the question that person answers.
- Three verdicts (approve, return, escalate) with the criteria for each, so two reviewers reach the same decision.
- Three copy-paste prompts and the wiring for Claude Code, Codex and Cursor.
- Four measures that tell you whether the triage itself is working.
This page puts reading evidence instead of code into practice; read that page first if approving code you did not read is still new to you.
Why a 50-point checklist for humans stopped working
Section titled “Why a 50-point checklist for humans stopped working”The old answer was a 50-point checklist applied by a person reading the diff. The checks were right; the reader was the problem.
Pull requests got bigger and slower to review at the same time. DX measured median pull request size growing from 44 to 72 lines between July 2025 and June 2026 (DX, Justin Reock, 17 June 2026). Faros AI’s Acceleration Whiplash report (April 2026, telemetry from 22,000 developers) found median time in review up 441.5% and 31.3% more pull requests merging with no review at all. A checklist that depends on reading every line ends one of two ways: the queue stalls, or the reading quietly stops.
The fix keeps every check and moves it: an agent runs the mechanical checks and reports evidence, and the human reads that report, the oracle changes and a short code sample, then decides. That is the move from Level 3, where you review diffs, to Level 4, where you write the specs and judge the evidence.
How does the six-step triage work?
Section titled “How does the six-step triage work?”Run the steps in order. Each can end the review early, so the expensive steps run only on pull requests that survived the cheap ones.
| Step | The question | What you look at | Budget | Stops the review when |
|---|---|---|---|---|
| 1. Spec delta | Did the agent change what we asked for, and only that? | The ticket and the plain-language list of behaviour changes | 2 min | A behaviour change nobody asked for, or no spec at all → return |
| 2. Evidence | Does every acceptance criterion have a check that ran on this commit? | The acceptance mapping, test output, runtime screenshots or traces | 3 min | Any criterion marked UNVERIFIED, or results from an older commit → return |
| 3. Risk flags | Does the change touch an escalation class, or did the review agent raise a blocker? | The review agent’s report and the list of touched paths | 1 min | An escalation class → escalate; a P0 or P1 finding → return |
| 4. Oracle changes | Did the change alter what “green” means? | The diff of tests, fixtures, snapshots, CI, lint and type config | 3 min | A check made looser without a reason in the spec → return |
| 5. Sampled hotspots | Does the code I did not read hide anything the evidence missed? | Up to three hunks, chosen by rule before you look | 5 min | A real problem → return, and add the missing check |
| 6. Decide | Approve, return or escalate? | Your notes from steps 1–5 | 1 min | — |
The 15-minute budget is our starting policy, not a research finding. If a pull request cannot be triaged in that time, that is information: it is too big, its evidence is thin, or it belongs to an escalation class.
-
Read the spec delta against the ticket. The pull request states, in sentences, which behaviour changed: “
GET /ordersreturns 50 items per page and anext_cursor; the oldpageparameter returns 400.” Compare it with what the ticket asked for. Scope creep, a misread requirement and a silent contract change all show here without opening a file. If the agent cannot explain in a few sentences the architecture it changed, return the pull request. -
Confirm the evidence, line by line. Every acceptance criterion needs one line: criterion, the check that proves it, and its result, for example
old page param returns 400 → orders.contract.test.ts:88 → pass. Check that the results are for the head commit, not an earlier push. For visual or operational changes, look for a screenshot or trace of each acceptance path, including the named error cases. The format of this manifest is defined once, on the evidence bundle. -
Scan the risk flags. The review agent’s report lists the escalation classes touched and its findings by severity. An escalation class routes the pull request to the named code reader; a P0 or P1 finding returns it to the authoring agent.
-
Read every oracle change in full. This is the one part of the diff you always read. A change to a test, fixture, snapshot, CI step, lint rule or
tsconfigchanges what every other check means: an assertion loosened from an exact value to “truthy”, a skipped test, or a regenerated snapshot can make a wrong implementation pass. If the tests are new, ask whether they would fail with the feature removed; how strong your oracle is covers the mechanical answer, mutation testing. -
Read a sample of hotspots, chosen before you look at the code. Pick up to three hunks. Choose the first two by rule, in this order: a hunk the review agent flagged below its confidence cutoff; error handling around external calls, retries or timeouts; concurrency or state transitions; a new dependency or a new abstraction; a hunk the spec delta does not explain; the largest non-test hunk. Choose the third at random (a random changed file, then its largest hunk), so that nothing about the pull request decides what gets read. Read those hunks properly.
-
Decide, and write the verdict on the pull request. Use the verdict table below, and record which hunks you sampled so the trust record shows what a human read.
Pick the random file with a command, not by eye, then read its largest hunk:
# Terminal, on the pull request branchgit diff --name-only main...HEAD -- . ':!*.test.*' ':!*.spec.*' ':!**/__snapshots__/**' | sort -R | head -n 1Approve, return or escalate: which verdict fits?
Section titled “Approve, return or escalate: which verdict fits?”Write the verdict and its reason as a pull request comment; “LGTM” is not a reason.
| Verdict | Criteria (all must hold) | Who acts next |
|---|---|---|
| Approve | The spec delta matches the ticket; every criterion has a passing check on the head commit; no oracle change loosens a check; no escalation class is touched; the sampled hunks raised no question the evidence did not answer | You merge, or auto-merge proceeds; progressive delivery watches production |
| Return | Any of: an unasked behaviour change, an UNVERIFIED criterion, stale results, a loosened check, a P0 or P1 finding, a problem in a sampled hunk, or a pull request over the team’s size budget | The authoring agent, with the failing items as its next prompt. Cap it at two rounds, then a human takes over |
| Escalate | Any of: an escalation class is touched; the review agent and the evidence disagree; the agent could not run a check it cites; the change makes a design decision the spec did not settle | The named code reader for that class, through CODEOWNERS |
The two-round cap on returns is borrowed from practice. At Stripe, the Minions agents are bounded to “at most two rounds of CI”, and they produce “over 1,300 Stripe pull requests” merged each week (Stripe engineering blog, Alistair Gray, Minions Parts 1 and 2, 9 and 19 February 2026). That is one company’s internal figure, not a benchmark. A pull request that is still failing after two returns has a spec problem, not a code problem. The bounded review-fix loop shows how to automate the return path with that stop condition.
When a sampled hunk finds a real problem, also write the check that would have caught it, so the next pull request with the same flaw fails before anyone samples it.
The review-agent checklist
Section titled “The review-agent checklist”The checklist below is the former 50-point human checklist, kept item for item and rewritten so every item asks an agent for evidence (a command and its exit status, a file:line, a search result) rather than an opinion. Commit it as .github/review/agent-checklist.md, and make the tech lead its code owner, so the review standard changes only through a reviewed pull request.
# Review-agent checklist
Review the changes on this branch against main. Do not edit any file.Run checks; never infer their results. Report only what could matter inproduction. No style comments unless the style hides a bug.
Mark every item PASS, FAIL or N/A with one line of evidence: a command and itsexit status, a file:line, or a search result. For every FAIL give severity(P0 blocks merge, P1 fix before merge, P2 fix soon, P3 optional), file:line,why it matters in production, and the smallest safe fix.
## 1. Scope and intent- Link the ticket or spec. If there is none, FAIL and stop.- List every behaviour change as a sentence. Flag any the ticket did not ask for.- List files the change did not need, and any refactor mixed into feature work.- List what the change intentionally does not do, and every assumption made.- Explain in three sentences the architecture this change touches. If you cannot, FAIL.- Check names, copy, comments and examples against the product's domain terms.
## 2. Context fit- For each new piece, name the existing local pattern it follows (file:line). FAIL a new architecture where a local pattern exists.- Search for an existing utility, type, hook, component or service this duplicates.- For each new abstraction, show the current call sites it deduplicates.- Check public interfaces stay backward-compatible unless the PR declares a break.- Check naming, error handling, logging, telemetry and analytics against the surrounding module and its existing helpers.- Check module boundaries and the CODEOWNERS owner of every touched path.
## 3. Correctness- Per changed function, list which of these a test covers: empty input, null/undefined, limits, duplicates, timeouts. Name the missing ones.- For each external call, state what happens on partial failure.- Find un-awaited async work, and fire-and-forget work without lifecycle handling.- Check state transitions are explicit and cannot skip a required step.- Check date, currency, locale and timezone logic is deterministic.- Check retries, idempotency and duplicate events wherever messages, webhooks or payments are handled.- Compare output shapes with the existing API contracts and types.- Say whether tests use realistic data or only toy fixtures.
## 4. Tests (the oracle)- List every changed test, fixture, snapshot, CI, lint and type-config file, each marked stricter, looser or neutral.- For each new test, name the bug it catches. FAIL a test that passes with the feature removed.- Confirm at least one failure-path test per changed behaviour.- Flag mocks inside the logic under test instead of at system boundaries.- Flag regenerated snapshots, sleeps, fixed waits and live network calls.- Run the project's test, type-check and lint commands. Paste each command and its exit status, and the commit SHA they ran on.
## 5. Security and privacy- Validate every new input at the boundary.- Check authorization separately from authentication on every new route or action.- Trace every user-controlled value into SQL, shell, file paths, HTML and URLs.- Check no secret is logged, returned, committed or shipped in the client bundle.- List any widened permission, OAuth scope, CORS, CSP or webhook trust, with its reason.- Check sensitive data is redacted in logs, analytics and error messages.- Check uploads, redirects and callbacks are limited to expected origins and types.- For every new dependency, confirm it exists on the registry under the intended name and is maintained.
## 6. Data and migrations- Check schema changes stay compatible during a rolling deploy.- Check migrations are idempotent or ship a rollback, and preserve existing data.- Check new queries use an index or a bounded scan where volume matters.- Check jobs and webhooks tolerate duplicate delivery.- Check deletes are soft, recoverable, or justified in the PR description.
## 7. Performance and operations- Flag unbounded loops, N+1 queries and repeated network calls.- Check expensive work is cached, batched, queued or paginated.- Report client bundle growth from new imports, if the build reports sizes.- Check errors carry enough context to debug without private data.- Check new critical paths have monitoring, alerting or analytics.- Check security-sensitive paths fail closed and UX paths fail gracefully.
## 8. Reviewability- Report changed lines excluding generated files. FAIL above 400 lines.- Flag formatting-only churn and generated code that was not simplified.- Flag comments that restate the code instead of explaining a decision.- Check screenshots or traces are attached for visual or operational changes.- Name the single riskiest hunk and why, in two sentences.
## 9. Escalation classesList every touched path in these classes: auth, permissions, tenancy, admin;billing, payments, refunds, pricing; schema and migrations; public API, SDK andwebhook contracts; privacy, deletion, export, consent; rate limits, abuseprevention, security headers; tests, CI, lint and type config.
## Output, in this order1. Findings, P0 first.2. Checklist results: PASS, FAIL or N/A with evidence, per item.3. Escalation classes touched, with paths.4. Suggested verdict: APPROVE, RETURN or ESCALATE, and the reason.The 400-line budget in section 8 is a starting value; set it to the size budget your team agreed in the review queue. The suggested verdict is a suggestion: the review agent never approves or merges.
Which changes does a human still read?
Section titled “Which changes does a human still read?”Section 9 of the checklist finds these classes. For each, a named person also reads the code, because the oracle is weak, feedback is slow, or the damage is hard to undo.
| Escalation class | Why evidence is not enough | The question the code reader answers |
|---|---|---|
| Authentication, authorization, permissions, tenancy, admin features | Tests prove the allowed paths; the failure is a path nobody tested | Can any caller reach data or actions it should not? |
| Money: billing, payments, refunds, pricing | Rounding and idempotency errors surface weeks later, in customer accounts | Is every amount computed once, rounded correctly, and safe to retry? |
| Schema and migrations that modify production data | Often irreversible, and run against data the fixtures do not resemble | Can it run twice, and how do we roll it back? |
| Public API, SDK and webhook contracts | The break shows up in someone else’s code, outside this repository’s tests | Which existing client breaks, and was that declared? |
| Privacy and compliance: deletion, export, consent, retention | The failure is a legal exposure, not a failing test | Does personal data go only where the policy says? |
| Abuse and incident controls: rate limits, abuse prevention, security headers | They matter under attack, which tests rarely simulate | Does it fail closed when someone pushes on it? |
| The oracle: tests, CI, lint and type config | It changes what “green” means for every other change | Does green still mean what it meant yesterday? |
| Anything no check covers, or where evidence and review agent disagree | With no oracle, the only evidence is the code | What would a test for this look like, and who writes it? |
Five rows match the escalation table on reading evidence instead of code; the other three come from the old human checklist. Keep a class only if it maps to a small set of paths; an escalation list that covers half the repository sends you back to Level 3. Enforce the list with CODEOWNERS plus the branch protection rule “Require review from Code Owners”:
/src/auth/ @acme/security-reviewers/src/billing/ @acme/payments-owners/db/migrations/ @acme/data-owners/tests/ @acme/tech-leads/.github/ @acme/tech-leadsHow do you run the review agent in Claude Code, Codex and Cursor?
Section titled “How do you run the review agent in Claude Code, Codex and Cursor?”The checklist and the triage are the same in all three tools. What differs is where the review agent runs and whether it can post to the pull request.
Run the checklist headless from your terminal or CI. claude -p sessions start in Manual permission mode, so tools outside the allowlist are denied unless the branch’s own project settings allow them. Section 4 of the checklist runs the test, type-check and lint commands and records the commit SHA, so all of them must be on the allowlist; replace the npm spellings with your project’s own:
# Terminal or CI, from the repository root (Claude Code 2.1.283)claude -p "$(cat .github/review/agent-checklist.md)" \ --allowedTools "Read,Grep,Glob,Bash(git diff *),Bash(git log *),Bash(git rev-parse *),Bash(npm test *),Bash(npm run typecheck *),Bash(npm run lint *)" \ --output-format json --max-budget-usd 2 > review.jsonFor a second pass on correctness, use the bundled /code-review in a session. For escalation-class pull requests, claude ultrareview 482 runs a cloud-hosted multi-agent review of pull request 482, with each finding independently reproduced. After three free runs on Pro and Max it bills usage credits (typically $5 to $25 per run); it is not available on Bedrock, Google Cloud or Foundry, or to Zero Data Retention organizations.
For step 4, the Anthropic plugin pr-review-toolkit focuses on tests and silent failures:
claude plugin install pr-review-toolkit@claude-plugins-officialThen, in a session on the branch: /pr-review-toolkit:review-pr tests errors. The managed Code Review service (research preview, Team and Enterprise) reads REVIEW.md, so copy the checklist there. Its check run always completes as neutral and never blocks merging: treat it as comments, not a gate.
codex exec review reviews the repository non-interactively and takes custom instructions as its prompt; - reads them from standard input. The checklist’s first line, “Review the changes on this branch against main”, sets the scope:
# Terminal or CI, from the repository root (Codex CLI 0.157.1)codex exec review - -c sandbox_mode="workspace-write" -o review.md < .github/review/agent-checklist.mdSection 4 of the checklist runs tests, and many test runners write caches or coverage files, which fail under a read-only sandbox; hence sandbox_mode="workspace-write" (exec review has no -s flag in Codex CLI 0.157.1). In CI, treat the branch as untrusted: keep secrets out of the job’s environment, and for forks let an unprivileged job run the tests while the agent only cites the results.
--base, --commit and --uncommitted are presets and cannot be combined with custom instructions (checked in Codex CLI 0.157.1). Add --output-schema verdict.schema.json when a CI step needs the verdict as JSON. On GitHub and GitLab, comment @codex review on a pull request; review rules live in AGENTS.md (checked 2026-08-28), so add one line there, “Apply .github/review/agent-checklist.md to every review”, and both paths use the same standard.
Open the agent on the pull request branch and paste the first prompt below; it tells the agent to read the checklist file, so no rule configuration is needed. Post the agent’s report as the first comment on the pull request.
On the pull request, Bugbot reviews for bugs and security issues, and PR Routing & Approval assigns reviewers by code ownership and can approve low-risk pull requests that meet your criteria (checked on cursor.com, 2026-08-28). Make the escalation table your approval criteria: auto-approval is allowed only when no escalation class is touched. See PR Routing & Approval in Cursor.
Run the review agent in a fresh session, not the one that wrote the code, which carries its assumptions into the review. For how review bots from each vendor and from third parties compare on noise and configuration, see AI code review bots compared. To split the review into separate passes for correctness, security, tests and spec compliance before human sign-off, see governing layered pull-request review.
Copy-paste prompts for agent PR review
Section titled “Copy-paste prompts for agent PR review”The third prompt is for round one and round two only. After the second return, a human reads the pull request or rewrites the spec.
How do you know the triage is working?
Section titled “How do you know the triage is working?”The triage needs evidence of its own. Track four numbers per change class, monthly:
| Measure | Definition | What a bad trend tells you |
|---|---|---|
| Sample yield | Share of sampled hunks where the human found a real problem the evidence and the review agent missed | Rising: the evidence has a gap. Write the missing check, then watch the yield fall |
| Escape rate by verdict | Production defects traced to pull requests approved through triage, divided by all triage approvals | Rising: a class is being approved too early; move it back to sampled or full reading |
| Return rounds | Average returns per pull request before approval | Above two: specs are unclear, not code quality |
| Review-agent precision | Findings the human accepted, divided by findings raised | Falling: the checklist produces noise; tighten it before people stop reading the report |
Sign-off stays explicit: the approver’s name is on every approved pull request, the CODEOWNERS owners sign off on escalation classes, and the tech lead changes the checklist and the escalation list only through reviewed pull requests. Canonical versions of these measures sit with the other team metrics in metrics frameworks for agentic engineering, and the staged rollout that moves a whole team to this protocol is helping the team stop reading every diff.
What breaks when you triage agent pull requests?
Section titled “What breaks when you triage agent pull requests?”The review agent buries you in nits. Forty findings, most about naming, and the one real bug is 31st. Recovery: keep the “no style comments” rule, ask only for P0 to P2, and when precision falls, remove the checklist items that produce the noise.
The author and the reviewer share the same blind spot. The same model, in the same session, with the same context misreads the spec twice. Recovery: run the review in a fresh session, and for escalation classes use a different tool or a deeper pass (claude ultrareview, @codex review, Bugbot). The review agent’s opinion is not evidence; only checks that ran are.
The results are for a different commit. The agent pushed a fix after running the tests, so the green checks belong to the previous commit. Recovery: require the commit SHA next to every command in the report, and make CI, not the agent, the source of the final test results.
Sampling drifts into skimming, or stops. Under pressure, “three hunks” becomes “I glanced at it”. Recovery: record the sampled hunks in the verdict comment, pick the random file with the command, not by eye, and count pull requests with no recorded sample as unsampled.
Returned pull requests loop. The authoring agent fixes one finding and breaks another. Recovery: stop at two returns. Rewrite the spec or its acceptance criteria before a third attempt; see writing acceptance criteria an agent cannot misread.
The escalation list swallows everything. Every pull request touches “shared” code, so every pull request escalates, and you are back to reading diffs. Recovery: escalate by path through CODEOWNERS, not by judgment, and add a class only through a reviewed change to that file.
A plausible dependency does not exist. The code imports a package name that looks right and was never published, or was published by an attacker. Recovery: keep the dependency item in section 5 of the checklist, and add the registry check from dependency verification for agent changes to CI.