Skip to content

Team PR review automation policy — evidence before approval

A team PR review automation policy is the organization-wide rule set that decides what review agents may do on a pull request: which deterministic checks are authoritative, which risk lane a change takes, what makes an AI finding blocking, who may approve, and what a failed reviewer means. Its core rule: no pull request merges on an AI opinion alone.

After a quarter of experiments, your organization runs three review bots. The platform team enabled Claude Code Review, the web team turned on Bugbot, and someone added @codex review to a few repositories. One pull request collects 41 comments that nobody reads. Another merges after a bot said “no issues found”, and it carried a migration that no person looked at. Then your security lead asks which of these bots is allowed to approve anything, and nobody can answer.

This page is for the CTO or VP Engineering who owns that answer and the tech leads who enforce it. It covers the policy only. Configuring each review agent is in layered pull-request review, comparing the bots is in AI code review bots compared, and running the daily queue is in the review queue.

Q9 · Quality gates Max-score evidence: deterministic checks, a focused specification and risk review, reproducible findings, risk-based routing, and a named human merge gate.

  • A one-page review-policy.md template you can adopt as-is and adjust per repository.
  • A risk-lane table that tells every repository which evidence, which reviewers, and which approver a change needs.
  • Enforcement in the forge (required checks and code owners), so the policy holds even when a bot is down.
  • Six metric definitions that show whether each AI reviewer is worth its cost.
  • Copy-paste prompts that inventory your current review bots and audit merged pull requests against the policy.

Why review automation needs a policy, not more bots

Section titled “Why review automation needs a policy, not more bots”

Agents raise the number and size of pull requests faster than review capacity grows. Faros AI’s “AI Engineering Report 2026” (April 2026; telemetry from 22,000 developers on Faros’s own platform, so a self-selected customer base) measured pull request size up 51% and pull requests merged per developer up 16.2%, while median time in review rose 441.5%, pull requests merged without any review rose 31.3%, and incidents per pull request rose 242.7%.

Adding reviewers does not fix that on its own. Each bot adds comments, and a comment is not evidence. Without a written policy, teams drift into one of two failure states: they approve without reading because the queue is too long, or they treat a bot’s silence as approval. The policy exists to stop both. It says which signals are facts, which are opinions that need proof, and which person carries the decision.

What must a review automation policy decide?

Section titled “What must a review automation policy decide?”

Settle these eight decisions once, at the organization level. Repositories may tighten a rule; they do not loosen one without an exception on file.

DecisionRecommended defaultOwner
Authoritative checksFormatting, types, lint, unit and integration tests, dependency audit, secret scan. Pass/fail, required in the forge.Platform team
Risk lanesThree lanes assigned by changed paths, never by the author’s own label.CTO, with security
AI reviewers per laneAt most one general reviewer, plus one security pass in the sensitive lane. Add a second only when metrics justify it.Tech leads
What an AI finding may blockOnly a finding with a reproducer, a failing test, or a cited policy clause. Everything else is advisory.Tech leads
Approval authorityA named human approves lanes 1 and 2. Machine approval, if allowed at all, only in lane 0.CTO
Reviewer failureA reviewer that errors, times out, or is not configured reports “evidence incomplete”. It never counts as a pass.Platform team
Data boundariesWhich repositories may send code to which review vendor, and under which retention terms.Security and legal
Measurement and reviewThe six metrics below, reviewed monthly; a reviewer below its precision floor for two months is retuned or removed.Engineering leadership

Lanes decide how much evidence a change needs before a person may approve it. Classify by changed paths through CODEOWNERS, so a change cannot pick its own lane.

LaneWhat it containsAutomated evidenceHuman gateMachine approval
0 · Low riskDocumentation, test-only changes, generated files with a generator checkAuthoritative checksAny team member, or none if the policy allows itOptional
1 · StandardApplication code outside lane 2Authoritative checks, one AI review against spec.md and the diffOne approver from the owning teamNo
2 · SensitiveAuthentication, authorization, payments, personal data, migrations, public API contracts, infrastructure, CI workflows, agent configuration (CLAUDE.md, AGENTS.md, .claude/, .codex/, .cursor/, MCP config), and review governance (CODEOWNERS, REVIEW.md, the policy file)Lane 1 evidence, a security review pass, and a rollback note in the PRA code owner for the affected system, by nameNo

Agent configuration sits in lane 2 on purpose. A change to AGENTS.md, a hook, or an MCP server entry changes how every later agent run behaves, including the review agents themselves, so it needs the same scrutiny as a CI workflow. Nested instruction files count as agent configuration too: Claude Code loads a CLAUDE.md from subdirectories, and Codex reads AGENTS.md per directory, so src/payments/AGENTS.md steers the review of the payments code. See shared hooks governance for the rules on hook changes.

Review governance sits in lane 2 for the same reason. CODEOWNERS decides every lane, REVIEW.md tells the review agents what to look for, and the policy file sets the rules both must satisfy. Left to the catch-all, they would be lane 1 files, and one approval from the default team could rewrite the lanes themselves.

Copy this file into a central repository (for example your engineering handbook) and link it from every repository’s REVIEW.md. The per-repository REVIEW.md implements the policy for its review agents; this file is what it must satisfy.

review-policy.md
# Pull request review policy (version 1.0, owner: VP Engineering)
## 1. Authoritative checks
Formatting, type check, lint, unit and integration tests, dependency audit and
secret scan run on every pull request and are required status checks on the
default branch. Their result is fact. No reviewer, human or AI, overrides a
failing check; the check is fixed or the policy owner grants a written exception.
## 2. Risk lanes
Lane is set by changed paths through CODEOWNERS, never by the author.
- Lane 0: docs, test-only, generated files with a generator check.
- Lane 1: application code outside lane 2.
- Lane 2: auth, authorization, payments, personal data, migrations, public API
contracts, infrastructure, CI workflows, agent configuration (CLAUDE.md,
AGENTS.md, .claude/, .codex/, .cursor/, MCP config), review governance
(CODEOWNERS, REVIEW.md, this policy file).
A pull request touching several lanes takes the highest one.
## 3. AI review
- Lane 1: one AI review against spec.md (or the linked issue) and the diff.
- Lane 2: lane 1 plus one security review pass.
- A second general AI reviewer is added only after a 30-day trial shows it
finds confirmed defects the first one missed.
## 4. Blocking findings
An AI finding blocks merge only if it has a reproducer, a failing test, or a
cited clause of this policy. Other findings are advisory. The author resolves
each blocking finding with a commit or a written rejection the approver accepts.
## 5. Approval
- Lane 0: any team member. Machine approval: [allowed | not allowed].
- Lane 1: one approver from the owning team.
- Lane 2: the code owner for the affected system, by name. No machine approval.
The approver answers for intent, risk and evidence, not for having read every
line.
## 6. Reviewer failure
A review agent that errors, times out or is missing reports "evidence
incomplete". The pull request waits or the approver records why it proceeds.
Silence is never approval.
## 7. Data boundaries
Repositories on the restricted list do not send code to external review
services. Current list and approved vendors: [link].
## 8. Measurement
Monthly: finding precision, duplicate rate, escaped defects, time in review,
unreviewed merges, review cost per merged PR. A reviewer below 50% precision for
two consecutive months is retuned or removed.

The 50% precision floor is a starting point, not a benchmark. Set it from your own shadow-mode data (step 4 below) and write down why.

Enforce the policy in the forge, not in the bot

Section titled “Enforce the policy in the forge, not in the bot”

Review bots comment; the forge decides. Put the authority in branch protection or a ruleset on the default branch, so it holds when a bot is down, misconfigured, or replaced:

  • Require the authoritative checks as status checks.
  • Require review from code owners, and choose one of two approval settings. The required approval count applies to the whole branch, not per lane:
    • Option A: zero required approvals plus required code-owner review. Lane 0 merges without a person; lanes 1 and 2 still need their owner. This works only with the CODEOWNERS layout below, where every lane 1 path has a default owner.
    • Option B: one required approval. Every lane, lane 0 included, needs one human approval.
  • Map paths to lanes in CODEOWNERS. The last matching pattern wins, and a pattern with no owner leaves its paths unowned. So the catch-all line gives lane 1 its default owner, ownerless lines after it take lane 0 paths out, and lane 2 lines come last:
.github/CODEOWNERS
# Lane 1: every path defaults to the engineering team
* @acme/engineering
# Lane 0: no owner, so no code-owner review is required
/docs/
/tests/
# Lane 2: sensitive paths need the owning team's approval
/src/auth/ @acme/identity
/src/billing/ @acme/payments
/db/migrations/ @acme/data-platform
/infra/ @acme/platform
/.github/ @acme/platform
/REVIEW.md @acme/platform
/.claude/ @acme/platform
/.codex/ @acme/platform
/.mcp.json @acme/platform
# No leading slash: matches in every directory, after /docs/ and /tests/
CLAUDE.md @acme/platform
AGENTS.md @acme/platform
.cursor/ @acme/platform

/.github/ covers the workflows and the CODEOWNERS file itself, so a pull request that edits either needs the platform team, not just any engineer. The last three lines have no leading slash, so they match CLAUDE.md, AGENTS.md and .cursor/ in every directory. Because they come after /docs/ and /tests/, a docs/AGENTS.md or src/payments/CLAUDE.md still goes to the platform team instead of falling into lane 0 or lane 1. GitHub applies the CODEOWNERS file from the pull request’s base branch, so a change to it cannot approve itself; it only takes effect after an owner approves it and it merges. Keep the file in .github/: GitHub looks there first, then at the root, then in docs/, and uses the first one it finds. In the central repository that holds review-policy.md, give that file a named owner the same way, for example /review-policy.md @acme/eng-leadership, with code-owner review required there too.

Replace @acme/... with your GitHub teams. Keep AI reviewer checks advisory in the forge until their precision is measured. A bot’s check becomes required only when the policy says its findings may block.

How Claude Code, Codex and Cursor fit the policy

Section titled “How Claude Code, Codex and Cursor fit the policy”

The three vendors differ most on one question: whether the review agent can approve. The policy answers that question, not the vendor default. Setup steps are in layered pull-request review and AI code review bots compared.

  • Managed Code Review (research preview, Team and Enterprise plans) reviews each GitHub pull request with several agents, posts severity-tagged inline comments, and adds a check run named Claude Code Review. It never approves or blocks, so blocking comes only from your policy and the forge. Anthropic’s docs put the average cost at $15–25 per review, billed as usage credits (checked 2026-09-26).
  • Request it on a PR with @claude review, or @claude review always to review every later push. It reads CLAUDE.md and a repository-root REVIEW.md of review-only rules (local /code-review does not read REVIEW.md); put the lane rules in REVIEW.md.
  • Neither Code Review nor claude ultrareview is available to organizations on Zero Data Retention. Record that in section 7 of the policy.
  • The check run always completes as neutral, even when a review fails or times out. To enforce section 6, add a small required CI job that fails with “evidence incomplete” when the Claude Code Review check run is missing or reports a failed or timed-out review on the head commit. Where findings may block, the same job reads the machine-readable severity line that Code Review writes at the end of the check run’s details text. Cap monthly spend for the Claude Code Review service at claude.ai/admin-settings/usage (Anthropic’s Code Review docs, checked 2026-09-26).
  • For the lane 2 security pass, run /security-review in a session before the pull request opens, or the anthropics/claude-code-security-review GitHub Action in CI.
  • On any plan, /code-review reviews a diff locally, and anthropics/claude-code-action@v1 runs a custom review in GitHub Actions. Follow AI in CI/CD for the permissions and trigger rules.

Roll the policy out without stalling delivery

Section titled “Roll the policy out without stalling delivery”
  1. Inventory what runs today. Use the first prompt below to list every review bot, REVIEW.md, CODEOWNERS file, and required check per repository. Expect to find bots enabled by individuals and default branches with no required checks.
  2. Publish version 1.0 of the policy. Adopt the template, fill in the lane 0 machine-approval choice, the restricted repository list, and the owners. Announce that it applies to agent-authored and human-authored pull requests alike.
  3. Map paths to lanes. Generate a CODEOWNERS proposal with the second prompt and have each owning team confirm its paths. Unclassified paths default to lane 1.
  4. Run AI reviewers in shadow mode for two to four weeks. They comment, nothing they say blocks. Label each finding as confirmed, rejected, or duplicate, which gives you precision per reviewer.
  5. Turn on enforcement. Make the authoritative checks and code-owner review required. Promote an AI reviewer’s findings to blocking only where its shadow-mode precision clears the floor you set.
  6. Audit monthly. Run the third prompt on a sample of merged pull requests, review the six metrics, and version the policy when you change a rule.

Measure whether each AI reviewer earns its cost

Section titled “Measure whether each AI reviewer earns its cost”

Report these per reviewer and per lane. Link them into your AI metrics panel rather than building a separate dashboard; the org-wide delivery metrics they sit beside are defined in metrics frameworks.

MetricDefinitionWhat it tells you
Finding precisionConfirmed findings (led to a change or accepted as real) ÷ all findingsWhether people should read this reviewer at all
Duplicate rateFindings already raised by a check or another reviewer ÷ all findingsWhether a second reviewer adds anything
Escaped defectsIncidents or bug reports traced to a merged PR, per lane, per monthWhether the lane’s evidence bar is high enough
Time in reviewMedian time from “ready for review” to merge, and median time to first human reviewWhether review is the bottleneck
Unreviewed mergesLane 1 and 2 PRs merged with no human approval ÷ all lane 1 and 2 mergesWhether the policy is enforced; target is zero
Review cost per merged PRAI review spend ÷ merged PRs, per repositoryWhether cost scales with value

Do not count comments, bot coverage, or “PRs reviewed by AI”. A reviewer that comments on every pull request and is right a third of the time costs attention it does not repay.

Run these in Claude Code, Codex, or Cursor’s agent with the gh CLI authenticated. Each one is read-only or writes only to a new branch.

How do you know the policy works without reading every diff?

Section titled “How do you know the policy works without reading every diff?”

The approver judges evidence, not lines. For each pull request, the policy guarantees a fixed evidence set: the authoritative checks passed, the lane is known from paths, blocking findings carry a reproducer or a failing test, and the PR description records intent, systems and data touched, reversibility, and open risks (the evidence bundle defines that record). The approver asks whether the evidence covers the riskiest behavior; if not, they return the pull request. The agent PR triage is the six-step routine for that judgment.

At the organization level, three signals prove the policy is in force: unreviewed merges in lanes 1 and 2 stay at zero, the monthly audit finds no silent passes, and escaped defects per lane do not rise as volume grows. When they do rise, raise the evidence bar for that lane before adding another bot. Moving a team from reading every diff to judging evidence is covered in trust transfer.

What breaks when an organization automates PR review?

Section titled “What breaks when an organization automates PR review?”

Before this page, place review in the full control map with governance and autonomy. After it, configure the review agents and run the server-side loop.