Team PR review automation policy — evidence before approval
A team PR review automation policy is the organization-wide rule set that decides what review agents may do on a pull request: which deterministic checks are authoritative, which risk lane a change takes, what makes an AI finding blocking, who may approve, and what a failed reviewer means. Its core rule: no pull request merges on an AI opinion alone.
After a quarter of experiments, your organization runs three review bots. The platform team enabled Claude Code Review, the web team turned on Bugbot, and someone added @codex review to a few repositories. One pull request collects 41 comments that nobody reads. Another merges after a bot said “no issues found”, and it carried a migration that no person looked at. Then your security lead asks which of these bots is allowed to approve anything, and nobody can answer.
This page is for the CTO or VP Engineering who owns that answer and the tech leads who enforce it. It covers the policy only. Configuring each review agent is in layered pull-request review, comparing the bots is in AI code review bots compared, and running the daily queue is in the review queue.
Q9 · Quality gates Max-score evidence: deterministic checks, a focused specification and risk review, reproducible findings, risk-based routing, and a named human merge gate.
What an org-wide review policy gives you
Section titled “What an org-wide review policy gives you”- A one-page
review-policy.mdtemplate you can adopt as-is and adjust per repository. - A risk-lane table that tells every repository which evidence, which reviewers, and which approver a change needs.
- Enforcement in the forge (required checks and code owners), so the policy holds even when a bot is down.
- Six metric definitions that show whether each AI reviewer is worth its cost.
- Copy-paste prompts that inventory your current review bots and audit merged pull requests against the policy.
Why review automation needs a policy, not more bots
Section titled “Why review automation needs a policy, not more bots”Agents raise the number and size of pull requests faster than review capacity grows. Faros AI’s “AI Engineering Report 2026” (April 2026; telemetry from 22,000 developers on Faros’s own platform, so a self-selected customer base) measured pull request size up 51% and pull requests merged per developer up 16.2%, while median time in review rose 441.5%, pull requests merged without any review rose 31.3%, and incidents per pull request rose 242.7%.
Adding reviewers does not fix that on its own. Each bot adds comments, and a comment is not evidence. Without a written policy, teams drift into one of two failure states: they approve without reading because the queue is too long, or they treat a bot’s silence as approval. The policy exists to stop both. It says which signals are facts, which are opinions that need proof, and which person carries the decision.
What must a review automation policy decide?
Section titled “What must a review automation policy decide?”Settle these eight decisions once, at the organization level. Repositories may tighten a rule; they do not loosen one without an exception on file.
| Decision | Recommended default | Owner |
|---|---|---|
| Authoritative checks | Formatting, types, lint, unit and integration tests, dependency audit, secret scan. Pass/fail, required in the forge. | Platform team |
| Risk lanes | Three lanes assigned by changed paths, never by the author’s own label. | CTO, with security |
| AI reviewers per lane | At most one general reviewer, plus one security pass in the sensitive lane. Add a second only when metrics justify it. | Tech leads |
| What an AI finding may block | Only a finding with a reproducer, a failing test, or a cited policy clause. Everything else is advisory. | Tech leads |
| Approval authority | A named human approves lanes 1 and 2. Machine approval, if allowed at all, only in lane 0. | CTO |
| Reviewer failure | A reviewer that errors, times out, or is not configured reports “evidence incomplete”. It never counts as a pass. | Platform team |
| Data boundaries | Which repositories may send code to which review vendor, and under which retention terms. | Security and legal |
| Measurement and review | The six metrics below, reviewed monthly; a reviewer below its precision floor for two months is retuned or removed. | Engineering leadership |
Assign every change to a risk lane
Section titled “Assign every change to a risk lane”Lanes decide how much evidence a change needs before a person may approve it. Classify by changed paths through CODEOWNERS, so a change cannot pick its own lane.
| Lane | What it contains | Automated evidence | Human gate | Machine approval |
|---|---|---|---|---|
| 0 · Low risk | Documentation, test-only changes, generated files with a generator check | Authoritative checks | Any team member, or none if the policy allows it | Optional |
| 1 · Standard | Application code outside lane 2 | Authoritative checks, one AI review against spec.md and the diff | One approver from the owning team | No |
| 2 · Sensitive | Authentication, authorization, payments, personal data, migrations, public API contracts, infrastructure, CI workflows, agent configuration (CLAUDE.md, AGENTS.md, .claude/, .codex/, .cursor/, MCP config), and review governance (CODEOWNERS, REVIEW.md, the policy file) | Lane 1 evidence, a security review pass, and a rollback note in the PR | A code owner for the affected system, by name | No |
Agent configuration sits in lane 2 on purpose. A change to AGENTS.md, a hook, or an MCP server entry changes how every later agent run behaves, including the review agents themselves, so it needs the same scrutiny as a CI workflow. Nested instruction files count as agent configuration too: Claude Code loads a CLAUDE.md from subdirectories, and Codex reads AGENTS.md per directory, so src/payments/AGENTS.md steers the review of the payments code. See shared hooks governance for the rules on hook changes.
Review governance sits in lane 2 for the same reason. CODEOWNERS decides every lane, REVIEW.md tells the review agents what to look for, and the policy file sets the rules both must satisfy. Left to the catch-all, they would be lane 1 files, and one approval from the default team could rewrite the lanes themselves.
Adopt the policy template
Section titled “Adopt the policy template”Copy this file into a central repository (for example your engineering handbook) and link it from every repository’s REVIEW.md. The per-repository REVIEW.md implements the policy for its review agents; this file is what it must satisfy.
# Pull request review policy (version 1.0, owner: VP Engineering)
## 1. Authoritative checksFormatting, type check, lint, unit and integration tests, dependency audit andsecret scan run on every pull request and are required status checks on thedefault branch. Their result is fact. No reviewer, human or AI, overrides afailing check; the check is fixed or the policy owner grants a written exception.
## 2. Risk lanesLane is set by changed paths through CODEOWNERS, never by the author.- Lane 0: docs, test-only, generated files with a generator check.- Lane 1: application code outside lane 2.- Lane 2: auth, authorization, payments, personal data, migrations, public API contracts, infrastructure, CI workflows, agent configuration (CLAUDE.md, AGENTS.md, .claude/, .codex/, .cursor/, MCP config), review governance (CODEOWNERS, REVIEW.md, this policy file).A pull request touching several lanes takes the highest one.
## 3. AI review- Lane 1: one AI review against spec.md (or the linked issue) and the diff.- Lane 2: lane 1 plus one security review pass.- A second general AI reviewer is added only after a 30-day trial shows it finds confirmed defects the first one missed.
## 4. Blocking findingsAn AI finding blocks merge only if it has a reproducer, a failing test, or acited clause of this policy. Other findings are advisory. The author resolveseach blocking finding with a commit or a written rejection the approver accepts.
## 5. Approval- Lane 0: any team member. Machine approval: [allowed | not allowed].- Lane 1: one approver from the owning team.- Lane 2: the code owner for the affected system, by name. No machine approval.The approver answers for intent, risk and evidence, not for having read everyline.
## 6. Reviewer failureA review agent that errors, times out or is missing reports "evidenceincomplete". The pull request waits or the approver records why it proceeds.Silence is never approval.
## 7. Data boundariesRepositories on the restricted list do not send code to external reviewservices. Current list and approved vendors: [link].
## 8. MeasurementMonthly: finding precision, duplicate rate, escaped defects, time in review,unreviewed merges, review cost per merged PR. A reviewer below 50% precision fortwo consecutive months is retuned or removed.The 50% precision floor is a starting point, not a benchmark. Set it from your own shadow-mode data (step 4 below) and write down why.
Enforce the policy in the forge, not in the bot
Section titled “Enforce the policy in the forge, not in the bot”Review bots comment; the forge decides. Put the authority in branch protection or a ruleset on the default branch, so it holds when a bot is down, misconfigured, or replaced:
- Require the authoritative checks as status checks.
- Require review from code owners, and choose one of two approval settings. The required approval count applies to the whole branch, not per lane:
- Option A: zero required approvals plus required code-owner review. Lane 0 merges without a person; lanes 1 and 2 still need their owner. This works only with the
CODEOWNERSlayout below, where every lane 1 path has a default owner. - Option B: one required approval. Every lane, lane 0 included, needs one human approval.
- Option A: zero required approvals plus required code-owner review. Lane 0 merges without a person; lanes 1 and 2 still need their owner. This works only with the
- Map paths to lanes in
CODEOWNERS. The last matching pattern wins, and a pattern with no owner leaves its paths unowned. So the catch-all line gives lane 1 its default owner, ownerless lines after it take lane 0 paths out, and lane 2 lines come last:
# Lane 1: every path defaults to the engineering team* @acme/engineering
# Lane 0: no owner, so no code-owner review is required/docs//tests/
# Lane 2: sensitive paths need the owning team's approval/src/auth/ @acme/identity/src/billing/ @acme/payments/db/migrations/ @acme/data-platform/infra/ @acme/platform/.github/ @acme/platform/REVIEW.md @acme/platform/.claude/ @acme/platform/.codex/ @acme/platform/.mcp.json @acme/platform# No leading slash: matches in every directory, after /docs/ and /tests/CLAUDE.md @acme/platformAGENTS.md @acme/platform.cursor/ @acme/platform/.github/ covers the workflows and the CODEOWNERS file itself, so a pull request that edits either needs the platform team, not just any engineer. The last three lines have no leading slash, so they match CLAUDE.md, AGENTS.md and .cursor/ in every directory. Because they come after /docs/ and /tests/, a docs/AGENTS.md or src/payments/CLAUDE.md still goes to the platform team instead of falling into lane 0 or lane 1. GitHub applies the CODEOWNERS file from the pull request’s base branch, so a change to it cannot approve itself; it only takes effect after an owner approves it and it merges. Keep the file in .github/: GitHub looks there first, then at the root, then in docs/, and uses the first one it finds. In the central repository that holds review-policy.md, give that file a named owner the same way, for example /review-policy.md @acme/eng-leadership, with code-owner review required there too.
Replace @acme/... with your GitHub teams. Keep AI reviewer checks advisory in the forge until their precision is measured. A bot’s check becomes required only when the policy says its findings may block.
How Claude Code, Codex and Cursor fit the policy
Section titled “How Claude Code, Codex and Cursor fit the policy”The three vendors differ most on one question: whether the review agent can approve. The policy answers that question, not the vendor default. Setup steps are in layered pull-request review and AI code review bots compared.
- Managed Code Review (research preview, Team and Enterprise plans) reviews each GitHub pull request with several agents, posts severity-tagged inline comments, and adds a check run named Claude Code Review. It never approves or blocks, so blocking comes only from your policy and the forge. Anthropic’s docs put the average cost at $15–25 per review, billed as usage credits (checked 2026-09-26).
- Request it on a PR with
@claude review, or@claude review alwaysto review every later push. It readsCLAUDE.mdand a repository-rootREVIEW.mdof review-only rules (local/code-reviewdoes not readREVIEW.md); put the lane rules inREVIEW.md. - Neither Code Review nor
claude ultrareviewis available to organizations on Zero Data Retention. Record that in section 7 of the policy. - The check run always completes as neutral, even when a review fails or times out. To enforce section 6, add a small required CI job that fails with “evidence incomplete” when the
Claude Code Reviewcheck run is missing or reports a failed or timed-out review on the head commit. Where findings may block, the same job reads the machine-readable severity line that Code Review writes at the end of the check run’s details text. Cap monthly spend for the Claude Code Review service at claude.ai/admin-settings/usage (Anthropic’s Code Review docs, checked 2026-09-26). - For the lane 2 security pass, run
/security-reviewin a session before the pull request opens, or theanthropics/claude-code-security-reviewGitHub Action in CI. - On any plan,
/code-reviewreviews a diff locally, andanthropics/claude-code-action@v1runs a custom review in GitHub Actions. Follow AI in CI/CD for the permissions and trigger rules.
- The Codex GitHub integration reviews pull requests on request (
@codex review,@codex security review) or automatically, and reads custom review rules fromAGENTS.md(OpenAI docs, verified 2026-08-28). Put the lane rules there. - For the lane 2 security pass, comment
@codex security reviewon the pull request. - In CI or on a laptop,
codex review --base mainreviews the branch againstmainand picks up review rules fromAGENTS.md. A custom-instructions prompt (an argument, or-for stdin) works only without--base,--commitor--uncommitted; with none of those flags set, the prompt says what to review (checked against codex 0.157.1). Its output is advisory text; your pipeline decides whether a finding fails a check. - Codex’s automatic approval review (“auto-review”) approves the agent’s own sandbox actions during a run. It is not pull request approval, and the policy should not treat it as review evidence.
- Bugbot reviews pull requests for bugs, security issues, and code quality problems (Cursor docs, verified 2026-08-28).
- PR Routing & Approval assigns reviewers by code ownership and commit history, and “can approve low-risk PRs when your criteria are met” (Cursor docs, verified 2026-08-28). This is the one feature here that approves. Allow it only in lane 0, and only if section 5 of the policy says so.
- For the lane 2 security pass, choose a security reviewer from AI code review bots compared; Bugbot is the general reviewer.
- Bugbot’s billing terms could not be verified from cursor.com on 2026-09-26; confirm current terms before you budget.
Roll the policy out without stalling delivery
Section titled “Roll the policy out without stalling delivery”- Inventory what runs today. Use the first prompt below to list every review bot,
REVIEW.md,CODEOWNERSfile, and required check per repository. Expect to find bots enabled by individuals and default branches with no required checks. - Publish version 1.0 of the policy. Adopt the template, fill in the lane 0 machine-approval choice, the restricted repository list, and the owners. Announce that it applies to agent-authored and human-authored pull requests alike.
- Map paths to lanes. Generate a
CODEOWNERSproposal with the second prompt and have each owning team confirm its paths. Unclassified paths default to lane 1. - Run AI reviewers in shadow mode for two to four weeks. They comment, nothing they say blocks. Label each finding as confirmed, rejected, or duplicate, which gives you precision per reviewer.
- Turn on enforcement. Make the authoritative checks and code-owner review required. Promote an AI reviewer’s findings to blocking only where its shadow-mode precision clears the floor you set.
- Audit monthly. Run the third prompt on a sample of merged pull requests, review the six metrics, and version the policy when you change a rule.
Measure whether each AI reviewer earns its cost
Section titled “Measure whether each AI reviewer earns its cost”Report these per reviewer and per lane. Link them into your AI metrics panel rather than building a separate dashboard; the org-wide delivery metrics they sit beside are defined in metrics frameworks.
| Metric | Definition | What it tells you |
|---|---|---|
| Finding precision | Confirmed findings (led to a change or accepted as real) ÷ all findings | Whether people should read this reviewer at all |
| Duplicate rate | Findings already raised by a check or another reviewer ÷ all findings | Whether a second reviewer adds anything |
| Escaped defects | Incidents or bug reports traced to a merged PR, per lane, per month | Whether the lane’s evidence bar is high enough |
| Time in review | Median time from “ready for review” to merge, and median time to first human review | Whether review is the bottleneck |
| Unreviewed merges | Lane 1 and 2 PRs merged with no human approval ÷ all lane 1 and 2 merges | Whether the policy is enforced; target is zero |
| Review cost per merged PR | AI review spend ÷ merged PRs, per repository | Whether cost scales with value |
Do not count comments, bot coverage, or “PRs reviewed by AI”. A reviewer that comments on every pull request and is right a third of the time costs attention it does not repay.
Copy-paste prompts for running the policy
Section titled “Copy-paste prompts for running the policy”Run these in Claude Code, Codex, or Cursor’s agent with the gh CLI authenticated. Each one is read-only or writes only to a new branch.
How do you know the policy works without reading every diff?
Section titled “How do you know the policy works without reading every diff?”The approver judges evidence, not lines. For each pull request, the policy guarantees a fixed evidence set: the authoritative checks passed, the lane is known from paths, blocking findings carry a reproducer or a failing test, and the PR description records intent, systems and data touched, reversibility, and open risks (the evidence bundle defines that record). The approver asks whether the evidence covers the riskiest behavior; if not, they return the pull request. The agent PR triage is the six-step routine for that judgment.
At the organization level, three signals prove the policy is in force: unreviewed merges in lanes 1 and 2 stay at zero, the monthly audit finds no silent passes, and escaped defects per lane do not rise as volume grows. When they do rise, raise the evidence bar for that lane before adding another bot. Moving a team from reading every diff to judging evidence is covered in trust transfer.
What breaks when an organization automates PR review?
Section titled “What breaks when an organization automates PR review?”Where to go next with review policy
Section titled “Where to go next with review policy”Before this page, place review in the full control map with governance and autonomy. After it, configure the review agents and run the server-side loop.