Skip to content

AI Code Review Bots Compared: Claude Code Review, Codex, Bugbot, CodeRabbit, Greptile and More

AI code review bots are second models that read a diff and post findings before a human approves it. The main options in September 2026 are Claude Code Review, claude-code-action, codex review, GitHub Copilot code review, Cursor’s Bugbot, CodeRabbit, Greptile and Alibaba’s open-source ocr. No single bot is enough: run them in layers and measure each on precision and cost per pull request.

Your team merges more pull requests every month, most of them written by agents, and review has become the queue everyone waits in. Someone installs a review bot, it posts 30 comments on the first pull request, and within a week people resolve its threads without reading them. This page is for developers who want a second opinion they can trust before pushing, and for tech leads who have to choose, configure and pay for the bots the whole team lives with.

What you’ll get from this review bot comparison

Section titled “What you’ll get from this review bot comparison”
  • A comparison table of eleven review options: where each runs, what triggers it, which file tunes it, and what a review costs.
  • A four-layer review workflow (local self-review, second model, pull request bot, human) with the order and the cost of each layer.
  • The claude-code-action@v1 review workflow adapted from Anthropic’s docs, a REVIEW.md that cuts noise, and a workflow that turns the neutral check run into a required status check.
  • The equivalent wiring for Codex and Cursor, and install commands for Greptile, CodeRabbit and ocr.
  • Three copy-paste prompts, four metrics that prove a bot earns its cost, and the traps that make each tool misfire.

This page compares the bots. The procedure a human follows on an agent’s pull request is reviewing an agent’s pull request without reading every line; read that first if you have not set up a review checklist yet.

Every tool below was checked on 2026-09-26 against the vendor’s docs, its GitHub repository or npm, or the installed CLIs (Claude Code 2.1.283, Codex CLI 0.157.1). Rows marked secondary rest on search extracts or on a check dated 2026-08-28, not on a vendor page read that day; confirm them on the vendor’s site before you roll them out.

ToolRuns whereTriggerTuned byCost per reviewEvidence
Claude Code Review (research preview, Team and Enterprise)Anthropic’s infrastructure, GitHubPR open, every push, or @claude reviewREVIEW.md, CLAUDE.md“averages $15-25”, billed as usage creditsVendor docs
/code-review (any plan) and claude ultrareviewYour terminal; ultrareview in Anthropic’s cloudYou run itCLAUDE.md (not REVIEW.md)Normal usage; ultrareview: 3 free runs on Pro and Max, then “typically $5 to $25”Vendor docs, CLI
claude-code-action@v1Your GitHub Actions runnerWorkflow events you chooseThe prompt, the plugin, CLAUDE.mdTokens on your API key or subscriptionVendor docs
codex review and openai/codex-action@v1Your terminal or runnerYou run it; workflow eventsCustom instructions, AGENTS.mdCodex usage limits or API tokensCLI help, action README
@codex review on GitHubOpenAI’s cloudPR commentAGENTS.mdCodex planSecondary (checked 2026-08-28)
GitHub Copilot code review (paid Copilot plans; not Copilot Free)GitHub; its agentic context gathering runs on Actions runnersRequest Copilot under Reviewers, or an automatic-review setting or ruleset.github/copilot-instructions.md, .github/instructions/*.instructions.md, AGENTS.mdAI credits: an estimated $0.05–1 per review at Lite effort, $0.25–5 at Balanced, plus Actions minutesgithub/docs
Cursor BugbotCursor’s cloud, GitHub and GitLabPR open or update, bugbot run.cursor/BUGBOT.mdNot verifiedSecondary
CodeRabbit (app and CLI)Vendor cloud; CLI locallyPR; coderabbit reviewVendor configNot verifiedCLI flags verified from CodeRabbit’s own skill
Greptile (app and CLI)Vendor cloud; CLI locallyPR; greptile review--instructionsNot verifiednpm greptile 3.6.0 README
Alibaba Open Code Review (ocr)Your terminal or runner, your modelYou run itRule templatesYour model’s tokensGitHub README, npm 1.12.9
Graphite Agent (formerly Diamond, Cursor-owned)Graphite’s cloudPRVendor configNot verifiedSecondary

Three differences matter more than the feature lists:

  • Who pays and how. The managed bots bill per review or per seat; the CLI and Actions routes bill tokens against a key you already own. Token prices live on the models hub, not here.
  • Whether it can block a merge. Claude Code Review’s check run “always completes with a neutral conclusion so it never blocks merging”. Copilot leaves a “Comment” review by default, which does not count toward required approvals; only with Copilot approvals (public preview) switched on can it approve. A gate is something you build in CI from the bot’s output.
  • What the bot reads for rules. Each bot has its own file. A rule you put in .cursor/rules never reaches Bugbot, and a rule in REVIEW.md never reaches the local /code-review.

One bot run by the same model that wrote the code shares that model’s blind spots. Layers work because each one catches a different class of problem at a different price, and each one is cheaper than the next. The tool-specific setups are in code review with Codex and automated code reviews in Claude Code; the team-level review contract, with focused passes for correctness, security, tests and spec compliance, is in govern layered pull-request review.

  1. Local self-review, before you push. Run the author tool’s own review in a fresh context: /code-review in Claude Code, codex review --uncommitted in Codex, or the Cursor agent with the first prompt below. It takes seconds to minutes within your normal usage and removes the obvious bugs before anyone else sees them.

  2. A second model, still local. Ask a different vendor’s model for an adversarial pass: /codex:review or /codex:adversarial-review from Claude Code, greptile review, coderabbit review --agent, or ocr review. A different model family disagrees with the author for reasons of its own, which is the point.

  3. The pull request bot, once per pull request. One bot owns the pull request: Claude Code Review, a claude-code-action workflow, @codex review, Copilot code review, or Bugbot. It posts inline comments with severities, and CI reads its output to decide whether the merge is blocked. Review on every push is an opt-in for high-risk pull requests, because it multiplies the cost.

  4. A human, on escalation classes only. A named person reads the code for authentication, payments, migrations and anything else on your escalation list, and signs off on everything else from the evidence. The escalation list and the verdicts are in agent PR review.

What a layered review costs per pull request. Layers 1 and 2 run on usage you already pay for. Layer 3 is the line item: with Claude Code Review at its published $15–25 average, a team merging 200 pull requests a month with one review each spends roughly $3,000–5,000 a month, and “after every push” mode multiplies that by the pushes per pull request. That is arithmetic on Anthropic’s average, not a measurement, so read the per-repository average cost in the admin settings after the first month. A claude-code-action or codex-action workflow bills only tokens plus runner minutes, and you set the model and effort, so it is the cheaper layer 3 when you can live without the managed verification step.

Set up a pull request review bot in each tool

Section titled “Set up a pull request review bot in each tool”

The workflow above is the same in every tool. The wiring differs, so pick your tab.

Managed Code Review (Team and Enterprise). An organization Owner enables it at claude.ai/admin-settings/claude-code, installs the Claude GitHub App, selects repositories, and chooses a trigger: once after PR creation, after every push, or Manual. To verify setup, open a test pull request; a check run named Claude Code Review appears within a few minutes. On a Manual repository, comment @claude review for one review, or @claude review always to subscribe the pull request to later pushes. It is not available to organizations with Zero Data Retention.

Your own workflow, on any plan with an API key or OAuth token. This workflow is adapted from Anthropic’s GitHub Actions docs (we add persist-credentials: false). Save it as .github/workflows/claude-code-review.yml:

name: Code Review
on:
pull_request:
types: [opened, synchronize, ready_for_review, reopened]
jobs:
review:
runs-on: ubuntu-latest
permissions:
contents: read
pull-requests: read
issues: read
id-token: write
steps:
- uses: actions/checkout@v6
with:
fetch-depth: 1
persist-credentials: false
- uses: anthropics/claude-code-action@v1
with:
anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }}
plugin_marketplaces: "https://github.com/anthropics/claude-code.git"
plugins: "code-review@claude-code-plugins"
prompt: "/code-review:code-review --comment ${{ github.repository }}/pull/${{ github.event.pull_request.number }}"
claude_args: '--allowedTools "mcp__github_inline_comment__create_inline_comment"'

Two lines decide where the review goes. --comment posts inline comments on the pull request; without it, findings stay in the run log. Keep the claude_args line: the action starts the inline-comment MCP server only when --allowedTools names it. The workflow skips drafts, closed pull requests and pull requests Claude already commented on. The job triggers on pull_request, not pull_request_target, so on public repositories GitHub withholds ANTHROPIC_API_KEY from fork pull requests and forks get no review. persist-credentials: false keeps the GITHUB_TOKEN out of .git/config while the agent runs on the pull request’s code. To authenticate with a subscription instead, run claude setup-token locally and pass the token as claude_code_oauth_token.

Local layers. /code-review reviews your branch and uncommitted changes as a background subagent; /code-review high widens coverage, low and medium report only the most confident findings, --fix applies findings, and --comment posts them to a pull request. For a deep pass before merge, claude ultrareview 482 reviews pull request 482 in the cloud, where “every reported finding is independently reproduced and verified”; --json prints the raw payload.

REVIEW.md sits at the repository root and reaches every agent that finds and verifies findings, so a rule there lands more reliably than the same rule in a long CLAUDE.md. Anthropic lists seven patterns; the four that cut the most noise are redefining Important, capping nits, skip rules and repository-specific checks, plus a re-review rule so a pull request converges instead of collecting new nits on every push. This version, adapted from Anthropic’s example, is a good first commit:

# Review instructions
## What Important means here
Reserve Important for findings that would break behavior, leak data,
or block a rollback: incorrect logic, unscoped database queries, PII in
logs or error messages, and migrations that aren't backward compatible.
Style, naming and refactoring suggestions are Nit at most.
## Cap the nits
Report at most five Nits per review. If you found more, say "plus N
similar items" in the summary. After the first review of a PR, post
Important findings only.
## Do not report
- Anything CI already enforces: lint, formatting, type errors
- Generated files under `src/gen/` and any `*.lock` file
## Always check
- New API routes have an integration test
- Log lines don't include email addresses, user IDs or request bodies
- Database queries are scoped to the caller's tenant

Keep it short. Anthropic’s own guidance is that “a long REVIEW.md dilutes the rules that matter most”.

Turn the neutral check run into a merge gate

Section titled “Turn the neutral check run into a merge gate”

The last line of the Claude Code Review check run’s details is a machine-readable severity count. Claude Code Review takes about 20 minutes on average, so the gate has to wait for it rather than poll once. This workflow runs when the check run completes and posts a commit status on the reviewed commit: success when the Important count is zero, failure otherwise. Save it as .github/workflows/claude-review-gate.yml on the default branch:

name: Claude review gate
on:
check_run:
types: [completed]
permissions:
checks: read
statuses: write
jobs:
gate:
if: github.event.check_run.name == 'Claude Code Review'
runs-on: ubuntu-latest
steps:
- name: Turn the severity count into a commit status
env:
GH_TOKEN: ${{ github.token }}
REPO: ${{ github.repository }}
RUN_ID: ${{ github.event.check_run.id }}
SHA: ${{ github.event.check_run.head_sha }}
run: |
# The severity line is JSON like {"normal": 2, "nit": 1, "pre_existing": 0}; "normal" = Important
important=$(gh api "repos/$REPO/check-runs/$RUN_ID" \
--jq '.output.text | split("bughunter-severity: ")[1] | split(" -->")[0] | fromjson | .normal') || important=""
if [ "$important" = "0" ]; then state=success; desc="No Important findings"
else state=failure; desc="Important findings: ${important:-unknown}"; fi
gh api "repos/$REPO/statuses/$SHA" -f state="$state" \
-f context="claude-review-gate" -f description="$desc"

Mark claude-review-gate as a required status check in branch protection or a ruleset; a check that fails but is not required blocks nothing. The workflow posts a commit status instead of relying on its own job result because, per GitHub’s docs, a check_run workflow runs against the default branch, so its job never appears on the pull request. Until the status arrives, the required check shows as expected and the merge stays blocked, which also covers a review that is still running. On a Manual repository, nothing is posted until someone comments @claude review.

A commit status belongs to one commit and does not carry forward. If the author pushes after the review, the gate on the new tip stays pending and the merge stays blocked until that commit is reviewed. That is the intended behaviour, because an unreviewed commit should not merge, but plan for it: comment @claude review on the final commit before merging, or use @claude review always (or the repository’s review-on-every-push trigger) on pull requests that change after review. Budget for the extra runs, since each review is billed.

For Greptile, greptile review status exits 0 when the commit has a completed review, 3 while one is running, 4/5 for a failed or cancelled review, and 1 when there is none or you are signed out, so a pre-push hook or CI step can require a finished review before anything merges. Treat every code other than 0 as “not reviewed”.

Add a second model: Greptile, CodeRabbit, ocr and the Codex plugin

Section titled “Add a second model: Greptile, CodeRabbit, ocr and the Codex plugin”

Each of these gives you a reviewer from outside your author tool. Install commands are verified against each project’s own README or npm entry on 2026-09-26, except CodeRabbit’s installer page, which is secondary.

Codex plugin for Claude Code (OpenAI, openai/codex-plugin-cc). It runs your local Codex install from inside a Claude Code session and counts against your Codex usage limits. /codex:review is “a normal read-only Codex review” and takes no focus text; /codex:adversarial-review is steerable and challenges the design.

# Inside Claude Code
/plugin marketplace add openai/codex-plugin-cc
/plugin install codex@openai-codex
/reload-plugins
/codex:setup
/codex:adversarial-review --base main challenge whether this retry design is safe under concurrent writes

/codex:setup --enable-review-gate adds a Stop hook that runs a Codex review on every Claude response and blocks the stop when it finds issues. The README warns that the gate “can create a long-running Claude/Codex loop and may drain usage limits quickly”, so enable it only in a session you are watching.

Greptile CLI (npm greptile 3.6.0, Node 22 or later). greptile review compares the merge base with HEAD; --plus and --apex raise the effort.

Terminal window
npm install -g greptile # or: brew install greptileai/tap/greptile
greptile # signs you in on first run
greptile init # admin only: enables the repository
greptile review -b main --instructions "focus on retry cancellation"

CodeRabbit CLI. Install it from CodeRabbit’s official installer or Homebrew, not from npm. CodeRabbit’s own coderabbitai/skills SKILL.md points to https://www.coderabbit.ai/cli. The install lines below are secondary: they come from search extracts of CodeRabbit’s docs, and the installer page was not re-read on 2026-09-26. Do not install the npm name: the npm package coderabbit is a “security holding package” (version 0.0.1-security.1). --agent returns findings an agent can act on, with severities from critical to info.

Terminal window
curl -fsSL https://cli.coderabbit.ai/install.sh | sh # or: brew install --cask coderabbit (secondary)
coderabbit auth login
coderabbit review --agent --uncommitted
coderabbit review --agent --base main

Alibaba Open Code Review (ocr), Apache-2.0, npm @alibaba-group/open-code-review 1.12.9. It pairs deterministic file selection and rule matching with an LLM agent, and works with any OpenAI- or Anthropic-compatible endpoint, so the code never leaves the provider you already use.

Terminal window
npm install -g @alibaba-group/open-code-review
ocr config provider && ocr config model
ocr review --from main --to feature-branch --format json --output result.json

Alibaba’s README reports higher precision and F1 than Claude Code with the same model, at about one ninth of the tokens, and says recall is lower, “a deliberate trade-off favoring precision over noise”. That is a vendor benchmark (200 pull requests from 50 repositories), not reproduced here. Choose ocr for a CI layer where noise costs you more than a missed nit.

As of 2026-09-26, anthropics/claude-code-action had 8,951 GitHub stars, openai/codex-action 1,248, openai/codex-plugin-cc 33,594 and alibaba/open-code-review 41,400; in Anthropic’s plugin directory, greptile showed 56,611 installs and coderabbit 32,361 (GitHub and claude.com/plugins, read for this site’s ecosystem catalogue). Stars measure attention, not review quality. Both plugins are in Anthropic’s official marketplace (claude-plugins-official, checked 2026-09-26):

# Inside Claude Code
/plugin install greptile@claude-plugins-official
/plugin install coderabbit@claude-plugins-official

The plugins carry context cost in every session. Before you keep one, run claude plugin details greptile or claude plugin details coderabbit in a terminal, which shows the plugin’s components and projected token cost, and compare /context before and after.

Copy-paste prompts for layered AI code review

Section titled “Copy-paste prompts for layered AI code review”

The third prompt produces the numbers you need for the precision metric below. In Claude Code with the GitHub CLI installed, the agent reads comments with gh api; in Codex and Cursor, use the GitHub MCP server or gh in the terminal.

How do you prove a review bot is worth its cost?

Section titled “How do you prove a review bot is worth its cost?”

A review bot is itself an unverified agent until you measure it. Track four numbers per bot per month, and review them with the team:

MetricDefinitionAct when
PrecisionFindings that led to a code change, divided by findings raisedIt falls below the level where people still read the comments; tighten REVIEW.md or BUGBOT.md, or lower the effort
EscapesProduction defects in code the bot reviewed, where a comment from the bot would have been enoughIt rises; add the missed class to the bot’s “Always check” list
Cost per merged PRBot spend for the month divided by merged pull requestsIt rises faster than merged PRs; move the bot from “every push” to “once per PR” or Manual
Time to first reviewMinutes from the PR opening to the bot’s first commentIt exceeds your review SLA; Claude Code Review averages 20 minutes, so do not make humans wait for it

The bots never own the decision. CI owns the gate (tests, types, lint, the severity-count check above), the bot owns the comments, and the approver named on the pull request owns the merge. The code owners in CODEOWNERS sign off on escalation classes, whatever the bots said. Canonical definitions for team metrics live in metrics frameworks for agentic engineering.

What breaks when you run AI code review bots?

Section titled “What breaks when you run AI code review bots?”

The bot drowns the real bug in nits. Thirty comments, one of them important, and the team stops reading. Recovery: cap nits in REVIEW.md or BUGBOT.md, add “after the first review, post Important findings only”, and track precision. If precision keeps falling, switch that bot to Manual.

Two bots argue on the same pull request. Claude Code Review and Bugbot both comment, disagree, and the author fixes one and reopens the other. Recovery: one bot owns layer 3. Run the second model locally (layer 2), where its output goes to the author, not to the pull request.

Push-triggered reviews multiply the bill. A pull request with 12 pushes gets 12 reviews. Recovery: use “once after PR creation” or Manual, and ask for @claude review always only on high-risk pull requests. Set a monthly spend cap for Claude Code Review at claude.ai/admin-settings/usage.

Forks and drafts get no review, and nobody notices. GitHub withholds secrets from fork pull requests, claude-code-action skips drafts, and Code Review reviews forks only on a comment command. Recovery: make the review job a required check with a clear “skipped” state, and have a maintainer comment @claude review on fork pull requests before approving.

The team treats a clean bot run as approval. “The bot found nothing” becomes the merge reason. Recovery: gate on tests and the severity count, leave Copilot approvals off unless a human still signs off on escalation classes, and keep a named human approver on every pull request.

A review loop drains usage. The Codex plugin’s review gate or an autofix bot keeps finding and fixing new issues. Recovery: enable the gate only in watched sessions, stop after two fix rounds, and escalate to a human with the open findings listed.