AI Code Review Bots Compared: Claude Code Review, Codex, Bugbot, CodeRabbit, Greptile and More
AI code review bots are second models that read a diff and post findings before a human approves it. The main options in September 2026 are Claude Code Review, claude-code-action, codex review, GitHub Copilot code review, Cursor’s Bugbot, CodeRabbit, Greptile and Alibaba’s open-source ocr. No single bot is enough: run them in layers and measure each on precision and cost per pull request.
Your team merges more pull requests every month, most of them written by agents, and review has become the queue everyone waits in. Someone installs a review bot, it posts 30 comments on the first pull request, and within a week people resolve its threads without reading them. This page is for developers who want a second opinion they can trust before pushing, and for tech leads who have to choose, configure and pay for the bots the whole team lives with.
What you’ll get from this review bot comparison
Section titled “What you’ll get from this review bot comparison”- A comparison table of eleven review options: where each runs, what triggers it, which file tunes it, and what a review costs.
- A four-layer review workflow (local self-review, second model, pull request bot, human) with the order and the cost of each layer.
- The
claude-code-action@v1review workflow adapted from Anthropic’s docs, aREVIEW.mdthat cuts noise, and a workflow that turns the neutral check run into a required status check. - The equivalent wiring for Codex and Cursor, and install commands for Greptile, CodeRabbit and
ocr. - Three copy-paste prompts, four metrics that prove a bot earns its cost, and the traps that make each tool misfire.
This page compares the bots. The procedure a human follows on an agent’s pull request is reviewing an agent’s pull request without reading every line; read that first if you have not set up a review checklist yet.
Which AI code review bot fits which job?
Section titled “Which AI code review bot fits which job?”Every tool below was checked on 2026-09-26 against the vendor’s docs, its GitHub repository or npm, or the installed CLIs (Claude Code 2.1.283, Codex CLI 0.157.1). Rows marked secondary rest on search extracts or on a check dated 2026-08-28, not on a vendor page read that day; confirm them on the vendor’s site before you roll them out.
| Tool | Runs where | Trigger | Tuned by | Cost per review | Evidence |
|---|---|---|---|---|---|
| Claude Code Review (research preview, Team and Enterprise) | Anthropic’s infrastructure, GitHub | PR open, every push, or @claude review | REVIEW.md, CLAUDE.md | “averages $15-25”, billed as usage credits | Vendor docs |
/code-review (any plan) and claude ultrareview | Your terminal; ultrareview in Anthropic’s cloud | You run it | CLAUDE.md (not REVIEW.md) | Normal usage; ultrareview: 3 free runs on Pro and Max, then “typically $5 to $25” | Vendor docs, CLI |
claude-code-action@v1 | Your GitHub Actions runner | Workflow events you choose | The prompt, the plugin, CLAUDE.md | Tokens on your API key or subscription | Vendor docs |
codex review and openai/codex-action@v1 | Your terminal or runner | You run it; workflow events | Custom instructions, AGENTS.md | Codex usage limits or API tokens | CLI help, action README |
@codex review on GitHub | OpenAI’s cloud | PR comment | AGENTS.md | Codex plan | Secondary (checked 2026-08-28) |
| GitHub Copilot code review (paid Copilot plans; not Copilot Free) | GitHub; its agentic context gathering runs on Actions runners | Request Copilot under Reviewers, or an automatic-review setting or ruleset | .github/copilot-instructions.md, .github/instructions/*.instructions.md, AGENTS.md | AI credits: an estimated $0.05–1 per review at Lite effort, $0.25–5 at Balanced, plus Actions minutes | github/docs |
| Cursor Bugbot | Cursor’s cloud, GitHub and GitLab | PR open or update, bugbot run | .cursor/BUGBOT.md | Not verified | Secondary |
| CodeRabbit (app and CLI) | Vendor cloud; CLI locally | PR; coderabbit review | Vendor config | Not verified | CLI flags verified from CodeRabbit’s own skill |
| Greptile (app and CLI) | Vendor cloud; CLI locally | PR; greptile review | --instructions | Not verified | npm greptile 3.6.0 README |
Alibaba Open Code Review (ocr) | Your terminal or runner, your model | You run it | Rule templates | Your model’s tokens | GitHub README, npm 1.12.9 |
| Graphite Agent (formerly Diamond, Cursor-owned) | Graphite’s cloud | PR | Vendor config | Not verified | Secondary |
Three differences matter more than the feature lists:
- Who pays and how. The managed bots bill per review or per seat; the CLI and Actions routes bill tokens against a key you already own. Token prices live on the models hub, not here.
- Whether it can block a merge. Claude Code Review’s check run “always completes with a neutral conclusion so it never blocks merging”. Copilot leaves a “Comment” review by default, which does not count toward required approvals; only with Copilot approvals (public preview) switched on can it approve. A gate is something you build in CI from the bot’s output.
- What the bot reads for rules. Each bot has its own file. A rule you put in
.cursor/rulesnever reaches Bugbot, and a rule inREVIEW.mdnever reaches the local/code-review.
How does layered AI code review work?
Section titled “How does layered AI code review work?”One bot run by the same model that wrote the code shares that model’s blind spots. Layers work because each one catches a different class of problem at a different price, and each one is cheaper than the next. The tool-specific setups are in code review with Codex and automated code reviews in Claude Code; the team-level review contract, with focused passes for correctness, security, tests and spec compliance, is in govern layered pull-request review.
-
Local self-review, before you push. Run the author tool’s own review in a fresh context:
/code-reviewin Claude Code,codex review --uncommittedin Codex, or the Cursor agent with the first prompt below. It takes seconds to minutes within your normal usage and removes the obvious bugs before anyone else sees them. -
A second model, still local. Ask a different vendor’s model for an adversarial pass:
/codex:reviewor/codex:adversarial-reviewfrom Claude Code,greptile review,coderabbit review --agent, orocr review. A different model family disagrees with the author for reasons of its own, which is the point. -
The pull request bot, once per pull request. One bot owns the pull request: Claude Code Review, a
claude-code-actionworkflow,@codex review, Copilot code review, or Bugbot. It posts inline comments with severities, and CI reads its output to decide whether the merge is blocked. Review on every push is an opt-in for high-risk pull requests, because it multiplies the cost. -
A human, on escalation classes only. A named person reads the code for authentication, payments, migrations and anything else on your escalation list, and signs off on everything else from the evidence. The escalation list and the verdicts are in agent PR review.
What a layered review costs per pull request. Layers 1 and 2 run on usage you already pay for. Layer 3 is the line item: with Claude Code Review at its published $15–25 average, a team merging 200 pull requests a month with one review each spends roughly $3,000–5,000 a month, and “after every push” mode multiplies that by the pushes per pull request. That is arithmetic on Anthropic’s average, not a measurement, so read the per-repository average cost in the admin settings after the first month. A claude-code-action or codex-action workflow bills only tokens plus runner minutes, and you set the model and effort, so it is the cheaper layer 3 when you can live without the managed verification step.
Set up a pull request review bot in each tool
Section titled “Set up a pull request review bot in each tool”The workflow above is the same in every tool. The wiring differs, so pick your tab.
Managed Code Review (Team and Enterprise). An organization Owner enables it at claude.ai/admin-settings/claude-code, installs the Claude GitHub App, selects repositories, and chooses a trigger: once after PR creation, after every push, or Manual. To verify setup, open a test pull request; a check run named Claude Code Review appears within a few minutes. On a Manual repository, comment @claude review for one review, or @claude review always to subscribe the pull request to later pushes. It is not available to organizations with Zero Data Retention.
Your own workflow, on any plan with an API key or OAuth token. This workflow is adapted from Anthropic’s GitHub Actions docs (we add persist-credentials: false). Save it as .github/workflows/claude-code-review.yml:
name: Code Reviewon: pull_request: types: [opened, synchronize, ready_for_review, reopened]jobs: review: runs-on: ubuntu-latest permissions: contents: read pull-requests: read issues: read id-token: write steps: - uses: actions/checkout@v6 with: fetch-depth: 1 persist-credentials: false - uses: anthropics/claude-code-action@v1 with: anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }} plugin_marketplaces: "https://github.com/anthropics/claude-code.git" plugins: "code-review@claude-code-plugins" prompt: "/code-review:code-review --comment ${{ github.repository }}/pull/${{ github.event.pull_request.number }}" claude_args: '--allowedTools "mcp__github_inline_comment__create_inline_comment"'Two lines decide where the review goes. --comment posts inline comments on the pull request; without it, findings stay in the run log. Keep the claude_args line: the action starts the inline-comment MCP server only when --allowedTools names it. The workflow skips drafts, closed pull requests and pull requests Claude already commented on. The job triggers on pull_request, not pull_request_target, so on public repositories GitHub withholds ANTHROPIC_API_KEY from fork pull requests and forks get no review. persist-credentials: false keeps the GITHUB_TOKEN out of .git/config while the agent runs on the pull request’s code. To authenticate with a subscription instead, run claude setup-token locally and pass the token as claude_code_oauth_token.
Local layers. /code-review reviews your branch and uncommitted changes as a background subagent; /code-review high widens coverage, low and medium report only the most confident findings, --fix applies findings, and --comment posts them to a pull request. For a deep pass before merge, claude ultrareview 482 reviews pull request 482 in the cloud, where “every reported finding is independently reproduced and verified”; --json prints the raw payload.
Local review (Codex CLI 0.157.1). codex review takes either one target flag or custom instructions, never both. In 0.157.1, passing --uncommitted, --base or --commit together with a prompt fails at parse time with the argument '--base <BRANCH>' cannot be used with '[PROMPT]':
# Terminal, from the repository root (Codex CLI 0.157.1)codex review --uncommittedcodex review --base maincodex review --commit 3f9c2ab# No target flag: the instructions alone drive the review, so name the scope in themcodex review "Review the changes on this branch against main. Focus on auth and SQL injection."A - in place of the prompt reads the instructions from standard input, not a diff, and the same no-target rule applies. For CI, codex exec review takes the same targets with the same restriction, plus --json, --output-schema FILE and -o FILE, which writes the final message to a file a later step can post or parse:
# CI step, after checking out the pull request with full history (Codex CLI 0.157.1)codex exec review --base origin/main -o review.mdYour own workflow. openai/codex-action@v1 installs the CLI and runs codex exec with your prompt; its README ships a complete pull request review example that posts the final message as a comment. For a review job, set permission-profile: ":read-only"; the action refuses to start if you also set sandbox, because “the profile and legacy sandbox systems do not compose”.
On GitHub (secondary). OpenAI’s docs, checked 2026-08-28, describe enabling Codex code review for a repository and then commenting @codex review (or @codex security review) on a pull request, with custom rules in AGENTS.md. Re-check the setup on OpenAI’s site before rolling it out.
Bugbot (secondary). Cursor describes Bugbot as reviewing “pull requests” and identifying “bugs, security issues, and code quality problems” (cursor.com, checked 2026-08-28). According to search extracts of Cursor’s help pages, you connect GitHub or GitLab under the Cursor dashboard’s integrations, enable Bugbot per repository, and it reviews on every pull request open and update. Comment bugbot run to re-run it. Project rules for Bugbot go in .cursor/BUGBOT.md; .cursor/rules/*.mdc files do not apply to it. Pricing is not verified here: check cursor.com before you budget.
Second model from Cursor. In Cursor, the quickest second model is a new agent chat on a model from a different vendor than the one that wrote the change, given the adversarial prompt below. For a reviewer outside Cursor, use a CLI: greptile skills install writes Greptile’s skills to .agents/skills/, which Cursor reads, and greptile review runs in Cursor’s terminal. npx skills add coderabbitai/skills installs CodeRabbit’s review skill the same way.
Graphite Agent (secondary). Graphite’s AI reviewer, formerly called Diamond, is reported as owned by Cursor since December 2025. Its stacked-PR CLI is npm @withgraphite/graphite-cli (1.8.6); there is no @graphite/cli.
Tune Claude Code Review with a REVIEW.md
Section titled “Tune Claude Code Review with a REVIEW.md”REVIEW.md sits at the repository root and reaches every agent that finds and verifies findings, so a rule there lands more reliably than the same rule in a long CLAUDE.md. Anthropic lists seven patterns; the four that cut the most noise are redefining Important, capping nits, skip rules and repository-specific checks, plus a re-review rule so a pull request converges instead of collecting new nits on every push. This version, adapted from Anthropic’s example, is a good first commit:
# Review instructions
## What Important means hereReserve Important for findings that would break behavior, leak data,or block a rollback: incorrect logic, unscoped database queries, PII inlogs or error messages, and migrations that aren't backward compatible.Style, naming and refactoring suggestions are Nit at most.
## Cap the nitsReport at most five Nits per review. If you found more, say "plus Nsimilar items" in the summary. After the first review of a PR, postImportant findings only.
## Do not report- Anything CI already enforces: lint, formatting, type errors- Generated files under `src/gen/` and any `*.lock` file
## Always check- New API routes have an integration test- Log lines don't include email addresses, user IDs or request bodies- Database queries are scoped to the caller's tenantKeep it short. Anthropic’s own guidance is that “a long REVIEW.md dilutes the rules that matter most”.
Turn the neutral check run into a merge gate
Section titled “Turn the neutral check run into a merge gate”The last line of the Claude Code Review check run’s details is a machine-readable severity count. Claude Code Review takes about 20 minutes on average, so the gate has to wait for it rather than poll once. This workflow runs when the check run completes and posts a commit status on the reviewed commit: success when the Important count is zero, failure otherwise. Save it as .github/workflows/claude-review-gate.yml on the default branch:
name: Claude review gateon: check_run: types: [completed]permissions: checks: read statuses: writejobs: gate: if: github.event.check_run.name == 'Claude Code Review' runs-on: ubuntu-latest steps: - name: Turn the severity count into a commit status env: GH_TOKEN: ${{ github.token }} REPO: ${{ github.repository }} RUN_ID: ${{ github.event.check_run.id }} SHA: ${{ github.event.check_run.head_sha }} run: | # The severity line is JSON like {"normal": 2, "nit": 1, "pre_existing": 0}; "normal" = Important important=$(gh api "repos/$REPO/check-runs/$RUN_ID" \ --jq '.output.text | split("bughunter-severity: ")[1] | split(" -->")[0] | fromjson | .normal') || important="" if [ "$important" = "0" ]; then state=success; desc="No Important findings" else state=failure; desc="Important findings: ${important:-unknown}"; fi gh api "repos/$REPO/statuses/$SHA" -f state="$state" \ -f context="claude-review-gate" -f description="$desc"Mark claude-review-gate as a required status check in branch protection or a ruleset; a check that fails but is not required blocks nothing. The workflow posts a commit status instead of relying on its own job result because, per GitHub’s docs, a check_run workflow runs against the default branch, so its job never appears on the pull request. Until the status arrives, the required check shows as expected and the merge stays blocked, which also covers a review that is still running. On a Manual repository, nothing is posted until someone comments @claude review.
A commit status belongs to one commit and does not carry forward. If the author pushes after the review, the gate on the new tip stays pending and the merge stays blocked until that commit is reviewed. That is the intended behaviour, because an unreviewed commit should not merge, but plan for it: comment @claude review on the final commit before merging, or use @claude review always (or the repository’s review-on-every-push trigger) on pull requests that change after review. Budget for the extra runs, since each review is billed.
For Greptile, greptile review status exits 0 when the commit has a completed review, 3 while one is running, 4/5 for a failed or cancelled review, and 1 when there is none or you are signed out, so a pre-push hook or CI step can require a finished review before anything merges. Treat every code other than 0 as “not reviewed”.
Add a second model: Greptile, CodeRabbit, ocr and the Codex plugin
Section titled “Add a second model: Greptile, CodeRabbit, ocr and the Codex plugin”Each of these gives you a reviewer from outside your author tool. Install commands are verified against each project’s own README or npm entry on 2026-09-26, except CodeRabbit’s installer page, which is secondary.
Codex plugin for Claude Code (OpenAI, openai/codex-plugin-cc). It runs your local Codex install from inside a Claude Code session and counts against your Codex usage limits. /codex:review is “a normal read-only Codex review” and takes no focus text; /codex:adversarial-review is steerable and challenges the design.
# Inside Claude Code/plugin marketplace add openai/codex-plugin-cc/plugin install codex@openai-codex/reload-plugins/codex:setup/codex:adversarial-review --base main challenge whether this retry design is safe under concurrent writes/codex:setup --enable-review-gate adds a Stop hook that runs a Codex review on every Claude response and blocks the stop when it finds issues. The README warns that the gate “can create a long-running Claude/Codex loop and may drain usage limits quickly”, so enable it only in a session you are watching.
Greptile CLI (npm greptile 3.6.0, Node 22 or later). greptile review compares the merge base with HEAD; --plus and --apex raise the effort.
npm install -g greptile # or: brew install greptileai/tap/greptilegreptile # signs you in on first rungreptile init # admin only: enables the repositorygreptile review -b main --instructions "focus on retry cancellation"CodeRabbit CLI. Install it from CodeRabbit’s official installer or Homebrew, not from npm. CodeRabbit’s own coderabbitai/skills SKILL.md points to https://www.coderabbit.ai/cli. The install lines below are secondary: they come from search extracts of CodeRabbit’s docs, and the installer page was not re-read on 2026-09-26. Do not install the npm name: the npm package coderabbit is a “security holding package” (version 0.0.1-security.1). --agent returns findings an agent can act on, with severities from critical to info.
curl -fsSL https://cli.coderabbit.ai/install.sh | sh # or: brew install --cask coderabbit (secondary)coderabbit auth logincoderabbit review --agent --uncommittedcoderabbit review --agent --base mainAlibaba Open Code Review (ocr), Apache-2.0, npm @alibaba-group/open-code-review 1.12.9. It pairs deterministic file selection and rule matching with an LLM agent, and works with any OpenAI- or Anthropic-compatible endpoint, so the code never leaves the provider you already use.
npm install -g @alibaba-group/open-code-reviewocr config provider && ocr config modelocr review --from main --to feature-branch --format json --output result.jsonAlibaba’s README reports higher precision and F1 than Claude Code with the same model, at about one ninth of the tokens, and says recall is lower, “a deliberate trade-off favoring precision over noise”. That is a vendor benchmark (200 pull requests from 50 repositories), not reproduced here. Choose ocr for a CI layer where noise costs you more than a missed nit.
As of 2026-09-26, anthropics/claude-code-action had 8,951 GitHub stars, openai/codex-action 1,248, openai/codex-plugin-cc 33,594 and alibaba/open-code-review 41,400; in Anthropic’s plugin directory, greptile showed 56,611 installs and coderabbit 32,361 (GitHub and claude.com/plugins, read for this site’s ecosystem catalogue). Stars measure attention, not review quality. Both plugins are in Anthropic’s official marketplace (claude-plugins-official, checked 2026-09-26):
# Inside Claude Code/plugin install greptile@claude-plugins-official/plugin install coderabbit@claude-plugins-officialThe plugins carry context cost in every session. Before you keep one, run claude plugin details greptile or claude plugin details coderabbit in a terminal, which shows the plugin’s components and projected token cost, and compare /context before and after.
Copy-paste prompts for layered AI code review
Section titled “Copy-paste prompts for layered AI code review”The third prompt produces the numbers you need for the precision metric below. In Claude Code with the GitHub CLI installed, the agent reads comments with gh api; in Codex and Cursor, use the GitHub MCP server or gh in the terminal.
How do you prove a review bot is worth its cost?
Section titled “How do you prove a review bot is worth its cost?”A review bot is itself an unverified agent until you measure it. Track four numbers per bot per month, and review them with the team:
| Metric | Definition | Act when |
|---|---|---|
| Precision | Findings that led to a code change, divided by findings raised | It falls below the level where people still read the comments; tighten REVIEW.md or BUGBOT.md, or lower the effort |
| Escapes | Production defects in code the bot reviewed, where a comment from the bot would have been enough | It rises; add the missed class to the bot’s “Always check” list |
| Cost per merged PR | Bot spend for the month divided by merged pull requests | It rises faster than merged PRs; move the bot from “every push” to “once per PR” or Manual |
| Time to first review | Minutes from the PR opening to the bot’s first comment | It exceeds your review SLA; Claude Code Review averages 20 minutes, so do not make humans wait for it |
The bots never own the decision. CI owns the gate (tests, types, lint, the severity-count check above), the bot owns the comments, and the approver named on the pull request owns the merge. The code owners in CODEOWNERS sign off on escalation classes, whatever the bots said. Canonical definitions for team metrics live in metrics frameworks for agentic engineering.
What breaks when you run AI code review bots?
Section titled “What breaks when you run AI code review bots?”The bot drowns the real bug in nits. Thirty comments, one of them important, and the team stops reading. Recovery: cap nits in REVIEW.md or BUGBOT.md, add “after the first review, post Important findings only”, and track precision. If precision keeps falling, switch that bot to Manual.
Two bots argue on the same pull request. Claude Code Review and Bugbot both comment, disagree, and the author fixes one and reopens the other. Recovery: one bot owns layer 3. Run the second model locally (layer 2), where its output goes to the author, not to the pull request.
Push-triggered reviews multiply the bill. A pull request with 12 pushes gets 12 reviews. Recovery: use “once after PR creation” or Manual, and ask for @claude review always only on high-risk pull requests. Set a monthly spend cap for Claude Code Review at claude.ai/admin-settings/usage.
Forks and drafts get no review, and nobody notices. GitHub withholds secrets from fork pull requests, claude-code-action skips drafts, and Code Review reviews forks only on a comment command. Recovery: make the review job a required check with a clear “skipped” state, and have a maintainer comment @claude review on fork pull requests before approving.
The team treats a clean bot run as approval. “The bot found nothing” becomes the merge reason. Recovery: gate on tests and the severity count, leave Copilot approvals off unless a human still signs off on escalation classes, and keep a named human approver on every pull request.
A review loop drains usage. The Codex plugin’s review gate or an autofix bot keeps finding and fixing new issues. Recovery: enable the gate only in watched sessions, stop after two fix rounds, and escalate to a human with the open findings listed.