Skip to content

Code Review Inside the Agent: Skills and Plugins

Code review skills and plugins review a diff inside the coding agent before a pull request reaches a human. Claude Code 2.1.283 ships a built-in /code-review; plugins and skills add PR comments, specialist reviewers, spec checks, security review and a second model. They find likely defects but approve nothing: tests, CI and a named human still gate the merge.

The agent finishes a 600-line branch, says the tests pass, and you open the pull request knowing nobody will read all of it closely. You have also installed three things called “code review” and no longer know which one /review runs. This page is for developers who want the agent’s work checked before they push, and for tech leads who want one review chain for the whole team instead of eight personal ones.

What you get from this page on in-agent code review

Section titled “What you get from this page on in-agent code review”
  • A decision table for eight review tools
  • A worked run of /code-review high on a feature branch, with the flags that change its behaviour
  • Install commands for Claude Code, Codex and Cursor, checked on 2026-09-26 against Claude Code 2.1.283, codex-cli 0.157.1 and the skills CLI 1.7.0
  • A three-stage workflow: self-review, a second model, then a PR bot, and what each stage hands to the next
  • Copy-paste prompts, a way to measure whether reviewers catch real bugs, and the traps that silently skip a review

Which code review skill or plugin should you use?

Section titled “Which code review skill or plugin should you use?”

“Code review” names at least eight different tools. They differ in what they read, where findings go and whether they edit your code.

ToolWho ships itWhat it checksWhere findings goRuns in
/code-review (built-in, alias /review)Anthropic, bundledCorrectness bugs in the diff, plus reuse, simplification and efficiency cleanups where the model’s review recipe covers themYour session; --fix applies them, --comment posts them to the PRClaude Code
code-review plugin, /code-review:code-reviewAnthropic, claude-plugins-officialA GitHub PR: CLAUDE.md compliance, obvious bugs, git history, earlier PR comments, code comments; drops findings scored below 80One PR comment posted through ghClaude Code
pr-review-toolkit, /pr-review-toolkit:review-prAnthropic, claude-plugins-officialSix specialist subagents: comments, tests, silent failures, type design, general quality, simplificationYour session, grouped Critical / Important / SuggestionClaude Code
Pocock code-reviewmattpocock/skillsTwo axes in parallel subagents: your documented standards, and whether the diff matches the spec or issueYour session, under ## Standards and ## SpecAny agent that reads skills
requesting-code-review / receiving-code-reviewobra/superpowersRequesting: a reviewer subagent checks each task against the plan. Receiving: the agent verifies feedback before acting on itYour sessionAny agent that reads skills
differential-reviewtrailofbits/skillsSecurity review of a diff with git blame, blast radius (callers) and test coverage of changed codeA Markdown report fileClaude Code, Codex
CodeRabbit code-review skillcoderabbitai/skillsRuns the CodeRabbit CLI on your changes; findings carry severities and fix instructionsYour session; the diff is sent to the CodeRabbit APIAny agent that reads skills
codex plugin, /codex:reviewOpenAI, openai/codex-plugin-ccA read-only Codex review of uncommitted changes or a branchYour Claude Code sessionClaude Code

gstack’s /review is a ninth option with a name collision; see the traps box below. For the managed PR bots (Claude Code Review, @codex review, Cursor Bugbot, CodeRabbit’s GitHub app), see AI code review bots.

Popularity, as of 2026-09-26. Installs in the claude.com plugin directory: code-review 438,525, pr-review-toolkit 114,856, coderabbit 32,361. All-time skills.sh installs, from the third-party LinklyAI/best-skills snapshot of 2026-09-26 (a secondary source; re-read skills.sh before you quote them): Pocock code-review 617,608, requesting-code-review 238,401, receiving-code-review 202,093. GitHub stars on the same date: garrytan/gstack 134.2k, openai/codex-plugin-cc 33.6k, trailofbits/skills 7.3k.

If you need to…UseWhy this one
check your own branch before you pushbuilt-in /code-reviewNo install, runs in the background, and --fix applies what it finds
check that the diff does what the ticket askedPocock code-reviewIt reports spec compliance as its own axis, apart from standards, so scope creep and missing requirements are not buried in style notes
review a change that touches auth, payments or input parsingdifferential-reviewIt classifies by risk, not size, and writes a report you can attach to the PR
get a second opinion from a different model family/codex:review (Claude Code) or the CodeRabbit skillThe reviewer does not start from the authoring model’s reasoning or its session history
run review between tasks of a long plan, or push back on a bad commentSuperpowers requesting-code-review / receiving-code-reviewThe first runs after every task, not once at the end; the second makes the agent verify each claim before agreeing
post a filtered review on a teammate’s PRcode-review pluginIts confidence filter keeps low-certainty nits off the PR

Pick one reviewer per stage: two skills firing on one request give two procedures and two reports.

How do you run /code-review high on a branch?

Section titled “How do you run /code-review high on a branch?”

Learn the built-in command first; every Claude Code session already has it. The full syntax in Claude Code 2.1.283 is /code-review [low|medium|high|xhigh|max|ultra] [--fix] [--comment] [pr#|branch|path].

  1. Finish the work on a branch. With no target, the review reads the branch’s commits ahead of its upstream plus any uncommitted changes. An empty diff has nothing to report.

  2. Run the review at high effort. In the Claude Code prompt:

    /code-review high

    At low and medium the review reports only the findings it is most confident in. high through max broaden coverage and may include findings it is less sure about. Use high on a feature branch and medium on a small fix.

  3. Keep working. The review runs as a background subagent with its own context window; the findings arrive in the conversation when it completes.

  4. Read the findings as claims, not facts. Each names a file location and a one-sentence summary. Have Claude reproduce each correctness finding with a failing test before fixing it (the prompt below does this).

  5. Apply or post. /code-review high --fix applies the findings to your working tree. /code-review high 1234 --comment reviews pull request 1234 and posts the findings as inline comments. A target can also be a file path or a branch name; the documented forms in 2.1.283 are pr#, branch and path.

Easy to miss:

  • The level is remembered. When you type no level, the review reuses the last level you typed, even from an earlier session, and prints a notice such as Reusing high effort, the level you typed last time.
  • Commit before --fix. It applies every finding to the working tree in one pass, so a clean commit gives you a single git restore back to where you were.
  • The local review reads CLAUDE.md, not REVIEW.md. REVIEW.md tunes only the managed Code Review service.
  • /simplify is a different job. It applies cleanups without hunting for bugs.

/code-review ultra escalates to ultrareview, a multi-agent review in Anthropic’s cloud. Pro and Max get three free runs, after which a run typically costs $5 to $25 in usage credits (Anthropic’s ultrareview docs, checked 2026-09-26). It is not available on Amazon Bedrock, Google Cloud’s Agent Platform (formerly Vertex AI), Microsoft Foundry or to organisations with Zero Data Retention; there, /code-review ultra runs a local review instead. From the shell, claude ultrareview [target] does the same and prints the findings.

To stop Claude or a scheduled task from starting the review on its own while you keep the command for yourself, add this to ~/.claude/settings.json:

{
"skillOverrides": {
"code-review": "user-invocable-only"
}
}

How do you install review skills in Claude Code, Codex and Cursor?

Section titled “How do you install review skills in Claude Code, Codex and Cursor?”

The built-in /code-review needs no install. Everything else comes as a plugin (with updates and namespaced commands) or as a portable skill from the skills CLI (npm skills 1.7.0).

Terminal window
# Anthropic's plugins (official marketplace, added on first interactive start)
claude plugin install code-review@claude-plugins-official --scope project
claude plugin install pr-review-toolkit@claude-plugins-official --scope project
# Second model: OpenAI's Codex plugin lives in OpenAI's own marketplace
claude plugin marketplace add openai/codex-plugin-cc
claude plugin install codex@openai-codex
# Security review of diffs
claude plugin marketplace add trailofbits/skills
claude plugin install differential-review@trailofbits
# Pocock's skills as a plugin, so his code-review gets a namespaced command
claude plugin install mattpocock-skills@claude-plugins-official
# Superpowers review skills as portable skills
npx skills add obra/superpowers --skill requesting-code-review receiving-code-review -a claude-code -y

Plugin commands are namespaced: /code-review:code-review, /pr-review-toolkit:review-pr tests errors, /mattpocock-skills:code-review, /codex:review, /differential-review:diff-review. Take Pocock’s skills through the plugin, because a portable copy named code-review competes with the built-in for the bare name. After installing the Codex plugin, run /codex:setup; it checks that Codex is installed and logged in. Codex usage counts against your Codex plan or API key.

Commit .claude/settings.json (project-scope plugins) and .agents/skills/, .claude/skills/ and skills-lock.json (portable skills). A review skill runs with the agent’s permissions, so read it first; skill supply-chain security has the checklist.

How do you chain self-review, a second model and a PR bot?

Section titled “How do you chain self-review, a second model and a PR bot?”

One reviewer catches what its model and prompt make it look for. This workflow uses three reviewers that fail differently, cheapest fixes first.

  1. Self-review in the authoring session. Run /code-review high (Claude Code) or /review (Codex) and fix only the findings you can reproduce, with the first prompt above. In Cursor, run claude in the integrated terminal and type /code-review high there, or ask the agent to review the branch diff against main with the same prompt. If the change touches auth, money or parsing, also ask for “a differential security review of this branch against main”; the differential-review skill writes a Markdown report, which you keep. The command form, /differential-review:diff-review, takes a PR URL, a commit SHA or a diff file, with --baseline <ref>.

  2. Check the diff against the spec. Run Pocock’s code-review with main as the fixed point (/mattpocock-skills:code-review in Claude Code, $code-review in Codex, “use the code-review skill” in Cursor). The Spec axis lists missing requirements, unrequested behaviour and requirements that look implemented but wrong, each with the spec line quoted. Resolve every Spec finding before you open the PR; after merge, each one is a new ticket.

  3. Ask a second model. The authoring model reviewing its own work shares its blind spots.

    /codex:review --base main --background
    /codex:status
    /codex:result

    /codex:review is read-only and takes no focus text. To challenge a specific decision, use /codex:adversarial-review --base main challenge whether the retry logic is safe under partial failure.

  4. Triage the combined findings with receiving-code-review. Paste the second model’s output into the authoring session with the prompt below. The agent verifies each claim against the code before it changes anything and pushes back on wrong ones with reasons.

  5. Open the PR and let the bot review it. A PR bot (Claude Code Review via @claude review, @codex review, Bugbot or CodeRabbit’s app) reviews the final diff with repository context. Anthropic’s managed Code Review is a research preview for Team and Enterprise that “averages $15-25” per review (Anthropic docs, checked 2026-09-26); setup for each bot is in AI code review bots.

  6. Hand the human an evidence bundle, not a diff. The PR description lists which reviewers ran, which findings were fixed with a reproducing test, and which were rejected and why. The human then follows reviewing an agent’s pull request.

In a long Superpowers plan, requesting-code-review moves step 1 inside the loop: after each task a reviewer subagent checks it against the task’s requirements, and the agent fixes Critical issues at once and Important ones before the next task. See Superpowers.

How do you know your review agents catch real bugs?

Section titled “How do you know your review agents catch real bugs?”

A review that reports nothing can mean a clean diff or a reviewer that did not look. Measure reviewers like a test suite: plant defects and count what comes back.

  • Seed known bugs. On a throwaway branch, plant three defects of the kind your incidents come from, run each reviewer and record which it names (the prompt below sets this up). Repeat when you change model, effort level or skill version.
  • Track the outcome of every finding. Label each one fixed, rejected or not reproduced. A reviewer whose findings are mostly rejected costs reviewer attention; lower its effort level or drop it.
  • Keep CI as the gate. Review agents advise; the tests, type checks, linters and security scanners in CI decide. A finding that matters becomes a test, so the same bug cannot return. Keep tests out of the agent’s reach with protecting the oracle.
  • Check that the review ran. Look for the report itself: the differential-review Markdown file, the CodeRabbit completion status, the findings message from /code-review. “No findings” without a report is not a result.

Who signs off. The developer who opens the PR owns the self-review and the reproduction evidence. The tech lead owns the team’s reviewer set, effort levels and canary results, and changes them through a pull request. A human approves the merge; no review agent here approves on its own.

Skills cost their name and description in every session and their body only when they fire; plugins can cost more up front. Measured on 2026-09-26:

What loadsWhenSize
pr-review-toolkit pluginEvery sessionabout 2,033 tokens (claude plugin details, Claude Code 2.1.283)
mattpocock-skills plugin (all of Pocock’s skills)Every sessionabout 1,609 tokens (claude plugin details, Claude Code 2.1.283)
Pocock code-review SKILL.mdWhen it fires6,589 bytes
Superpowers requesting-code-review + its code-reviewer.md templateWhen it fires2,977 + 6,449 bytes
Superpowers receiving-code-review SKILL.mdWhen it fires6,203 bytes
CodeRabbit code-review SKILL.mdWhen it fires7,518 bytes
Trail of Bits differential-review SKILL.md (plus four reference files it opens as needed)When it fires7,439 bytes (about 33 KB with all references)
gstack review/SKILL.mdWhen it firesabout 73 KB

At roughly 3 to 4 bytes per token (Claude 4.7 and later tokenizers produce about 30% more tokens for the same text), most review skills cost about 1,600 to 3,200 tokens when they fire, and gstack’s about 18,000 to 24,000. The built-in /code-review reads the diff in its own subagent context. Check an install with claude plugin details PLUGIN_NAME and /context.

What breaks when review runs inside the agent?

Section titled “What breaks when review runs inside the agent?”

The review finds nothing on a large change. Usually the target was wrong: the branch had no upstream, the changes were already pushed, or the review read only uncommitted files. Recovery: name the branch as the target, such as /code-review high feature/rate-limit, and check that the findings message names the files you expected. A git range is not a documented target form in 2.1.283.

Two reviewers fire on one request. “Review this” triggers several review skills at once, or one you did not expect. Recovery: invoke by exact name, keep one general reviewer per repository, and remove the others with npx skills remove --skill SKILL_NAME or claude plugin uninstall PLUGIN_NAME.

--fix changes more than you expected. It applies every finding at once, including the ones you would not have reproduced. Recovery: git diff to see the edits, git restore or git revert to undo them, and commit before every --fix run.

The agent agrees with every comment. It applies a wrong PR-bot suggestion and breaks untested behaviour. Recovery: route all review feedback through receiving-code-review with the second prompt above, and require a failing test before any fix.

The review loop does not end. The Codex plugin’s optional review gate runs from a Stop hook and blocks Claude from stopping while it finds issues; OpenAI’s README warns it “may drain usage limits quickly”. Recovery: /codex:setup --disable-review-gate (the flag in the openai/codex-plugin-cc README, checked 2026-09-26), and enable it only in sessions you watch.

Cloud review is unavailable. On Bedrock, Google Cloud, Foundry or under Zero Data Retention, /code-review ultra quietly runs a local review. Recovery: treat it as a local review and do not record it as an ultrareview in the evidence bundle.

The code-review plugin skips your PR. It drops closed, draft, trivial or already-reviewed PRs, and posts nothing when no finding reaches 80. Recovery: mark the PR ready for review, or use the built-in /code-review to see low-confidence findings too.

Where to go next with in-agent code review

Section titled “Where to go next with in-agent code review”

For tool-specific setups, see review automation in Claude Code, code review in Codex and code review in Cursor. For the wider ranking of practice skills, see the most-installed development-practice skills.