Skip to content

Run a bounded PR review-fix loop

A bounded PR review-fix loop is an agent workflow that reads new review comments and failed CI checks on a pull request, classifies each one, reproduces the actionable ones, applies the smallest in-scope fix, reruns the required gates, and stops for a named human. The bound is what matters: a ledger, an attempt limit, and explicit stop conditions.

Scorecard question (Q19): How are review comments and failed CI checks resolved on pull requests?

Maximum-score answer: A bounded loop triages new findings, fixes verified causes, reruns required checks, and stops ready for named human sign-off.

This page is for developers who own agent-written pull requests, and for tech leads who set the rules those loops run under. Your pull request has 11 review comments from a review bot and two colleagues, the integration job is red, and half the comments are about code that the last push already changed. Copying each comment into a chat session works for one PR. It does not survive a week of them, and an unbounded @bot fix this loop will happily rewrite a test until it passes.

  • A ledger file that records every feedback item, its classification, the head commit it was checked against, and how it was resolved.
  • A snapshot script that collects unresolved review threads and required checks with the GitHub CLI (gh).
  • Three copy-paste prompts: triage, fix, and the stop report.
  • Headless invocations for Claude Code and Codex, and the Cursor equivalent.
  • A way to check the loop’s output from evidence, without rereading the whole diff.

One pass handles every new item once, then stops. The loop never starts a redesign on its own; a finding that needs one goes back to planning.

snapshot (head SHA, threads, checks)
→ classify each new item
→ reproduce actionable defects
→ patch inside the approved scope
→ rerun affected + required gates
→ reply once per item, update the ledger
→ stop: all clear | human decision needed | attempt limit reached

The loop’s state lives in a file on the PR branch’s working copy, not in the agent’s memory. Keep it out of the commit (add .pr-loop/ to .gitignore) and paste its summary into the PR:

{
"pr": 412,
"base_sha": "9c1e4d2",
"items": [
{
"id": "thread:PRRT_kwDOA1b2c3",
"source": "review-thread",
"checked_at_sha": "4f7a0b1",
"class": "defect",
"repro": "npm test -- tests/billing/proration.test.ts",
"status": "fixed",
"attempts": 1,
"fix_sha": "a81c9e0"
},
{
"id": "check:integration",
"source": "required-check",
"checked_at_sha": "4f7a0b1",
"class": "needs-human",
"reason": "Fix requires a schema migration, outside plan.md",
"status": "escalated",
"attempts": 0
}
]
}

Three fields do the work. The stable id (a review-thread node ID or a check name) stops the loop from answering the same comment twice. checked_at_sha marks findings stale when a newer push changed the code they cite. attempts enforces the limit per item, so one stubborn failure cannot consume the budget.

What may the agent change, and when must it stop?

Section titled “What may the agent change, and when must it stop?”

Write this table into AGENTS.md or CLAUDE.md so the loop reads it on every pass.

The agent mayThe agent must stop and ask when
Read unresolved threads, check results, and failed job logsA finding conflicts with intent.md, spec.md, or plan.md
Reproduce a reported failure locallyThe failure does not reproduce from the available evidence
Edit files inside the scope plan.md approvesThe fix needs a migration, a new dependency, production access, or files outside scope
Rerun the documented test and lint commandsThe same item fails after three attempts
Reply to a thread with the evidence for its fixA reviewer asks for a product, security, legal, or architecture decision
Push to the PR branchAnything else: merging, deploying, or changing branch protection is never in scope

The protected branch, required reviews, and release approval stay outside the loop. A green check is evidence for the reviewer, not permission to merge.

  1. Open a draft pull request after local gates pass. Link intent.md, spec.md, and plan.md (the artifact chain), and attach the test evidence. The evidence bundle page defines what that PR description must contain.

  2. Wait for results instead of busy polling. gh pr checks 412 --watch --required blocks until the required checks finish. Reviews arrive when people or bots post them, so run the pass after a review is submitted, not on a timer.

  3. Take a snapshot. Save this as scripts/pr-loop-snapshot.sh and run it from the PR branch. It records the head commit, unresolved review threads, and required checks for this pass:

    #!/usr/bin/env bash
    set -euo pipefail
    PR="${1:?usage: pr-loop-snapshot.sh PR_NUMBER}"
    OUT=".pr-loop/pass-$(date +%Y%m%dT%H%M%S)"
    mkdir -p "$OUT"
    gh pr view "$PR" --json headRefOid,url,reviews,comments > "$OUT/pr.json"
    # gh exits 1 when a check failed and 8 while checks are pending; keep the snapshot either way
    gh pr checks "$PR" --required --json name,bucket,link,workflow > "$OUT/checks.json" || true
    gh api graphql -F owner='{owner}' -F repo='{repo}' -F pr="$PR" -f query='
    query($owner: String!, $repo: String!, $pr: Int!) {
    repository(owner: $owner, name: $repo) {
    pullRequest(number: $pr) {
    # first 100 threads only; paginate with pageInfo/endCursor for more
    reviewThreads(first: 100) {
    nodes { id isResolved isOutdated path line
    comments(first: 20) { nodes { author { login } body } } }
    }
    }
    }
    }' > "$OUT/threads.json"
    echo "$OUT"

    isOutdated is GitHub telling you the cited lines changed since the comment. The query reads at most 100 threads; a PR with more needs pageInfo { hasNextPage endCursor }, an $endCursor: String variable passed as after:, and gh api graphql --paginate, or the rest are silently dropped. The bucket field groups check states into pass, fail, pending, skipping, and cancel.

  4. Classify before editing. Run the triage prompt below on the snapshot. Every item gets exactly one class: defect, question, duplicate, stale, note, or needs-human. Only defect items move on.

  5. Reproduce each defect. The agent writes the narrowest command that fails because of the reported cause, such as one test file or one lint rule, and records it in the ledger’s repro field. An item that does not reproduce becomes needs-human, not a speculative patch.

  6. Apply the smallest in-scope fix. One commit per item, with the item id in the commit message. The fix never edits a test to make it pass, never lowers a coverage threshold, and never adds a lint suppression.

  7. Rerun the gates on the new head. First the repro command, which must now pass, then every required repository gate. Save the command, exit status, and the new head SHA in the ledger.

  8. Reply once, then stop. Post one reply per fixed thread with the commit and the passing command. Leave threads for the reviewer to resolve. End the pass with the stop report, and start another pass only after new feedback arrives.

Prompts for the triage, fix, and stop steps

Section titled “Prompts for the triage, fix, and stop steps”

Replace npm run lint and npm test with your repository’s documented gates. The prompts assume the ledger lives at .pr-loop/ledger.json. Save the three prompts as .pr-loop/prompts/triage.txt, fix.txt, and report.txt; the commands below read them from there.

Run the loop in Claude Code, Codex, or Cursor

Section titled “Run the loop in Claude Code, Codex, or Cursor”

The loop contract, the ledger, and the prompts are the same in every tool. What differs is how you run a pass unattended and which review bot feeds it.

Run a pass headless from the PR branch. claude -p starts in the Manual permission mode, so name the mode and the allowed tools explicitly and cap the spend:

Terminal window
SNAP=$(scripts/pr-loop-snapshot.sh 412)
[ -f .pr-loop/ledger.json ] || echo '{"items":[]}' > .pr-loop/ledger.json
# Triage: no Bash, so no commits and no test runs; it may write only files
claude -p "$(cat .pr-loop/prompts/triage.txt)" \
--permission-mode manual \
--allowedTools "Read,Grep,Glob,Edit,Write" \
--max-budget-usd 2 \
--output-format json > "$SNAP/triage.json"
test -z "$(git status --porcelain -- . ':!.pr-loop')" || { echo "triage edited source files; stop" >&2; false; }
# Fix: runs only after you have read the ledger diff
claude -p "$(cat .pr-loop/prompts/fix.txt)" \
--permission-mode acceptEdits \
--allowedTools "Read,Grep,Glob,Edit,Bash(npm test *),Bash(npm run lint),Bash(gh run view *),Bash(git add *),Bash(git commit *)" \
--max-budget-usd 5 \
--output-format json > "$SNAP/result.json"

Start from a clean working tree. The ledger line creates an empty ledger on the first pass. The triage pass gets no Bash, so it cannot run tests or commit; it can still write files, because it has to update the ledger, which is why plan mode does not fit this step. The git status line is the actual guard: it fails if triage changed anything outside .pr-loop/. Read the ledger diff before you start the fix pass. git push is left off the allowed list, so you push after reading the stop report. Before the first push, /code-review in an interactive session runs a local review pass.

On GitHub, anthropics/claude-code-action@v1 answers a comment containing the trigger phrase (@claude by default). By default it ignores commenters without write access; leave allowed_non_write_users unset for a job that can push. The Claude Code CI/CD guide covers the workflow file. Managed Code Review (Team and Enterprise, research preview) posts findings as inline comments, and its Claude Code Review check run always completes with a neutral conclusion, so it never approves or blocks the PR. Treat its comments as one more input to this loop, not a replacement for it. Replying to one of its comments does not make it respond. If the repository runs a review after every push (@claude review always), the next run resolves the thread once the issue is fixed; otherwise comment @claude review.

If the agent should read PR state through a tool connection instead of gh, GitHub’s GitHub MCP server exposes pull requests and Actions logs, and its URL accepts a /readonly suffix per toolset. In Claude Code, claude mcp add --transport http github https://api.githubcopilot.com/mcp/ -H "Authorization: Bearer $GITHUB_PAT_TOKEN" adds it; the shell expands the token when you run the command. In Codex, codex mcp add github --url https://api.githubcopilot.com/mcp/ --bearer-token-env-var GITHUB_PAT_TOKEN keeps the token in an environment variable. The gh CLI costs fewer tokens for the same reads.

How do you check the loop’s output without reading the whole diff?

Section titled “How do you check the loop’s output without reading the whole diff?”

Verify the pass from its evidence, and read code only where the evidence points.

  • Every fix has a before and after. Each fixed item names a repro command. Run one or two yourself on the fix commit’s parent and on the head: it should fail, then pass.

  • Required checks are green on the head the report names. Compare the report’s SHA with gh pr view 412 --json headRefOid. A report about an older SHA is stale.

  • Nothing left the approved scope. List files changed by the loop’s commits and compare them with plan.md:

    Terminal window
    git log --format=%H --grep='^fix(review):' origin/main..HEAD \
    | xargs -r git show --name-only --format= | sort -u

    Any path outside the plan’s scope, any test file, or any CI config in that list gets read line by line.

  • The oracle did not move. A fix that touched a test, a snapshot, or a threshold is a needs-human item by definition, even if the agent classified it as a defect.

  • A named human signs off. The loop ends at READY FOR REVIEW. The reviewer approves, and the release follows the production approval gate.

For the full triage protocol a reviewer follows on an agent-written PR, see reviewing an agent’s pull request without reading every line.

What breaks in a PR review-fix loop, and how do you recover?

Section titled “What breaks in a PR review-fix loop, and how do you recover?”

The loop answers the same comment on every pass. The ledger has no stable id, or the agent keys items by comment text. Recover: key by thread node ID or check name, delete the duplicate replies, and rerun triage against the current ledger.

A fix targets code the previous push already changed. The item was checked against an older head. Recover: mark items whose checked_at_sha differs from the current headRefOid as stale, take a new snapshot, and triage again.

A flaky test gets “fixed”. The repro passed on its first run and the agent patched anyway, or it added a retry. Recover: require the repro to fail before any edit, revert the commit, and route the test to your flaky-test owner as needs-human.

The checks stay red after three attempts. The cause is outside the scope the agent was given, or the environment differs from CI. Recover: stop the item, attach the last failed log (gh run view RUN_ID --log-failed), and escalate. Do not raise the attempt limit for that item.

A patch drifts from the plan. The scope check shows files outside plan.md, or the fix introduces a new abstraction. Recover: revert the commit and return the item to planning. A review comment that needs a design change is a new plan, not a fix.

Green checks trigger a merge or deploy. Someone wired auto-merge to the loop’s result. Recover: remove auto-merge from the loop’s workflow, keep merge behind branch protection and required human review, and treat rollback as the safety net.

  • Every processed item has a stable id, one class, and the head SHA it was checked against.
  • Every fixed defect has a repro command that failed before the fix and passes after it.
  • Fixes stay inside plan.md’s scope and never weaken a test, threshold, or lint rule.
  • Required checks ran on the head SHA named in the stop report.
  • Each item has an attempt limit, and the loop stops with a report instead of retrying.
  • The loop cannot merge or deploy, and a named human signs off.

Where to go next with PR review and release

Section titled “Where to go next with PR review and release”

The loop consumes findings from layered PR review and hands a clean PR to the release gate. The developer answer key maps every scorecard question to its page.