Run a bounded PR review-fix loop
A bounded PR review-fix loop is an agent workflow that reads new review comments and failed CI checks on a pull request, classifies each one, reproduces the actionable ones, applies the smallest in-scope fix, reruns the required gates, and stops for a named human. The bound is what matters: a ledger, an attempt limit, and explicit stop conditions.
Scorecard question (Q19): How are review comments and failed CI checks resolved on pull requests?
Maximum-score answer: A bounded loop triages new findings, fixes verified causes, reruns required checks, and stops ready for named human sign-off.
This page is for developers who own agent-written pull requests, and for tech leads who set the rules those loops run under. Your pull request has 11 review comments from a review bot and two colleagues, the integration job is red, and half the comments are about code that the last push already changed. Copying each comment into a chat session works for one PR. It does not survive a week of them, and an unbounded @bot fix this loop will happily rewrite a test until it passes.
What a bounded review-fix loop gives you
Section titled “What a bounded review-fix loop gives you”- A ledger file that records every feedback item, its classification, the head commit it was checked against, and how it was resolved.
- A snapshot script that collects unresolved review threads and required checks with the GitHub CLI (
gh). - Three copy-paste prompts: triage, fix, and the stop report.
- Headless invocations for Claude Code and Codex, and the Cursor equivalent.
- A way to check the loop’s output from evidence, without rereading the whole diff.
What does the loop do on each pass?
Section titled “What does the loop do on each pass?”One pass handles every new item once, then stops. The loop never starts a redesign on its own; a finding that needs one goes back to planning.
snapshot (head SHA, threads, checks) → classify each new item → reproduce actionable defects → patch inside the approved scope → rerun affected + required gates → reply once per item, update the ledger → stop: all clear | human decision needed | attempt limit reachedThe loop’s state lives in a file on the PR branch’s working copy, not in the agent’s memory. Keep it out of the commit (add .pr-loop/ to .gitignore) and paste its summary into the PR:
{ "pr": 412, "base_sha": "9c1e4d2", "items": [ { "id": "thread:PRRT_kwDOA1b2c3", "source": "review-thread", "checked_at_sha": "4f7a0b1", "class": "defect", "repro": "npm test -- tests/billing/proration.test.ts", "status": "fixed", "attempts": 1, "fix_sha": "a81c9e0" }, { "id": "check:integration", "source": "required-check", "checked_at_sha": "4f7a0b1", "class": "needs-human", "reason": "Fix requires a schema migration, outside plan.md", "status": "escalated", "attempts": 0 } ]}Three fields do the work. The stable id (a review-thread node ID or a check name) stops the loop from answering the same comment twice. checked_at_sha marks findings stale when a newer push changed the code they cite. attempts enforces the limit per item, so one stubborn failure cannot consume the budget.
What may the agent change, and when must it stop?
Section titled “What may the agent change, and when must it stop?”Write this table into AGENTS.md or CLAUDE.md so the loop reads it on every pass.
| The agent may | The agent must stop and ask when |
|---|---|
| Read unresolved threads, check results, and failed job logs | A finding conflicts with intent.md, spec.md, or plan.md |
| Reproduce a reported failure locally | The failure does not reproduce from the available evidence |
Edit files inside the scope plan.md approves | The fix needs a migration, a new dependency, production access, or files outside scope |
| Rerun the documented test and lint commands | The same item fails after three attempts |
| Reply to a thread with the evidence for its fix | A reviewer asks for a product, security, legal, or architecture decision |
| Push to the PR branch | Anything else: merging, deploying, or changing branch protection is never in scope |
The protected branch, required reviews, and release approval stay outside the loop. A green check is evidence for the reviewer, not permission to merge.
Run one pass of the loop
Section titled “Run one pass of the loop”-
Open a draft pull request after local gates pass. Link
intent.md,spec.md, andplan.md(the artifact chain), and attach the test evidence. The evidence bundle page defines what that PR description must contain. -
Wait for results instead of busy polling.
gh pr checks 412 --watch --requiredblocks until the required checks finish. Reviews arrive when people or bots post them, so run the pass after a review is submitted, not on a timer. -
Take a snapshot. Save this as
scripts/pr-loop-snapshot.shand run it from the PR branch. It records the head commit, unresolved review threads, and required checks for this pass:#!/usr/bin/env bashset -euo pipefailPR="${1:?usage: pr-loop-snapshot.sh PR_NUMBER}"OUT=".pr-loop/pass-$(date +%Y%m%dT%H%M%S)"mkdir -p "$OUT"gh pr view "$PR" --json headRefOid,url,reviews,comments > "$OUT/pr.json"# gh exits 1 when a check failed and 8 while checks are pending; keep the snapshot either waygh pr checks "$PR" --required --json name,bucket,link,workflow > "$OUT/checks.json" || truegh api graphql -F owner='{owner}' -F repo='{repo}' -F pr="$PR" -f query='query($owner: String!, $repo: String!, $pr: Int!) {repository(owner: $owner, name: $repo) {pullRequest(number: $pr) {# first 100 threads only; paginate with pageInfo/endCursor for morereviewThreads(first: 100) {nodes { id isResolved isOutdated path linecomments(first: 20) { nodes { author { login } body } } }}}}}' > "$OUT/threads.json"echo "$OUT"isOutdatedis GitHub telling you the cited lines changed since the comment. The query reads at most 100 threads; a PR with more needspageInfo { hasNextPage endCursor }, an$endCursor: Stringvariable passed asafter:, andgh api graphql --paginate, or the rest are silently dropped. Thebucketfield groups check states intopass,fail,pending,skipping, andcancel. -
Classify before editing. Run the triage prompt below on the snapshot. Every item gets exactly one class:
defect,question,duplicate,stale,note, orneeds-human. Onlydefectitems move on. -
Reproduce each defect. The agent writes the narrowest command that fails because of the reported cause, such as one test file or one lint rule, and records it in the ledger’s
reprofield. An item that does not reproduce becomesneeds-human, not a speculative patch. -
Apply the smallest in-scope fix. One commit per item, with the item
idin the commit message. The fix never edits a test to make it pass, never lowers a coverage threshold, and never adds a lint suppression. -
Rerun the gates on the new head. First the
reprocommand, which must now pass, then every required repository gate. Save the command, exit status, and the new head SHA in the ledger. -
Reply once, then stop. Post one reply per fixed thread with the commit and the passing command. Leave threads for the reviewer to resolve. End the pass with the stop report, and start another pass only after new feedback arrives.
Prompts for the triage, fix, and stop steps
Section titled “Prompts for the triage, fix, and stop steps”Replace npm run lint and npm test with your repository’s documented gates. The prompts assume the ledger lives at .pr-loop/ledger.json. Save the three prompts as .pr-loop/prompts/triage.txt, fix.txt, and report.txt; the commands below read them from there.
Run the loop in Claude Code, Codex, or Cursor
Section titled “Run the loop in Claude Code, Codex, or Cursor”The loop contract, the ledger, and the prompts are the same in every tool. What differs is how you run a pass unattended and which review bot feeds it.
Run a pass headless from the PR branch. claude -p starts in the Manual permission mode, so name the mode and the allowed tools explicitly and cap the spend:
SNAP=$(scripts/pr-loop-snapshot.sh 412)[ -f .pr-loop/ledger.json ] || echo '{"items":[]}' > .pr-loop/ledger.json
# Triage: no Bash, so no commits and no test runs; it may write only filesclaude -p "$(cat .pr-loop/prompts/triage.txt)" \ --permission-mode manual \ --allowedTools "Read,Grep,Glob,Edit,Write" \ --max-budget-usd 2 \ --output-format json > "$SNAP/triage.json"test -z "$(git status --porcelain -- . ':!.pr-loop')" || { echo "triage edited source files; stop" >&2; false; }
# Fix: runs only after you have read the ledger diffclaude -p "$(cat .pr-loop/prompts/fix.txt)" \ --permission-mode acceptEdits \ --allowedTools "Read,Grep,Glob,Edit,Bash(npm test *),Bash(npm run lint),Bash(gh run view *),Bash(git add *),Bash(git commit *)" \ --max-budget-usd 5 \ --output-format json > "$SNAP/result.json"Start from a clean working tree. The ledger line creates an empty ledger on the first pass. The triage pass gets no Bash, so it cannot run tests or commit; it can still write files, because it has to update the ledger, which is why plan mode does not fit this step. The git status line is the actual guard: it fails if triage changed anything outside .pr-loop/. Read the ledger diff before you start the fix pass. git push is left off the allowed list, so you push after reading the stop report. Before the first push, /code-review in an interactive session runs a local review pass.
On GitHub, anthropics/claude-code-action@v1 answers a comment containing the trigger phrase (@claude by default). By default it ignores commenters without write access; leave allowed_non_write_users unset for a job that can push. The Claude Code CI/CD guide covers the workflow file. Managed Code Review (Team and Enterprise, research preview) posts findings as inline comments, and its Claude Code Review check run always completes with a neutral conclusion, so it never approves or blocks the PR. Treat its comments as one more input to this loop, not a replacement for it. Replying to one of its comments does not make it respond. If the repository runs a review after every push (@claude review always), the next run resolves the thread once the issue is fixed; otherwise comment @claude review.
Run a pass with codex exec. The permission profile :workspace (beta) lets Codex write only inside the checkout, and --output-schema forces a stop report you can parse. Save this minimal schema as .pr-loop/report.schema.json first:
{ "type": "object", "additionalProperties": false, "required": ["head_sha", "items", "verdict"], "properties": { "head_sha": { "type": "string" }, "items": { "type": "array", "items": { "type": "object", "additionalProperties": false, "required": ["id", "class", "status", "repro", "fix_sha"], "properties": { "id": { "type": "string" }, "class": { "type": "string", "enum": ["defect", "question", "duplicate", "stale", "note", "needs-human"] }, "status": { "type": "string" }, "repro": { "type": ["string", "null"] }, "fix_sha": { "type": ["string", "null"] } } } }, "verdict": { "type": "string" } }}Then run the pass:
SNAP=$(scripts/pr-loop-snapshot.sh 412)[ -f .pr-loop/ledger.json ] || echo '{"items":[]}' > .pr-loop/ledger.json
# Triage first; read the ledger diff before the fix passcodex exec -c default_permissions=":workspace" \ -o "$SNAP/triage.md" \ "$(cat .pr-loop/prompts/triage.txt)"test -z "$(git status --porcelain -- . ':!.pr-loop')" || { echo "triage edited source files; stop" >&2; false; }
cat .pr-loop/ledger.json | codex exec \ -c default_permissions=":workspace" \ --output-schema .pr-loop/report.schema.json \ -o "$SNAP/report.json" \ "$(cat .pr-loop/prompts/fix.txt)"In Codex CLI 0.157.1, the :workspace profile mounts .git read-only, so step 4 of the fix prompt cannot commit: the pass edits files and reports, and fix_sha stays null. Commit the fixes yourself after reading the report, one commit per item id, and fill in fix_sha. The triage pass runs under :workspace too, because it writes the ledger; :read-only would block that write. The git status line after it fails if triage touched anything outside .pr-loop/. Piped stdin is appended to the prompt as a <stdin> block. Before pushing, codex review --base main runs a second, independent review of the branch against main. In Codex CLI 0.157.1, --base cannot be combined with custom review instructions; the CLI rejects the pair, so put review focus in AGENTS.md instead.
For red CI on pushes to your own branches, the Codex CI/CD guide shows openai/codex-action@v1 split into a read-only trigger, a job that can write only the checkout, and a job that opens a draft PR. Use that split rather than giving one job both the secret and push rights.
Bugbot reviews pull requests and posts findings; treat its comments as one more feedback source in the snapshot. Run the triage and fix prompts in Cursor’s agent on the PR branch: run the triage prompt first and review the ledger before you allow any edits. For an unattended pass, the Cursor CLI’s print mode (-p) runs the same prompts headless.
Cursor’s Bugbot and PR Routing & Approval features were last verified on 2026-08-28; check the current setup on the Bugbot page before enabling write access for any automated fix.
If the agent should read PR state through a tool connection instead of gh, GitHub’s GitHub MCP server exposes pull requests and Actions logs, and its URL accepts a /readonly suffix per toolset. In Claude Code, claude mcp add --transport http github https://api.githubcopilot.com/mcp/ -H "Authorization: Bearer $GITHUB_PAT_TOKEN" adds it; the shell expands the token when you run the command. In Codex, codex mcp add github --url https://api.githubcopilot.com/mcp/ --bearer-token-env-var GITHUB_PAT_TOKEN keeps the token in an environment variable. The gh CLI costs fewer tokens for the same reads.
How do you check the loop’s output without reading the whole diff?
Section titled “How do you check the loop’s output without reading the whole diff?”Verify the pass from its evidence, and read code only where the evidence points.
-
Every fix has a before and after. Each
fixeditem names areprocommand. Run one or two yourself on the fix commit’s parent and on the head: it should fail, then pass. -
Required checks are green on the head the report names. Compare the report’s SHA with
gh pr view 412 --json headRefOid. A report about an older SHA is stale. -
Nothing left the approved scope. List files changed by the loop’s commits and compare them with
plan.md:Terminal window git log --format=%H --grep='^fix(review):' origin/main..HEAD \| xargs -r git show --name-only --format= | sort -uAny path outside the plan’s scope, any test file, or any CI config in that list gets read line by line.
-
The oracle did not move. A fix that touched a test, a snapshot, or a threshold is a
needs-humanitem by definition, even if the agent classified it as a defect. -
A named human signs off. The loop ends at READY FOR REVIEW. The reviewer approves, and the release follows the production approval gate.
For the full triage protocol a reviewer follows on an agent-written PR, see reviewing an agent’s pull request without reading every line.
What breaks in a PR review-fix loop, and how do you recover?
Section titled “What breaks in a PR review-fix loop, and how do you recover?”The loop answers the same comment on every pass. The ledger has no stable id, or the agent keys items by comment text. Recover: key by thread node ID or check name, delete the duplicate replies, and rerun triage against the current ledger.
A fix targets code the previous push already changed. The item was checked against an older head. Recover: mark items whose checked_at_sha differs from the current headRefOid as stale, take a new snapshot, and triage again.
A flaky test gets “fixed”. The repro passed on its first run and the agent patched anyway, or it added a retry. Recover: require the repro to fail before any edit, revert the commit, and route the test to your flaky-test owner as needs-human.
The checks stay red after three attempts. The cause is outside the scope the agent was given, or the environment differs from CI. Recover: stop the item, attach the last failed log (gh run view RUN_ID --log-failed), and escalate. Do not raise the attempt limit for that item.
A patch drifts from the plan. The scope check shows files outside plan.md, or the fix introduces a new abstraction. Recover: revert the commit and return the item to planning. A review comment that needs a design change is a new plan, not a fix.
Green checks trigger a merge or deploy. Someone wired auto-merge to the loop’s result. Recover: remove auto-merge from the loop’s workflow, keep merge behind branch protection and required human review, and treat rollback as the safety net.
Checklist for a Q19-ready review-fix loop
Section titled “Checklist for a Q19-ready review-fix loop”- Every processed item has a stable
id, one class, and the head SHA it was checked against. - Every fixed defect has a repro command that failed before the fix and passes after it.
- Fixes stay inside
plan.md’s scope and never weaken a test, threshold, or lint rule. - Required checks ran on the head SHA named in the stop report.
- Each item has an attempt limit, and the loop stops with a report instead of retrying.
- The loop cannot merge or deploy, and a named human signs off.
Where to go next with PR review and release
Section titled “Where to go next with PR review and release”The loop consumes findings from layered PR review and hands a clean PR to the release gate. The developer answer key maps every scorecard question to its page.