From issue to pull request with no hands on the keyboard
An issue-to-PR pipeline turns a labelled GitHub issue into a draft pull request without anyone typing: an agent (Claude Code, Codex or Cursor) implements the issue in an isolated run, CI re-runs the gates independently, a fix loop capped at three attempts repairs failures, and an evidence bundle leaves the pull request ready for a human merge.
You have a backlog of well-specified issues: rename a field, add a filter to an endpoint, fix a bug with a known reproduction. Each one can eat most of an hour of setup, typing and waiting for CI, and the agent could do the typing. The problem is everything around the typing. Somebody has to start the run, babysit it, push the branch, open the pull request, rerun a flaky check, and decide whether the green tick means anything.
This tutorial builds the version where nobody does those things by hand. It is written for the developer who wires it up and the tech lead who decides which issues may enter it. Every file is complete and meant to be copied into your repository.
What you get from an issue-to-PR pipeline
Section titled “What you get from an issue-to-PR pipeline”- A label,
agent:ready, that starts an agent run on GitHub-hosted runners or in a Cursor Cloud Agent, never on a laptop. - A task contract: an issue form, a prompt the run always uses, and one script that defines “done”.
- A split of authority: the agent edits files, the workflow commits, CI judges and a human merges.
- A review-fix loop that stops after three attempts and hands the pull request to a named person.
- An evidence bundle comment on every agent pull request, produced by CI rather than by the agent.
- Five numbers that tell a tech lead whether the pipeline produces mergeable work.
How does the issue-to-PR pipeline fit together?
Section titled “How does the issue-to-PR pipeline fit together?”The pipeline has six stages. The column that matters most is the third: at no stage does the component that writes the code also decide that the code is good.
| Stage | What runs | Who holds authority | What stops it |
|---|---|---|---|
| 1. Label | A maintainer applies agent:ready | A human with write access | The label is missing, or the labeller lacks write access |
| 2. Isolated run | The agent on a fresh runner or Cloud Agent VM | The agent may edit files only | Turn cap, dollar cap, job timeout, or a stop condition in the prompt |
| 3. Commit and draft PR | A workflow job with a GitHub App token | The workflow, never the agent | An empty patch labels the issue agent:needs-human |
| 4. Gates | The base branch’s scripts/agent-gates.sh, run in CI against the pull request’s code | CI, in a job with no secrets and no write token | Any failing gate, including a touched protected path |
| 5. Review-fix loop | The agent again, with gate logs and review comments as input | Capped at three Agent-Fix: commits | The cap, or a stop condition, labels the PR agent:needs-human |
| 6. Merge | A human reads the evidence bundle and approves | The issue owner, plus CODEOWNERS for sensitive paths | A ruleset that requires one approval and the evidence check |
The design choice behind stages 3 and 4 is worth stating once: no job that runs agent-written code holds a token that can write. The agent job checks out without stored git credentials and hands the Claude Code action the job’s own GITHUB_TOKEN, which is read-only under contents: read. It passes a patch to a second job, and only that job mints a write token. In CI, the job that runs the pull request’s code has no secrets, and the job that comments and marks the pull request ready never checks out or runs that code. A prompt-injected agent can therefore damage its own working copy and spend its budget, but it cannot push to a branch, comment as the bot, or touch another repository. It does hold the model API key while it runs code it has just written, so give the pipeline its own key with a spending limit.
Which tool should run the agent step?
Section titled “Which tool should run the agent step?”Only the agent step differs between the three tools. The task contract, gates, evidence and loop are identical, so you can switch tools or run two side by side for comparison.
| Claude Code | Codex | Cursor | |
|---|---|---|---|
| Runs in | anthropics/claude-code-action@v1 on a GitHub-hosted runner | openai/codex-action@v1 (tag v1.12) on a GitHub-hosted runner | A Cloud Agent, which Cursor says runs “in isolated VMs in the cloud with full development environments” (checked 2026-08-28) |
| Started by | The workflow, in automation mode (a prompt input) | The workflow | The workflow, through @cursor/sdk 1.0.32 |
| Containment | --allowedTools allowlist, --max-turns, --max-budget-usd | permission-profile: ":workspace", which has no network access, and safety-strategy: drop-sudo | Cursor’s VM; the SDK exposes no tool allowlist for cloud agents |
| Who opens the PR | Your workflow | Your workflow | Cursor (autoCreatePR: true); your workflow labels it |
| Billing | Actions minutes plus API tokens, or a subscription through CLAUDE_CODE_OAUTH_TOKEN | Actions minutes plus API tokens (an OPENAI_API_KEY secret) | Cursor usage on the account that owns CURSOR_API_KEY |
All three use the tool’s default model unless you set one. Current defaults and prices live on the models hub; start there rather than pinning a model in the workflow.
Set up the repository once
Section titled “Set up the repository once”-
Create three labels.
agent:readystarts a run,agent-prmarks the pull requests the loop owns, andagent:needs-humanstops the loop. -
Create a small GitHub App for pushes. Give it Contents: read and write, Pull requests: read and write and Issues: read and write, install it on the repository, and store its client ID as the Actions variable
AGENT_APP_CLIENT_IDand its private key as the secretAGENT_APP_PRIVATE_KEY. The App exists because GitHub does not start workflows from events that the defaultGITHUB_TOKENcreates: a pull request opened with it would never run CI. Treat the App as an agent identity with its own owner; agent identity and secrets covers rotation and scope. -
Add the secret for your tool.
ANTHROPIC_API_KEY(orCLAUDE_CODE_OAUTH_TOKEN, generated withclaude setup-token),OPENAI_API_KEY, orCURSOR_API_KEY. For Claude Code you do not need the Claude GitHub App for this pipeline: the workflows pass the job’s read-onlyGITHUB_TOKENasgithub_token. Without that input the action authenticates through the App, whose token can write repository contents. -
Protect the default branch with a ruleset. Require a pull request, one approving review, a review from code owners, and the status checks
evidenceand your normal CI. Block force pushes. Then route the change classes a human must read throughCODEOWNERS, including every file that decides what “passing” means:# .github/CODEOWNERS: paths the pipeline may change but never merge alone/src/auth/ @acme/security/src/billing/ @acme/payments/migrations/ @acme/data/tests/contract/ @acme/tech-leads/.github/ @acme/tech-leads# The rules the agent is judged by, and the settings its runs load/scripts/agent-gates.sh @acme/tech-leads/scripts/evidence.sh @acme/tech-leads/package.json @acme/tech-leads/package-lock.json @acme/tech-leads/tsconfig*.json @acme/tech-leads/vite.config.* @acme/tech-leads/vitest.config.* @acme/tech-leads/eslint.config.* @acme/tech-leads/.claude/ @acme/tech-leads/.codex/ @acme/tech-leads/.cursor/ @acme/tech-leads/.mcp.json @acme/tech-leads/AGENTS.md @acme/tech-leads/CLAUDE.md @acme/tech-leads -
Ignore the agent’s scratch directory. Add
.agent/to.gitignore. The task file, gate logs and the agent’s self-report live there and must never reach a commit.
Write the task contract: issue form, gate script and result schema
Section titled “Write the task contract: issue form, gate script and result schema”Three files define what the agent is asked to do and what “done” means. They change rarely, and they are the files a tech lead reviews when the pipeline misbehaves.
The issue form makes the fields the prompt depends on mandatory. DORA’s advice on small batches applies directly here: “AI can easily generate massive blocks of code, which are hard to review and test. Enforcing the discipline of small batches counteracts this risk” (DORA, Google Cloud blog, 2025-12-10). One outcome per issue keeps each pull request reviewable from its evidence alone.
name: Agent-ready taskdescription: One outcome, provable by checks. Label agent:ready only after review.body: - type: textarea id: outcome attributes: label: Outcome description: One sentence a reviewer can verify. validations: { required: true } - type: textarea id: criteria attributes: label: Acceptance criteria description: "One per line: Given / when / then -> check: <test file or command>" validations: { required: true } - type: textarea id: scope attributes: label: Allowed paths description: Globs the agent may change, for example src/invoices/** validations: { required: true } - type: textarea id: out-of-scope attributes: label: Out of scope description: What must not change, even if it looks related.The gate script is the single definition of “done”. The agent runs it inside its run, and CI runs it again independently; only the CI result counts. CI always runs the base branch’s copy of the script against the pull request’s code, so an agent that edits the script, package.json (whose lint, typecheck and test scripts the gates call) or a test config changes nothing about how it is judged, and fails the oracle guard for trying. That guard is how the pipeline protects the oracle from the agent that is being judged by it.
#!/usr/bin/env bash# scripts/agent-gates.sh <base-ref># Writes .agent/gates.md and exits 0 only when every gate passes.set -uo pipefailBASE="${1:?usage: agent-gates.sh <base-ref>}"# The protected-path pattern lives next to the prompts, in .github/agent/, which is itself protected.PROTECTED=$(cat "$(dirname "$0")/../.github/agent/protected-paths.txt")mkdir -p .agentout=.agent/gates.mdprintf '| Gate | Result | Command |\n| --- | --- | --- |\n' > "$out"fail=0
gate() { local name="$1"; shift if "$@" > ".agent/$name.log" 2>&1; then echo "| $name | pass | \`$*\` |" >> "$out" else echo "| $name | **fail (exit $?)** | \`$*\` |" >> "$out"; fail=1 fi}
# Diff from the merge base, not the base tip: commits that landed on the base# branch after the PR forked are not the agent's. No merge base fails closed.if ! mb=$(git merge-base "$BASE" HEAD); then echo "| oracle-guard | **fail** | no merge base with $BASE |" >> "$out"; exit 1fi# Committed, uncommitted and new files, so it works before and after the commit.changed=$( { git diff --name-only "$mb"; git ls-files --others --exclude-standard; } | sort -u)protected=$(printf '%s\n' "$changed" | grep -E "$PROTECTED" || true)if [ -n "$protected" ]; then echo "| oracle-guard | **fail** | touched: $(echo $protected) |" >> "$out"; fail=1else echo "| oracle-guard | pass | no protected path touched |" >> "$out"fi
gate lint npm run lintgate typecheck npm run typecheckgate test npm testexit $failThe protected-path pattern is one extended regular expression in .github/agent/protected-paths.txt, a file inside a protected directory. The gate script and the fix loop both read it, so the list lives in one place:
^(\.github/|\.claude/|\.codex/|\.cursor/|\.mcp\.json$|AGENTS\.md$|CLAUDE\.md$|tests/contract/|migrations/|scripts/(agent-gates|evidence)\.sh$|package(-lock)?\.json$|tsconfig[^/]*\.json$|(vite|vitest|eslint)\.config\.[^/]+$)Add your own test and lint configuration files to it, and keep it in step with the CODEOWNERS block above.
The result schema gives every run the same machine-readable self-report. Codex enforces it through --output-schema; Claude Code and Cursor are asked to follow it in the prompt. The workflow prints the self-report in the pull request body under a heading that says it is a claim, not evidence.
{ "type": "object", "additionalProperties": false, "required": ["status", "stop_reason", "items"], "properties": { "status": { "enum": ["done", "stopped"] }, "stop_reason": { "type": "string" }, "items": { "type": "array", "items": { "type": "object", "additionalProperties": false, "required": ["claim", "check", "result"], "properties": { "claim": { "type": "string" }, "check": { "type": "string" }, "result": { "enum": ["pass", "fail", "not-run"] } } } } }}Save it as .github/agent/result.schema.json.
Copy-paste prompts for the unattended run
Section titled “Copy-paste prompts for the unattended run”The pipeline uses two prompts, stored in the repository so that a change to them goes through review like any other code. Both treat the issue and the review feedback as data, because anyone who can edit an issue can put text in front of the agent.
The third prompt is for the human who applies the label. Run it in any of the three tools, locally, before you label an issue; it catches the issues that would burn three fix attempts and still end in agent:needs-human.
Trigger the agent run from a labelled issue
Section titled “Trigger the agent run from a labelled issue”The workflow below is complete for Claude Code. The Codex and Cursor versions replace only the step marked AGENT STEP, as shown in the tabs after it.
name: agent-issueon: issues: types: [labeled]
concurrency: group: agent-issue-${{ github.event.issue.number }} cancel-in-progress: false
jobs: agent: if: github.event.label.name == 'agent:ready' runs-on: ubuntu-latest timeout-minutes: 45 permissions: contents: read # the only token in this job is read-only steps: - uses: actions/checkout@v7 with: fetch-depth: 0 persist-credentials: false # the agent gets no git credentials - uses: actions/setup-node@v7 with: { node-version: 22, cache: npm } - run: npm ci # install before the agent runs; Codex's :workspace has no network
- name: Build the task file (issue text goes through env, never through the shell) env: ISSUE_NUMBER: ${{ github.event.issue.number }} ISSUE_TITLE: ${{ github.event.issue.title }} ISSUE_BODY: ${{ github.event.issue.body }} BASE_REF: ${{ github.event.repository.default_branch }} run: | mkdir -p .agent { cat .github/agent/implement.md printf '\nBase ref for scripts/agent-gates.sh: origin/%s\n' "$BASE_REF" printf '\n<issue number="%s">\n# %s\n\n%s\n</issue>\n' \ "$ISSUE_NUMBER" "$ISSUE_TITLE" "$ISSUE_BODY" } > .agent/task.md
# AGENT STEP - name: Run Claude Code uses: anthropics/claude-code-action@v1 with: anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }} github_token: ${{ github.token }} # read-only here; without it the action mints a Claude App token that can push prompt: "Read .agent/task.md and follow it exactly." claude_args: >- --max-turns 60 --max-budget-usd 10 --allowedTools "Read,Edit,Write,Grep,Glob,Bash(npm run *),Bash(npm test *),Bash(scripts/agent-gates.sh *),Bash(git diff *),Bash(git status *)"
- name: Hand the change to the next job as a patch run: | git add -A git diff --cached --binary > agent.patch - uses: actions/upload-artifact@v7 with: name: agent-output path: | agent.patch .agent/result.json include-hidden-files: true if-no-files-found: warn
open-pr: needs: agent runs-on: ubuntu-latest permissions: contents: read steps: - uses: actions/create-github-app-token@v3 id: app with: client-id: ${{ vars.AGENT_APP_CLIENT_ID }} private-key: ${{ secrets.AGENT_APP_PRIVATE_KEY }} - uses: actions/checkout@v7 with: token: ${{ steps.app.outputs.token }} - uses: actions/download-artifact@v8 with: { name: agent-output, path: /tmp/agent }
- name: Commit, push and open a draft pull request env: GH_TOKEN: ${{ steps.app.outputs.token }} ISSUE: ${{ github.event.issue.number }} TITLE: ${{ github.event.issue.title }} RUN_URL: ${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }} run: | if [ ! -s /tmp/agent/agent.patch ]; then gh issue comment "$ISSUE" --body "The agent produced no change. Run: $RUN_URL" gh issue edit "$ISSUE" --add-label agent:needs-human --remove-label agent:ready exit 1 fi branch="agent/issue-$ISSUE" git switch -c "$branch" git apply --index /tmp/agent/agent.patch git -c user.name="agent-pipeline" -c user.email="agent-pipeline@users.noreply.github.com" \ commit -m "Implement #$ISSUE: $TITLE" -m "Agent-Run: $RUN_URL" git push -u origin "$branch" { echo "Closes #$ISSUE" echo echo "### Agent's own report (a claim, not evidence)" jq -r '"Status: \(.status). \(.stop_reason)", (.items[] | "- \(.claim) -> `\(.check)`: \(.result)")' \ /tmp/agent/.agent/result.json 2>/dev/null || echo "No self-report was written." echo echo "CI posts the evidence bundle as a comment. Run: $RUN_URL" } > body.md gh pr create --draft --head "$branch" --title "$TITLE (#$ISSUE)" \ --body-file body.md --label agent-pr gh issue edit "$ISSUE" --remove-label agent:ready
report-failure: needs: agent if: always() && (needs.agent.result == 'failure' || needs.agent.result == 'cancelled') runs-on: ubuntu-latest permissions: issues: write steps: - name: Hand the issue to a human instead of stalling silently env: GH_TOKEN: ${{ github.token }} GH_REPO: ${{ github.repository }} ISSUE: ${{ github.event.issue.number }} RUN_URL: ${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }} run: | gh issue comment "$ISSUE" --body "The agent run did not finish (budget, timeout or an action error). Run: $RUN_URL" gh issue edit "$ISSUE" --add-label agent:needs-human --remove-label agent:readyRemoving agent:ready at the end means a second label event does not start a duplicate run on the same issue. To rerun, delete the branch and apply the label again. The report-failure job covers the other exit: when the agent job fails or times out, open-pr never runs, so without it the issue would keep agent:ready with nothing to show for it. It uses the default GITHUB_TOKEN, because a comment and a label change on the issue need to start no workflow.
The workflow above is the Claude Code version. The action runs in automation mode because it has a prompt input, so it does not wait for an @claude mention. Before Claude starts, the action checks that the triggering user has write access and is not a bot, so a label applied by a user with only triage rights fails the run.
--max-budget-usd is listed in claude --help (2.1.283) and works in print mode, which is how the action runs Claude. --max-turns is accepted but no longer listed in --help (checked 2026-09-26), so rely on the budget and the job’s timeout-minutes as the hard limits.
The allowlist gives Claude no git commit, no git push and no GitHub tools, and the job gives the action nothing to push with: github_token: ${{ github.token }} is read-only under contents: read, and without id-token: write the action cannot exchange an OIDC token for a Claude GitHub App token. Keep it that way; if you drop github_token, the action falls back to that App token, which can write contents. One entry in the allowlist deserves a second look: Bash(npm run *) runs whatever package.json says, including a script the agent has just added. That is why package.json is on the protected-path list and in CODEOWNERS.
If you prefer a conversation over a pipeline, the action’s tag mode reacts to a label too: its label_trigger input names the label that starts it. In that mode Claude pushes a claude/ branch and replies with a link to a pre-filled pull request page rather than opening the pull request, per the action’s FAQ (checked 2026-09-26). That keeps a human click in the path, which is the opposite of what this page builds.
Claude Code routines cannot replace this workflow: their GitHub triggers accept pull request and release events only, not issue events (routines documentation, checked 2026-09-26). A routine fits the review side instead; see Claude Code routines.
Replace the AGENT STEP with the Codex action. The job’s permissions stay contents: read:
# AGENT STEP - name: Run Codex uses: openai/codex-action@v1 with: openai-api-key: ${{ secrets.OPENAI_API_KEY }} permission-profile: ":workspace" prompt-file: .agent/task.md output-schema-file: .github/agent/result.schema.json output-file: .agent/result.json:workspace is the permission profile the action’s README recommends for workflows that edit the checkout. It grants no network access, which is why npm ci runs before this step. It cannot be combined with the sandbox input: the action fails before Codex starts if you set both. The default safety-strategy: drop-sudo removes the runner user’s sudo before Codex runs, and the README advises running the action as the last privileged step in a job. That is one more reason the commit happens in a separate job.
output-schema-file passes --output-schema to codex exec, so the final message must match the schema and lands in .agent/result.json. Codex uses its default model unless you set the model input.
Two alternatives start Codex without a workflow. In Linear, you can assign an issue to Codex or mention @Codex (checked 2026-08-28; see Codex in Slack and Linear), and codex cloud exec --env ENV_ID "…" submits a Codex cloud task from a terminal (experimental in CLI 0.157.1). Both run in Codex cloud, so the gates and evidence steps still have to run in your CI once the pull request exists. The action’s own options are covered in the Codex GitHub Action.
Cursor runs the agent in a Cloud Agent and opens the pull request itself, so the Cursor version replaces both jobs with one. Save this script as .github/agent/cursor-run.mjs. It uses @cursor/sdk 1.0.32, whose type definitions were checked on 2026-09-26; the SDK needs Node 22.13 or later.
import { readFileSync, appendFileSync } from 'node:fs';import { Agent } from '@cursor/sdk';
const agent = await Agent.create({ apiKey: process.env.CURSOR_API_KEY, cloud: { repos: [{ url: `https://github.com/${process.env.GITHUB_REPOSITORY}`, startingRef: process.env.BASE_REF }], autoCreatePR: true, metadata: { issue: process.env.ISSUE_NUMBER }, },});const run = await agent.send(readFileSync('.agent/task.md', 'utf8'));const result = await run.wait();const prUrl = result.git?.branches?.[0]?.prUrl ?? '';appendFileSync(process.env.GITHUB_OUTPUT, `agent_id=${agent.agentId}\npr_url=${prUrl}\n`);if (result.status !== 'finished' || !prUrl) process.exit(1);Then use this job in place of agent and open-pr:
agent-cursor: if: github.event.label.name == 'agent:ready' runs-on: ubuntu-latest timeout-minutes: 60 permissions: { contents: read } steps: - uses: actions/checkout@v7 with: { persist-credentials: false } - uses: actions/setup-node@v7 with: { node-version: 22 } # Reuse the "Build the task file" step from the Claude Code workflow here. - run: npm install --no-save @cursor/sdk@1.0.32 - id: cursor env: CURSOR_API_KEY: ${{ secrets.CURSOR_API_KEY }} BASE_REF: ${{ github.event.repository.default_branch }} ISSUE_NUMBER: ${{ github.event.issue.number }} run: node .github/agent/cursor-run.mjs - uses: actions/create-github-app-token@v3 id: app with: client-id: ${{ vars.AGENT_APP_CLIENT_ID }} private-key: ${{ secrets.AGENT_APP_PRIVATE_KEY }} - name: Hand the pull request to the loop env: GH_TOKEN: ${{ steps.app.outputs.token }} PR_URL: ${{ steps.cursor.outputs.pr_url }} AGENT_ID: ${{ steps.cursor.outputs.agent_id }} run: | gh pr comment "$PR_URL" --body "cursor-agent-id: $AGENT_ID" gh pr edit "$PR_URL" --add-label agent-prThe label is applied with the App token for the same reason as the push in the other tabs: a label added with GITHUB_TOKEN would not start the loop workflow. The agent ID comment is what the fix loop uses to resume the same Cloud Agent.
Unlike the two actions, the SDK offers no tool allowlist for cloud agents (its tools option is local-only in 1.0.32), so the Cloud Agent’s VM and your gates are the containment. Cursor Automations, which can start a Cloud Agent from a Source control trigger (checked 2026-08-28), are the no-code alternative; the SDK version keeps the trigger, label and caps in your repository where they are reviewed. See Cursor Cloud Agents and Automations.
Run the gates and publish the evidence bundle
Section titled “Run the gates and publish the evidence bundle”The loop workflow owns every pull request labelled agent-pr. It re-runs the gates from scratch and posts the result as the evidence bundle. The agent’s claim that the gates passed is never used as evidence.
The work is split across two jobs because the pull request’s code is the agent’s output and must be treated as untrusted:
gates(reported as theevidencecheck) checks out the pull request’s head intopr/withpersist-credentials: false, checks out the base branch’sscripts/and.github/agent/intotrusted/, and runstrusted/scripts/agent-gates.shinsidepr/. Its only permission iscontents: read, and it references no secret, sonpm cilifecycle scripts and the tests the agent wrote have nothing to steal and no later privileged step to hijack.publishruns after it and never checks out or executes the pull request’s code. It downloads the gate log, builds the bundle with the base branch’sevidence.shfrom the GitHub API, comments withGITHUB_TOKEN, and only then mints the App token forgh pr ready.
The gate log itself was written in a job that ran the agent’s code, so the bundle treats it as a log: the pass or fail it reports comes from the gates job’s result, not from the file.
#!/usr/bin/env bash# scripts/evidence.sh <pr-number> <gates-result> <gates-md>: prints the evidence bundle as Markdown.# Reads the pull request through the GitHub API only; it never checks out or runs the PR's code.set -euo pipefailPR="$1"; RESULT="$2"; GATES="$3"pr=$(gh pr view "$PR" --json headRefOid,additions,deletions,changedFiles,files,closingIssuesReferences)issues=$(jq -r '[.closingIssuesReferences[].number | "#\(.)"] | join(", ")' <<<"$pr")attempts=$(gh api "repos/$GITHUB_REPOSITORY/pulls/$PR/commits" --paginate \ --jq '.[] | select(.commit.message | test("(^|\n)Agent-Fix:")) | .sha' | wc -l | tr -d ' ')sensitive=$(jq -r '.files[].path | select(test("^(src/auth/|src/billing/|migrations/)"))' <<<"$pr")cat <<EOF## Evidence bundle**Closes:** ${issues:-none} · **Head:** \`$(jq -r '.headRefOid[0:7]' <<<"$pr")\` · **Gates:** $RESULT · **Fix attempts:** $attempts of 3
### Gate log (written by the job that ran the PR's code; the check result above is what counts)$(cat "$GATES" 2>/dev/null || echo "No gate log was uploaded.")
### Size$(jq -r '"\(.changedFiles) files, +\(.additions) -\(.deletions)"' <<<"$pr")
### Sensitive paths (a named human reads these)${sensitive:-none}EOFThis is a minimal bundle. The full contract, with the spec delta, acceptance results per criterion, screenshots and provenance, is on the evidence bundle; extend this script toward it once the pipeline runs.
name: agent-loopon: pull_request: types: [opened, synchronize, reopened, labeled] pull_request_review: types: [submitted]
concurrency: group: agent-loop-${{ github.event.pull_request.number }} cancel-in-progress: false
jobs: gates: name: evidence # the required check: PR code runs here, with no secrets and no write token if: >- contains(github.event.pull_request.labels.*.name, 'agent-pr') && (github.event.action != 'labeled' || github.event.label.name == 'agent-pr') runs-on: ubuntu-latest timeout-minutes: 30 permissions: contents: read steps: - name: Check out the PR's code (untrusted, it is the agent's output) uses: actions/checkout@v7 with: ref: ${{ github.event.pull_request.head.sha }} path: pr fetch-depth: 0 persist-credentials: false - name: Check out the gate rules from the base branch (trusted) uses: actions/checkout@v7 with: ref: ${{ github.event.pull_request.base.ref }} path: trusted persist-credentials: false sparse-checkout: | scripts .github/agent - uses: actions/setup-node@v7 with: { node-version: 22, cache: npm, cache-dependency-path: pr/package-lock.json } - name: Run the base branch's gates against the PR's code working-directory: pr env: BASE: origin/${{ github.event.pull_request.base.ref }} run: | npm ci ../trusted/scripts/agent-gates.sh "$BASE" - uses: actions/upload-artifact@v7 if: always() with: name: gate-logs path: pr/.agent/ include-hidden-files: true if-no-files-found: warn
publish: needs: gates if: always() && (needs.gates.result == 'success' || needs.gates.result == 'failure') runs-on: ubuntu-latest permissions: contents: read pull-requests: write steps: - name: Check out trusted scripts only (the PR's code is never checked out here) uses: actions/checkout@v7 with: ref: ${{ github.event.pull_request.base.ref }} persist-credentials: false sparse-checkout: scripts - uses: actions/download-artifact@v8 continue-on-error: true with: { name: gate-logs, path: /tmp/gates } - name: Post the evidence bundle env: GH_TOKEN: ${{ github.token }} PR: ${{ github.event.pull_request.number }} RESULT: ${{ needs.gates.result }} run: | scripts/evidence.sh "$PR" "$RESULT" /tmp/gates/gates.md > evidence.md gh pr comment "$PR" --body-file evidence.md --edit-last --create-if-none - uses: actions/create-github-app-token@v3 id: app if: needs.gates.result == 'success' && github.event.pull_request.draft with: client-id: ${{ vars.AGENT_APP_CLIENT_ID }} private-key: ${{ secrets.AGENT_APP_PRIVATE_KEY }} - name: Mark ready for review (App token, so review workflows start) if: needs.gates.result == 'success' && github.event.pull_request.draft env: GH_TOKEN: ${{ steps.app.outputs.token }} run: gh pr ready "${{ github.event.pull_request.number }}"The gates job reports its check under the name evidence; require that name in the ruleset. Marking the pull request ready with the App token matters: review bots that start on ready_for_review, such as the Claude Code review workflow, skip drafts and never see an event that GITHUB_TOKEN creates.
Close the loop with a bounded review-fix job
Section titled “Close the loop with a bounded review-fix job”The fix jobs run when the gates fail or a human submits a “changes requested” review. They count earlier Agent-Fix: commit trailers on the branch, so the attempt counter survives reruns and needs no stored state. On the fourth attempt the workflow labels the pull request agent:needs-human and stops.
The same trust split applies here. fix-plan never checks out the pull request’s code: it counts attempts through the API, stops the loop when the pull request touches a protected path, and builds the fix task from the base branch’s fix.md. Only fix runs the agent on the pull request’s code, with a read-only token, and only push-fix, which runs nothing from the branch, holds the App token. Because fix-plan stops on any change to package.json or the lockfile, the npm ci in fix installs exactly what the base branch would. The same state machine, with its authority table, is described in running a bounded PR review-fix loop.
Add these jobs to agent-loop.yml:
fix-plan: needs: gates if: >- always() && contains(github.event.pull_request.labels.*.name, 'agent-pr') && !contains(github.event.pull_request.labels.*.name, 'agent:needs-human') && (needs.gates.result == 'failure' || github.event.review.state == 'changes_requested') runs-on: ubuntu-latest permissions: contents: read pull-requests: write outputs: attempt: ${{ steps.plan.outputs.attempt }} steps: - name: Check out the fix prompt and the protected-path list from the base branch uses: actions/checkout@v7 with: ref: ${{ github.event.pull_request.base.ref }} persist-credentials: false sparse-checkout: .github/agent - id: plan env: GH_TOKEN: ${{ github.token }} PR: ${{ github.event.pull_request.number }} run: | n=$(gh api "repos/$GITHUB_REPOSITORY/pulls/$PR/commits" --paginate \ --jq '.[] | select(.commit.message | test("(^|\n)Agent-Fix:")) | .sha' | wc -l | tr -d ' ') touched=$(gh pr view "$PR" --json files --jq '.files[].path' \ | grep -E -f .github/agent/protected-paths.txt || true) if [ -n "$touched" ]; then gh pr edit "$PR" --add-label agent:needs-human gh pr comment "$PR" --body "Protected paths changed: $(echo $touched). The loop stops here; a human decides." echo "attempt=stop" >> "$GITHUB_OUTPUT" elif [ "$n" -ge 3 ]; then gh pr edit "$PR" --add-label agent:needs-human gh pr comment "$PR" --body "Fix limit reached (3 attempts). Handing over to the issue owner." echo "attempt=stop" >> "$GITHUB_OUTPUT" else echo "attempt=$((n + 1))" >> "$GITHUB_OUTPUT" fi - uses: actions/download-artifact@v8 if: steps.plan.outputs.attempt != 'stop' && needs.gates.result == 'failure' with: { name: gate-logs, path: /tmp/gates } - name: Build the fix task (feedback is data) if: steps.plan.outputs.attempt != 'stop' env: GH_TOKEN: ${{ github.token }} PR: ${{ github.event.pull_request.number }} BASE_REF: ${{ github.event.pull_request.base.ref }} run: | mkdir -p task { cat .github/agent/fix.md printf '\nBase ref for scripts/agent-gates.sh: origin/%s\n' "$BASE_REF" echo '<feedback>' cat /tmp/gates/gates.md 2>/dev/null for f in /tmp/gates/*.log; do [ -f "$f" ] && { echo "## $f"; tail -n 80 "$f"; }; done gh pr view "$PR" --json reviews \ --jq '.reviews[] | select(.state == "CHANGES_REQUESTED") | "- review by \(.author.login): \(.body)"' gh api "repos/$GITHUB_REPOSITORY/pulls/$PR/comments" \ --jq '.[] | "- \(.path):\(.line // .original_line) \(.user.login): \(.body)"' echo '</feedback>' } > task/task.md - uses: actions/upload-artifact@v7 if: steps.plan.outputs.attempt != 'stop' with: { name: fix-task, path: task/task.md }
fix: needs: fix-plan if: always() && needs.fix-plan.result == 'success' && needs.fix-plan.outputs.attempt != 'stop' runs-on: ubuntu-latest timeout-minutes: 30 permissions: contents: read # the agent runs PR code here, so the job holds no write token steps: - uses: actions/checkout@v7 with: ref: ${{ github.event.pull_request.head.ref }} fetch-depth: 0 persist-credentials: false - uses: actions/download-artifact@v8 with: { name: fix-task, path: .agent } - uses: actions/setup-node@v7 with: { node-version: 22, cache: npm } - run: npm ci # safe to run only because fix-plan stops on any change to package.json or the lockfile
# AGENT STEP: the same step as in agent-issue.yml
- run: git add -A && git diff --cached --binary > agent.patch - uses: actions/upload-artifact@v7 with: { name: fix-output, path: agent.patch }
push-fix: needs: [fix-plan, fix] if: always() && needs.fix.result == 'success' runs-on: ubuntu-latest permissions: { contents: read } steps: - uses: actions/create-github-app-token@v3 id: app with: client-id: ${{ vars.AGENT_APP_CLIENT_ID }} private-key: ${{ secrets.AGENT_APP_PRIVATE_KEY }} - uses: actions/checkout@v7 with: ref: ${{ github.event.pull_request.head.ref }} token: ${{ steps.app.outputs.token }} - uses: actions/download-artifact@v8 with: { name: fix-output, path: /tmp/fix } - env: GH_TOKEN: ${{ steps.app.outputs.token }} PR: ${{ github.event.pull_request.number }} N: ${{ needs.fix-plan.outputs.attempt }} run: | if [ ! -s /tmp/fix/agent.patch ]; then gh pr edit "$PR" --add-label agent:needs-human gh pr comment "$PR" --body "Fix attempt $N produced no change. Handing over." exit 0 fi git apply --index /tmp/fix/agent.patch git -c user.name="agent-pipeline" -c user.email="agent-pipeline@users.noreply.github.com" \ commit -m "Address review and gate feedback" -m "Agent-Fix: $N" git pushThe push starts a new synchronize run, which re-runs the gates, so each attempt is judged exactly like the first one.
Use the same Run Claude Code step in the fix job. The loop needs a reviewer as well as gates: the code-review plugin workflow from the Claude Code GitHub Actions documentation posts inline comments when a pull request is opened, updated or marked ready, and skips drafts. Those comments reach the fix prompt through the pulls/…/comments call above. Managed Code Review is the no-workflow option (research preview, Team and Enterprise); its check run “always completes with a neutral conclusion”, so it informs the loop but never blocks it.
Cloud sessions also offer auto-fix (/autofix-pr from a terminal on the pull request’s branch): Claude watches the pull request and pushes a fix when one is clear, and asks you when a comment is ambiguous. Its documentation describes no attempt limit (checked 2026-09-26) and notes that it cannot react to merge conflicts, so keep the capped job above as the loop of record.
Use the same Run Codex step in the fix job. For the review side, @codex review on a pull request requests a Codex review on GitHub, and custom review rules live in AGENTS.md (checked 2026-08-28). Put your CODEOWNERS classes in those rules so every review names the change class it found.
Because :workspace has no network, the fix job fetches the review comments before the agent step and puts them in .agent/task.md. Codex never needs a GitHub token for the loop.
The Cursor fix step resumes the Cloud Agent that opened the pull request, so it keeps the context of its first run. Save it as .github/agent/cursor-fix.mjs:
import { readFileSync } from 'node:fs';import { Agent } from '@cursor/sdk';
const agent = await Agent.resume(process.env.AGENT_ID, { apiKey: process.env.CURSOR_API_KEY });const run = await agent.send(readFileSync('.agent/task.md', 'utf8'));const result = await run.wait();if (result.status !== 'finished') process.exit(1);In fix-plan, read AGENT_ID from the cursor-agent-id: comment with gh pr view "$PR" --json comments and pass it on as a job output. In fix, run the script instead of the agent step; it needs the fix-task artifact, not the pull request’s code. Drop the push-fix job: the Cloud Agent pushes to its own branch. Record the pull request’s head SHA with gh pr view "$PR" --json headRefOid before and after the run, and have a follow-up job with pull-requests: write label agent:needs-human when it did not change. That check keeps the loop honest even if a follow-up run ends without a commit. Count attempts with a label per run instead of the Agent-Fix: trailer, since the commit message is Cursor’s.
For review, Bugbot “reviews pull requests and identifies bugs, security issues, and code quality problems” (checked 2026-08-28); its comments reach the fix prompt the same way. See Bugbot.
How do you know the pipeline produces mergeable work?
Section titled “How do you know the pipeline produces mergeable work?”The gates prove each pull request. These five numbers prove the pipeline, and they come from labels and commit trailers you already have. The tech lead reviews them weekly for the first month.
| Measure | Definition | Healthy direction |
|---|---|---|
| First-pass gate rate | Agent pull requests whose first evidence run passed, divided by all agent pull requests | Rising; a low rate means the issues or the prompt need work, not the model |
| Fix attempts per merged PR | Count of Agent-Fix: trailers on merged agent pull requests | Mostly 0 or 1 |
| Hand-over rate | Pull requests that received agent:needs-human, divided by all agent pull requests | Falling, but never zero; zero means the stop conditions are too loose |
| Human edit rate | Agent pull requests where a human pushed a commit before merge | Falling |
| Post-merge revert rate | Agent pull requests reverted or followed by a fix within 14 days | At or below your human-authored baseline |
Sign-off stays with people. The issue owner approves the merge after reading the evidence bundle, not the diff; reading evidence instead of code shows what to read and in what order. CODEOWNERS forces a named reader for authentication, money, schema, migrations and the oracle itself. The tech lead owns the caps (three fix attempts, the dollar limit, the job timeouts) and which issue classes may receive agent:ready.
Start with ten issues labelled by one person. Compare each pull request’s cycle time from label to merge, and the hand-over rate, with how long the team estimated those ten issues, then widen the label to the team. How those reviews fit into a team’s day is covered in managing the review queue.
What breaks in an issue-to-PR pipeline?
Section titled “What breaks in an issue-to-PR pipeline?”The pull request opens, but no CI runs on it. The push or the pull request was made with GITHUB_TOKEN, and GitHub does not start workflows from events that token creates. Recovery: push, label and mark ready with the GitHub App token, as every workflow on this page does, then close and reopen the pull request to trigger CI once.
The run fails at once with a permission error. The person who applied agent:ready has triage rights but not write access, and both actions check the triggering user before the agent starts. Recovery: restrict who can label, or reapply the label as a maintainer. Do not reach for allowed_non_write_users or allow-users to work around it; that opens the run to anyone who can file an issue.
Codex fails installing a package or reaching a service. The :workspace profile has no network access. Recovery: install every dependency before the agent step, run integration tests that need services in the evidence job instead, and keep the agent’s gate run to what works offline.
The gates pass because the agent changed a test. A fix loop that is rewarded for green will find the cheapest way to green. Recovery: CI runs the base branch’s agent-gates.sh, so the agent cannot rewrite the rules it is judged by, and the oracle guard fails any change to a protected path, including the gate scripts, package.json and the test and lint configs. Add every directory whose tests define behaviour to protected-paths.txt and to CODEOWNERS.
The loop ping-pongs between two failures. Attempt 2 undoes attempt 1 and attempt 3 redoes it. Recovery: the cap stops it at three, and the fix prompt tells the agent to stop when a gate fails for the same reason twice. When it happens often, the issue was not agent-ready; send it back through shaping the backlog.
An issue or a review comment carries instructions. Anyone who can edit an issue can write “also update the deploy key” into it. Recovery: the workflow passes issue text through environment variables, wraps it in tags the prompt marks as data, and gives the agent no write token, no GitHub tools and, in Codex, no network. The CI jobs that run the pull request’s code hold no secrets. The wider model is on the agent threat model.
The issue keeps agent:ready and nothing happens. The agent job ran out of budget, hit its timeout or failed inside the action, so open-pr was skipped. Recovery: the report-failure job comments the run URL on the issue and swaps agent:ready for agent:needs-human. Read the run log before relabelling; a budget stop usually means the issue was too large.
A human pushed to the branch and the next fix overwrote their intent. The loop does not know who is steering. Recovery: when a human takes over, they apply agent:needs-human, which every fix job checks before it starts.
The base branch moved and the pull request now conflicts. No gate fails, so no fix starts. Recovery: add a scheduled job that rebases open agent-pr branches, or let the owner resolve conflicts by hand; the pipeline is for the first draft, not the last mile of a busy branch.
Spend creeps up. A vague issue uses the full turn and dollar budget three times over. Recovery: the first-pass gate rate shows it before the invoice does. Lower --max-budget-usd for the fix job, and look at the issues with the most attempts rather than at the model.