Rollouts and PR Routing & Approval in Cursor
Cursor’s PR Routing & Approval assigns reviewers from code ownership and commit history and can approve low-risk pull requests when criteria you define are met. It is safe only when CI, not the agent, computes the risk class, a GitHub ruleset owns the merge, every approval is audited, and a rollout monitor watches the change after merge.
This page is for the tech lead or CTO deciding what an agent may approve, and for the developer who wires it up. Your team’s agents open 30 pull requests a day. Half of them change docs, copy or an isolated component, and two senior engineers spend their mornings approving them while the billing change waits. Auto-approval by rule is the obvious fix. It is also the fastest way to merge a payments bug with a green check, unless “low risk” is computed by something the agent cannot edit.
The pressure is measurable. Faros AI’s AI Engineering Report 2026 (April 2026, vendor telemetry from 22,000 developers at Faros customers, a self-selected base) found the pull request merge rate per developer up 16.2%, median time in review up 441.5%, incidents per pull request up 242.7%, and 31.3% more pull requests merging without any review. Unreviewed merges already happen; this page turns them into rule-based merges with an audit trail.
What a rule-based approval setup gives you
Section titled “What a rule-based approval setup gives you”- A division of labour in which CI computes the risk class, Cursor approves, and GitHub enforces, so no single actor can approve its own work.
- Approval criteria you paste into PR Routing & Approval, and the starting risk rules behind them.
- A ruleset and
CODEOWNERSconfiguration that keeps high-risk paths with a named human even when an approval arrives. - A workflow step that arms auto-merge for
lowpull requests only, pinned to the commit that was checked. - An audit query, a weekly sample, and three metrics that tell you when to widen or narrow the rule.
- An acceptance test for the rollout monitor, whether that is Cursor’s Rollouts or your own canaries.
What do PR Routing & Approval and Rollouts do?
Section titled “What do PR Routing & Approval and Rollouts do?”Cursor documents the two features at different depths, and this page only relies on what could be checked. cursor.com could not be fetched from the writing environment on 2026-09-26, so the table records the last verified reading.
| Feature | What Cursor documents | Last checked | What this page assumes |
|---|---|---|---|
| PR Routing & Approval | “It assigns reviewers based on code ownership and commit history, and can approve low-risk PRs when your criteria are met.” (approval-agents) | 2026-08-28 | You write the criteria. Which signals they can read (labels, check status, paths) is confirmed in Cursor’s documentation before you rely on any of them. |
| Bugbot | “Bugbot reviews pull requests and identifies bugs, security issues, and code quality problems.” (bugbot) | 2026-08-28 | Its findings are input to the standard class, not an approval. See Bugbot: learned rules and Autofix. |
| Rollouts | Search results for Cursor’s changelog (2026-09-23 entry) describe a Rollouts feature for the period after merge; neither the changelog nor its documentation could be fetched on 2026-09-26. | not verified | This page names none of its settings. It gives you the acceptance test any rollout monitor must pass before you trust it. |
Who decides that a pull request is low risk?
Section titled “Who decides that a pull request is low risk?”The approval is only as trustworthy as the thing that labels the pull request. Split the job across three actors, and give each one exactly one power:
| Actor | Owns | Must never |
|---|---|---|
| CI (the evidence bundle check) | Computes the risk class from the diff and labels the pull request risk:low, risk:standard or risk:high | Read its policy or checker from the pull request branch |
| Cursor PR Routing & Approval | Requests reviewers by ownership; approves a pull request that meets the criteria | Decide the class from the agent’s own description, or sit in CODEOWNERS or a ruleset bypass list |
| GitHub ruleset | Blocks the merge until checks pass, approvals are current, and code owners have approved their paths | Be editable by the agent that opened the pull request |
The agent that wrote the change may declare a class in its evidence bundle, but the declaration can only raise the floor CI computes, never lower it. That one rule is what makes “approve low-risk” mean something.
Set up rule-based approval for low-risk pull requests
Section titled “Set up rule-based approval for low-risk pull requests”The order matters: the classifier comes first, the enforcement second, and the approver last. An approver switched on before the other two approves whatever the agent claims.
-
Land the evidence bundle gate. Follow the evidence bundle setup until the
evidencecheck is required onmainand every agent pull request carries onerisk:*label. The checker and policy load from the base commit, so a pull request cannot weaken the gate that judges it. -
Make the label trustworthy. The bundle workflow adds a label but does not remove an old one, so a pull request that was
lowand then touchedsrc/billing/can carry bothrisk:lowandrisk:high. Replace the labelling step with one that clears the others first:# .github/workflows/evidence-bundle.yml — replaces the "Label the risk class" step- name: Label the risk classif: always() && steps.check.outputs.risk != ''env:GH_TOKEN: ${{ github.token }}PR: ${{ github.event.pull_request.number }}RISK: ${{ steps.check.outputs.risk }}run: |for c in low standard high; do[ "$c" = "$RISK" ] || gh pr edit "$PR" --remove-label "risk:$c" || truedonegh pr edit "$PR" --add-label "risk:$RISK" -
Protect
mainwith a ruleset. In the repository’s Settings, open Rulesets under Code and automation, target the default branch, and turn on: require a pull request with one approval; dismiss stale approvals when new commits are pushed; require review from code owners; require theevidencecheck and your test job to pass; block force pushes. Leave Cursor’s integration off the bypass list. GitHub’s documentation describes the stale-approval rule precisely: if the diff changes after an approval, “the approving review is dismissed as stale, and the pull request cannot be merged until someone approves the work again”. -
Keep high-risk paths with named people. With code-owner review required, a pull request that touches an owned path needs that owner’s approval, whatever else approved it. Put every escalation class in
CODEOWNERS, and never list Cursor’s identity as an owner:# .github/CODEOWNERS.github/ @acme/platformscripts/check-evidence.mjs @acme/platform**/auth/** @acme/security**/billing/** @acme/payments**/payments/** @acme/payments**/migrations/** @acme/data**/*.sql @acme/data**/*.test.* @acme/platformThe last line routes every test change to a person, because a test edit changes what “green” means for every other change. The globs match the escalation classes on reading evidence instead of code.
-
Write the approval criteria in Cursor. Paste the criteria below into PR Routing & Approval. Then check in Cursor’s documentation which of these signals the criteria can read. A condition Cursor cannot evaluate is not dropped; it moves to the ruleset or to CI.
Approve a pull request only when ALL of these are true:1. It carries exactly one risk label, and that label is risk:low.2. The required checks, including "evidence", passed on the current head commit.3. No changed file matches a path in .github/CODEOWNERS.4. No test, snapshot, CI workflow, lint config, tsconfig or lockfile changed.5. No new dependency was added to package.json or any other manifest.6. The diff is at most 300 changed lines across at most 10 files.7. The evidence bundle lists every acceptance criterion with result: pass.If any condition fails, do not approve. Request reviewers by code ownershipand commit history, and leave one comment naming the condition that failed.Never approve on the basis of the pull request description alone. -
Run in shadow mode for two weeks. Do not arm auto-merge yet. Cursor approves, a person still merges, and the person records every pull request where they would have decided differently. Move on to step 7 when the disagreements are zero or all in the safe direction (Cursor declined something a person would have approved).
-
After two clean weeks, arm auto-merge for
lowonly. Enable Allow auto-merge in the repository settings, then add these steps to the end of the evidence job. Two details carry the safety.--match-head-commitpins the merge to the commit CI checked. A GitHub App token does the merging because, per GitHub’s documentation, events triggered by the repository’sGITHUB_TOKEN“will not create a new workflow run”, so a merge made with it would not start your deploy workflow.- name: Create a merge tokenid: app-tokenif: always() && steps.check.outputs.risk != ''uses: actions/create-github-app-token@v3with:client-id: ${{ vars.MERGE_APP_CLIENT_ID }}private-key: ${{ secrets.MERGE_APP_PRIVATE_KEY }}- name: Arm auto-merge for low risk onlyif: always() && steps.check.outputs.risk != ''env:GH_TOKEN: ${{ steps.app-token.outputs.token }}PR: ${{ github.event.pull_request.number }}HEAD: ${{ github.event.pull_request.head.sha }}run: |if [ "${{ steps.check.outcome }}" = "success" ] && [ "${{ steps.check.outputs.risk }}" = "low" ]; thengh pr merge "$PR" --auto --squash --match-head-commit "$HEAD"elsegh pr merge "$PR" --disable-auto || truefi--automerges “only after necessary requirements are met”, so the ruleset still decides. Theelsebranch disarms auto-merge when a later push raises the class. Grant the App Contents and Pull requests write access on this repository only.
What risk rules should you start with?
Section titled “What risk rules should you start with?”Start narrow. A low class that covers half the repository puts code into production that nobody read and nothing checked. Widen it one row at a time, each time the audit below shows a clean month for the rows you already have.
| Change | Starting class | Why |
|---|---|---|
| Docs, comments, READMEs | low | A mistake is visible, cheap and reversible |
UI copy and translation strings, with a screenshot in runtime | low | The runtime evidence is the check |
| Isolated component behind a feature flag that is off, with runtime evidence | low | The flag is the blast-radius limit |
| Behaviour change in code with tests, stricter or new tests only | standard | Needs a reviewer on the bundle and a review agent’s findings |
| Dependency added or upgraded, lockfile change | standard | Invented and typosquatted packages; see dependency verification |
CI workflows, infrastructure, tsconfig, lint config | standard or higher | Changes the oracle or the environment for every other change |
| Auth, money, schema, migrations, a loosened test | high | A named code owner reads the code; staged rollout after merge |
Two rules hold for every row. A change with no check covering it is never low, whatever its path, because the only evidence is the code. And an agent’s pull request that edits the risk policy, CODEOWNERS or the checker is high by construction, because those files are owned by @acme/platform.
How do you prove the setup works before you trust it?
Section titled “How do you prove the setup works before you trust it?”Test the gate the way you test any oracle: with cases that must pass and cases that must fail. Open four throwaway pull requests from an agent against a test branch protected by the same ruleset, and write down what happened.
| Test pull request | Expected result |
|---|---|
Fix a typo in docs/setup.md | Labelled risk:low, approved by Cursor, merged automatically, deploy workflow runs |
One-line change in src/billing/refund.ts with the bundle declaring class: low | CI raises it to risk:high, Cursor does not approve, @acme/payments is requested, auto-merge stays off |
| The docs pull request again, with a second commit pushed after Cursor approved | The approval is dismissed as stale, nothing merges until a new approval on the new head |
A docs change plus an edit to tests/orders.test.ts | Not low; a code owner for tests is requested |
If any row behaves differently, fix the configuration before step 7 goes live. The second row is the one that matters most: it proves the agent’s own claim cannot lower the class.
How do you audit what Cursor approved?
Section titled “How do you audit what Cursor approved?”Every approval is a GitHub review with an author, a state, a timestamp and the commit it was made on. That is your audit trail, and it lives outside Cursor. This query lists merged low pull requests from the last 30 days, who approved each one, and whether the approval was made on the commit that merged:
# Terminal, repository root. Needs gh and jq.gh pr list --state merged --label risk:low --limit 200 \ --search "merged:>=$(date -d '30 days ago' +%F)" \ --json number,title,url,mergedAt,headRefOid,reviews \| jq -r '.[] | . as $pr | [.reviews[] | select(.state == "APPROVED")] as $ok | [$pr.number, $pr.mergedAt[:10], ($ok | map(.author.login) | unique | join(",")), (if any($ok[]; .commit.oid == $pr.headRefOid) then "head" else "STALE" end), $pr.title] | @tsv'On macOS, replace date -d '30 days ago' +%F with date -v-30d +%F. Any row marked STALE means an approval did not cover the code that shipped; check the ruleset first.
Then run a weekly sample. A person picks five merged low pull requests at random, reads the diff in full, and records in the trust log whether reading the code found anything the evidence missed. Sampling is what keeps “nobody reads low” honest.
Track three numbers per month and act on them:
| Metric | Definition | Action |
|---|---|---|
| Auto-approval share | Merged pull requests approved only by Cursor ÷ all merged pull requests | Context for the next two; a rising share is only good while they stay flat |
| Auto-approved revert rate | Auto-approved merges reverted or hotfixed within 14 days ÷ auto-approved merges | Above the rate for human-approved standard merges: narrow the low rules |
| Sample miss rate | Sampled pull requests where reading the code found a defect ÷ sampled pull requests | Any miss: write the check that would have caught it, then decide whether the row stays low |
How do rollouts catch what approval missed?
Section titled “How do rollouts catch what approval missed?”Approval by rule accepts that some defects reach main. The rollout is what keeps them small. Match the post-merge safety net to the class, as the evidence bundle’s routing table does:
| Class | After merge |
|---|---|
low | Normal deploy, with a monitor watching error rate and latency for the changed service |
standard | Behind a feature flag or a canary, promoted when the monitor reports healthy |
high | Staged rollout and a production approval gate |
Whatever monitors the rollout (Cursor’s Rollouts, your CD system’s canary analysis, or an SLO alert), run this acceptance test on a staging environment before you let it decide anything:
- Ship a build that returns HTTP 500 on 5% of requests to one endpoint. The monitor must report a regression within the window you set.
- Ship a change to a service that receives no traffic in staging. A result of “no data” or “inconclusive” must stop promotion, not pass it.
- Take the revert path. The revert is a pull request like any other, so it must pass the same ruleset; confirm a revert of a
lowchange is itselflowand merges without waiting for a person. - Measure the time from the bad deploy to the revert reaching production. That number is your real blast-radius budget for
low.
Progressive delivery in depth, including flags, canaries and SLO-triggered rollback, is on progressive delivery for agent-written changes.
How do Claude Code and Codex handle the same approval?
Section titled “How do Claude Code and Codex handle the same approval?”The design on this page is the same in all three tools; what differs is who presses Approve.
PR Routing & Approval approves pull requests that meet your criteria and requests reviewers by ownership and history. The ruleset and CODEOWNERS above still decide whether the approval is enough to merge.
Claude Code’s Code Review posts findings on the pull request but does not approve: its check run “always completes with a neutral conclusion so it never blocks merging through branch protection rules”. The approval for low comes from a person reading the bundle, or from your own bot identity driven by the same criteria. See review automation with Claude Code.
Codex reviews pull requests with @codex review and automatic reviews, and OpenAI’s documentation described no pull request approval feature when last checked (2026-08-28). Use the same ruleset, labels and auto-merge step, with a person or your own bot approving low. See the Codex GitHub Action.
Copy-paste prompts for approval and rollout
Section titled “Copy-paste prompts for approval and rollout”PR_NUMBER is the merged pull request the monitor flagged.
What breaks when Cursor approves pull requests?
Section titled “What breaks when Cursor approves pull requests?”An approval covers a commit that is no longer the head. The agent pushes a fix after the approval, and the pull request merges on the old approval. Recovery: turn on stale-approval dismissal in the ruleset, keep --match-head-commit in the auto-merge step, and look for STALE rows in the audit query.
The criteria read the agent’s own claim. If Cursor’s criteria look at the bundle’s class field instead of the CI label, an agent that writes class: low on a billing change gets approved. Recovery: criteria name the risk:low label only, and the second test pull request above runs again after every criteria change.
Two risk labels on one pull request. A pull request labelled low on its first push and high on its second keeps both unless the workflow removes the old one. Recovery: the labelling step in step 2, plus criterion 1 (“exactly one risk label”).
Cursor’s identity sits in CODEOWNERS or the bypass list. Then its approval satisfies the code-owner rule for sensitive paths. Recovery: list teams of people as owners, audit the bypass list, and check the reviews author in the audit query for approvals on owned paths.
Auto-merge happens, deploy does not. The merge ran with GITHUB_TOKEN, so no push workflow started. Recovery: merge with a GitHub App token as in step 7.
A new directory handles money and matches no glob. src/checkout/ arrives, matches nothing, and billing changes come through as low. Recovery: review the policy whenever a top-level directory is added, and keep a fixture per sensitive directory so a rename fails the checker’s tests.
“Inconclusive” is read as “healthy”. A monitor with no traffic to judge reports nothing, and the rollout proceeds. Recovery: rerun the second rollout acceptance test, and require a positive “healthy” before promotion.
Everything becomes low, or nothing does. The first sends unread code to production; the second leaves the seniors approving typos. Recovery: read the three metrics monthly, and move one row of the risk table at a time.