From intent to production: a worked end-to-end pipeline
An intent-to-production pipeline carries one agent change from a written intent to full production traffic with no manual step except the approvals its risk class demands. CI computes the class, a GitHub ruleset and auto-merge merge by that class, the class sets the rollout budget, an SLO breach rolls back automatically, and each incident becomes an eval.
Your issue-to-PR pipeline works, and its pull requests arrive with green gates and an evidence bundle. Then they wait: a person merges on Thursday, someone deploys on Friday, a regression pages on-call on Saturday, and the lesson ends up in a chat thread no agent will read. The code was automated; the path to production and back was not.
This tutorial is for the developer who wires the pipeline, the tech lead who sets the classes and caps, and the CTO who signs off on what may reach production without a person reading it. It follows one change, cursor pagination on GET /orders, through every stage, and every file is meant to be copied.
What the intent-to-production pipeline gives you
Section titled “What the intent-to-production pipeline gives you”- A stage map from intent to production, with the file and the human role at each stage.
- The integration fixes that make the issue-to-PR pipeline and the evidence bundle work as one.
merge-by-class.yml, which arms auto-merge forlowpull requests only, from a job that never touches pull request code.- A
deploy.ymlplan job that picks the blast-radius budget from the merged pull request’s class and waits for a named approver onhigh. slo-rollback.ymlandopen-revert.sh: an SLO alert restores the previous release, opens an incident issue and a revert pull request, and freezes deploys.- An incident-to-eval loop, three copy-paste prompts, five acceptance drills and the numbers that show the chain works.
What does the pipeline look like from intent to production?
Section titled “What does the pipeline look like from intent to production?”The pipeline has eight stages. Read the last column twice: people stay in the loop where their judgment adds something, not at every diff.
| Stage | What happens | Repository file | Who decides | Human role |
|---|---|---|---|---|
| 1. Intent | A product need becomes an agent-ready issue with acceptance criteria, allowed paths and a guardrail metric | .github/ISSUE_TEMPLATE/agent-task.yml | The issue owner | Writes or approves the intent, applies agent:ready |
| 2. Build | The agent implements it in an isolated run and a draft pull request opens | agent-issue.yml, agent-loop.yml | The workflow | None while the loop is inside its caps |
| 3. Classify | CI recomputes the risk class from the diff and labels it | evidence-bundle.yml, check-evidence.mjs | CI, from the base branch’s policy | Owns the policy file |
| 4. Merge | low auto-merges after an approval; standard waits for a reviewer on the bundle; high waits for a code owner | merge-by-class.yml, the ruleset, CODEOWNERS | The ruleset | Approves on the bundle; code owners read high code |
| 5. Release | The class picks the exposure: full deploy, canary or staged rollout | deploy.yml, scripts/deploy.sh | The deploy plan | Named approver promotes high |
| 6. Watch | Guardrails and SLO alerts judge the change on real traffic | Rollout analysis, alert rules | The controller or the alert | None |
| 7. Roll back | Traffic returns to the previous release; a revert pull request and an incident issue open | slo-rollback.yml, scripts/open-revert.sh | The alert, through a dispatch | On-call reads the incident issue |
| 8. Learn | The incident becomes an eval and, if needed, a policy change | evals/incidents/, evidence-policy.yml | The pipeline, then a code owner | Signs off the eval and the policy change |
The worked example moves through those stages in one morning. The times are illustrative, not a measurement.
| Time | Event |
|---|---|
| 09:02 | Issue #418, “Paginate GET /orders with a cursor”, is labelled agent:ready; pull request #423 opens with its bundle at 09:31. |
| 10:07 | CI classified it standard (behaviour change in covered code, new tests), so a reviewer approved on the bundle and merged. deploy.yml starts a 5% canary. |
| 10:24 | The canary’s 5xx ratio breaches its guardrail. The controller aborts, the alert dispatches slo-breach, incident issue #431 and revert pull request #432 open, and deploys freeze. |
| 11:40 | An eval task from #431 runs through the same pipeline; the eval fails on the bad commit, passes on the revert, and a code owner approves it. |
Step 1: capture intent the pipeline can route
Section titled “Step 1: capture intent the pipeline can route”The issue form from the issue-to-PR pipeline already asks for the outcome, acceptance criteria and allowed paths. Add two fields that the later stages need. The guardrail comes from the service’s SLO document and is written before the change exists, because an agent that picks its threshold after seeing canary numbers is grading its own work.
# Append to the body of .github/ISSUE_TEMPLATE/agent-task.yml - type: dropdown id: expected-risk attributes: label: Expected risk class description: CI computes the real class from the diff. This can only raise it. options: [low, standard, high] validations: { required: true } - type: textarea id: guardrail attributes: label: Guardrail description: "Service, SLO document and the metric that triggers rollback, e.g. orders-api, docs/slo/orders-api.md, 5xx ratio < 0.01 on the canary" validations: { required: true }The first prompt turns a product note into an issue that fills those fields. It works unchanged in Claude Code, Codex and Cursor.
Step 2: join the issue-to-PR pipeline and the evidence bundle
Section titled “Step 2: join the issue-to-PR pipeline and the evidence bundle”The two pages were written to stand alone, so running them together needs three changes. Without them, every agent pull request fails the bundle check, the ruleset cannot tell two checks apart, or a label step fails.
-
Put the bundle in the agent’s pull request. The issue-to-PR workflow builds the pull request body from the agent’s self-report, and the evidence checker looks for a YAML block that starts with
evidence_bundle:. Add a step to.github/agent/implement.mdso the agent writes one:8. Write .agent/bundle.yml in the format of the evidence_bundle block in.github/pull_request_template.md, starting with the line"evidence_bundle: 1". Copy the guardrail from the issue intorollout.guardrails. Declare the risk class from the issue or higher.Add
.agent/bundle.ymlto the artifact paths of theagentjob, then append it in theopen-prjob, after the self-report and beforegh pr create:Terminal window if [ -s /tmp/agent/.agent/bundle.yml ]; then{ echo; echo '## Evidence bundle'; echo; echo '~~~yaml'cat /tmp/agent/.agent/bundle.ymlecho '~~~'; } >> body.mdfiDo the same in
push-fixwithgh pr edit "$PR" --body-file, so a fix that changes the tests also updates the bundle. Theeditedtrigger onevidence-bundle.ymlthen re-runs the check. -
Give each check its own name. The
gatesjob inagent-loop.ymlreports asevidence, and so does theevidencejob inevidence-bundle.yml. Rename the first one toname: agent-gates, keep its job IDgatesso theneeds:references still work, and require both names in the ruleset. GitHub’s documentation does not say which result a ruleset reads when two jobs report the same check name, so do not leave it to chance. -
Create the new labels.
revert,incidentandincident:agent, next to therisk:*labels the evidence bundle already uses.ghdoes not create missing labels.
Step 3: merge by risk class with a ruleset and auto-merge
Section titled “Step 3: merge by risk class with a ruleset and auto-merge”The ruleset decides whether a merge is allowed, CODEOWNERS decides who must read, and a small workflow asks GitHub to merge as soon as the ruleset allows it, for low only. Configure the ruleset on the default branch (Settings, then Rulesets):
| Rule | Setting | The hole it closes |
|---|---|---|
| Require a pull request before merging | 1 approval, dismiss stale approvals, require review from code owners | GitHub dismisses an approval “as stale” when the diff changes after it, so a fix pushed after approval cannot ride on it |
| Require status checks to pass | evidence, agent-gates and your test job; pin each to the GitHub Actions app as its expected source | Any integration with write access can set a status; pinning the source means only your workflows can turn a check green |
| Allowed merge methods | Squash only | A squash merge has one parent, so the git revert in step 5 needs no -m and cannot fail on a merge commit |
| Block force pushes | On | History cannot be rewritten under an approved head |
| Bypass list | Empty, and never the release or agent app | An app on the bypass list merges without any of the above |
Then add the workflow. It runs on workflow_run, after evidence-bundle finishes, so it executes from the default branch and never checks out the pull request. That is what lets it hold the release app’s token safely.
name: merge-by-classon: workflow_run: workflows: [evidence-bundle] types: [completed]
permissions: {}
jobs: route: if: github.event.workflow_run.event == 'pull_request' runs-on: ubuntu-latest steps: - uses: actions/create-github-app-token@v3 id: app with: client-id: ${{ vars.RELEASE_APP_CLIENT_ID }} private-key: ${{ secrets.RELEASE_APP_PRIVATE_KEY }} - name: Arm auto-merge for low risk only, pinned to the checked commit env: GH_TOKEN: ${{ steps.app.outputs.token }} GH_REPO: ${{ github.repository }} HEAD_SHA: ${{ github.event.workflow_run.head_sha }} CONCLUSION: ${{ github.event.workflow_run.conclusion }} run: | pr=$(gh api "repos/$GH_REPO/commits/$HEAD_SHA/pulls" \ --jq '[.[] | select(.state == "open")] | first // {} | .number // empty') [ -n "$pr" ] || { echo "No open pull request for $HEAD_SHA"; exit 0; } classes=$(gh pr view "$pr" --json labels \ --jq '[.labels[].name | select(startswith("risk:"))] | join(",")') if [ "$CONCLUSION" = "success" ] && [ "$classes" = "risk:low" ]; then gh pr merge "$pr" --auto --squash --match-head-commit "$HEAD_SHA" else gh pr merge "$pr" --disable-auto || true fiThree details carry the safety. --match-head-commit merges only the commit that was classified. The else branch disarms auto-merge when a later push raises the class. And the merge uses a GitHub App token because, per GitHub’s documentation, events created with the repository’s GITHUB_TOKEN do not start new workflow runs, so a merge made with it would never start deploy.yml. Create the release app with Contents, Pull requests and Issues write access on this repository only, turn on Allow auto-merge in the repository settings, and keep the app out of CODEOWNERS.
The ruleset still requires one approval, even for low. Who gives it is the one place the three tools differ.
Claude Code does not approve pull requests. Its managed Code Review posts findings, and its check run “always completes with a neutral conclusion so it never blocks merging through branch protection rules” (Claude Code documentation, checked 2026-09-26). Feed those findings to the standard reviewer, and have a person approve low pull requests on the bundle. If you build your own approver identity, confirm on a test branch that the ruleset counts its approval. Setup is on review automation with Claude Code.
Codex reviews pull requests when you comment @codex review, and custom review rules live in AGENTS.md (checked 2026-08-28). OpenAI’s documentation described no pull request approval feature when last checked (2026-08-28), so a person approves low on the bundle, as in the Claude Code tab. Put your risk classes in the review rules so every Codex review names the class it saw.
PR Routing & Approval “assigns reviewers based on code ownership and commit history, and can approve low-risk PRs when your criteria are met” (cursor.com, checked 2026-08-28). Make the criteria name the risk:low label, never the bundle’s self-declared class, and keep Cursor’s identity out of CODEOWNERS and the bypass list. The criteria text, a two-week shadow mode and an audit query are on Rollouts and PR Routing & Approval in Cursor. Cursor’s newer post-merge release features could not be verified on 2026-09-26, so this pipeline does not depend on them.
Step 4: release each class with its blast-radius budget
Section titled “Step 4: release each class with its blast-radius budget”The deploy workflow reads the class of the pull request that produced the commit, not a class the agent declared. The budgets (first exposure, promotion steps, bake time and rollback trigger per class) are defined on progressive delivery for agent-written changes; this workflow only chooses between them.
# .github/workflows/deploy.yml: plan and approval jobs. Build and push your image before "deploy".name: deployon: push: branches: [main]
concurrency: group: deploy-production cancel-in-progress: false
permissions: {}
jobs: plan: runs-on: ubuntu-latest permissions: contents: read pull-requests: read outputs: risk: ${{ steps.plan.outputs.risk }} steps: - id: plan env: GH_TOKEN: ${{ github.token }} GH_REPO: ${{ github.repository }} run: | # An open revert means production runs an older release on purpose. if [ "$(gh pr list --label revert --state open --json number --jq length)" != "0" ]; then echo "An open revert pull request freezes deploys. Merge or close it first."; exit 1 fi # No merged pull request, or no class label, means high. read -r pr risk < <(gh api "repos/$GH_REPO/commits/$GITHUB_SHA/pulls" \ --jq '[.[] | select(.merged_at != null)] | first // {} | "\(.number // "none") \(([.labels[]?.name | select(startswith("risk:"))] | first) // "risk:high")"') echo "Pull request $pr, class ${risk#risk:}" echo "risk=${risk#risk:}" >> "$GITHUB_OUTPUT"
approve-high: needs: plan if: needs.plan.outputs.risk == 'high' runs-on: ubuntu-latest environment: production-approval # required reviewers, no secrets steps: - run: echo "Staged rollout of $GITHUB_SHA approved"
deploy: needs: [plan, approve-high] if: >- always() && needs.plan.result == 'success' && (needs.approve-high.result == 'success' || needs.approve-high.result == 'skipped') runs-on: ubuntu-latest environment: production permissions: contents: read steps: - uses: actions/checkout@v7 with: { persist-credentials: false } - name: Deploy with this class's budget env: RISK: ${{ needs.plan.outputs.risk }} run: scripts/deploy.sh "$RISK" "$GITHUB_SHA"scripts/deploy.sh is yours: it maps low to a normal rollout, standard to the analysed canary and high to the staged rollout. Keep the mapping in that one script so a budget change is one reviewed diff. The prompt below drafts it with the step 5 rollback script, in any of the three tools.
The production-approval environment holds the production approval gate for high: up to six required reviewers, with self-review prevented, so the person who merged cannot also release. On the Free, Pro and Team plans, required reviewers for environments are available only in public repositories (GitHub documentation, checked 2026-09-26), so check your plan before you design around them.
Step 5: roll back on an SLO breach without waiting for a person
Section titled “Step 5: roll back on an SLO breach without waiting for a person”A canary controller catches a bad release below full exposure and has already moved traffic back. An SLO alert catches a low change that went straight to 100%, and nothing has moved yet. Both send the same event to GitHub, and one workflow handles the rest.
Your alerting system (Alertmanager and Grafana alerting can both send a webhook) calls a small relay, which sends a repository dispatch carrying the service, the commit SHA the service reports as its version, the alert name, and the exposure, canary or full:
# In the relay. GITHUB_DISPATCH_TOKEN comes from the relay's secret store.curl -sS -X POST "https://api.github.com/repos/acme/shop/dispatches" \ -H "Accept: application/vnd.github+json" \ -H "Authorization: Bearer $GITHUB_DISPATCH_TOKEN" \ -d '{"event_type":"slo-breach","client_payload":{"service":"orders-api","sha":"'"$VERSION_SHA"'","alert":"OrdersErrorBudgetBurn","stage":"canary"}}'Issue that token from a GitHub App installed on this repository alone, with the minimum permission GitHub’s REST reference lists for creating a repository dispatch event. A repository_dispatch workflow always runs from the default branch, so the payload can choose values but never code.
name: slo-rollbackon: repository_dispatch: types: [slo-breach]
concurrency: group: slo-rollback-${{ github.event.client_payload.service }} cancel-in-progress: false
permissions: {}
jobs: restore: # A canary abort has already restored traffic; a full deploy has not. if: github.event.client_payload.stage == 'full' runs-on: ubuntu-latest environment: production-rollback # branch rule only, no reviewers: a rollback must not wait permissions: contents: read steps: - uses: actions/checkout@v7 with: { persist-credentials: false } - name: Return the service to its previous release, with no build env: SERVICE: ${{ github.event.client_payload.service }} run: | [[ "$SERVICE" =~ ^[a-z0-9-]{1,40}$ ]] || { echo "Invalid service"; exit 1; } scripts/rollback.sh "$SERVICE"
record: needs: restore if: always() # open the incident and the revert even if the restore failed runs-on: ubuntu-latest permissions: contents: read steps: - uses: actions/create-github-app-token@v3 id: app with: client-id: ${{ vars.RELEASE_APP_CLIENT_ID }} private-key: ${{ secrets.RELEASE_APP_PRIVATE_KEY }} - uses: actions/checkout@v7 with: fetch-depth: 0 token: ${{ steps.app.outputs.token }} - name: Open the incident issue and the revert pull request env: GH_TOKEN: ${{ steps.app.outputs.token }} SHA: ${{ github.event.client_payload.sha }} SERVICE: ${{ github.event.client_payload.service }} ALERT: ${{ github.event.client_payload.alert }} STAGE: ${{ github.event.client_payload.stage }} RESTORE: ${{ needs.restore.result }} HUMAN_OWNER: ${{ vars.ONCALL_OWNER }} RUN_URL: ${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }} run: scripts/open-revert.shscripts/rollback.sh wraps your platform’s own undo, for example kubectl rollout undo deployment/"$SERVICE" for a plain Kubernetes Deployment. It must not build anything: a rollback that waits for CI is a second incident. The production-rollback environment holds the same deploy credentials as production, with a deployment-branch rule and no required reviewers.
The script opens the incident issue first, so the revert’s bundle can link to it, then commits the revert and opens its pull request with a bundle generated from the diff:
#!/usr/bin/env bash# scripts/open-revert.sh: run by slo-rollback.yml on the default branch with the release# app's token in GH_TOKEN. Opens the incident issue and a revert pull request for the# merge that breached an SLO. It never builds or runs the reverted code.set -euo pipefail[[ "$SHA" =~ ^[0-9a-f]{7,40}$ ]] || { echo "Invalid sha: $SHA"; exit 1; }[[ "$SERVICE" =~ ^[a-z0-9-]{1,40}$ ]] || { echo "Invalid service: $SERVICE"; exit 1; }SHA=$(git rev-parse --verify "$SHA^{commit}")
# Provenance: the pull request that merged the commit, and whether an agent wrote it.pr_json=$(gh api "repos/$GITHUB_REPOSITORY/commits/$SHA/pulls" \ --jq '[.[] | select(.merged_at != null)] | first // {}')pr=$(jq -r '.number // empty' <<<"$pr_json")labels=$(jq -r '[.labels[]?.name] | join(",")' <<<"$pr_json")# Same files, same computed floor; the reverted class is at least that floor, so declare it.class=$(jq -r '[.labels[]?.name | select(startswith("risk:"))] | first // "risk:high"' <<<"$pr_json")class=${class#risk:}kind=incident[[ ",$labels," == *",agent-pr,"* ]] && kind=incident:agent
cat > /tmp/incident.md <<MD**Alert:** $ALERT on \`$SERVICE\` · **Stage:** $STAGE · **Restore job:** $RESTORE**Pull request:** ${pr:+#$pr} (\`${SHA:0:12}\`) · **Labels:** ${labels:-none} · **Run:** $RUN_URL
Close this issue only when each box is ticked with a link:- [ ] Eval or test in \`evals/incidents/\` that fails on \`${SHA:0:12}\` and passes on the revert- [ ] Policy change, if the risk class or the guardrail let this through- [ ] Loop level decision, if an agent loop produced the changeMDissue_url=$(gh issue create --title "SLO breach on $SERVICE after ${pr:+#$pr }${SHA:0:12}" \ --label "$kind" --body-file /tmp/incident.md)
# The revert, and one check a reviewer can trust: every file the commit touched is back to its parent.branch="revert/${SHA:0:12}"git switch -c "$branch"git -c user.name="release-bot" -c user.email="release-bot@users.noreply.github.com" \ revert --no-edit "$SHA"mapfile -t files < <(git diff --name-only "$SHA^" "$SHA")if git diff --quiet "$SHA^" HEAD -- "${files[@]}"; then restored=pass; code=0; else restored=fail; code=1; fi
{ echo "Reverts ${pr:+#$pr }(\`${SHA:0:12}\`) after an SLO breach. Incident: $issue_url" echo echo '~~~yaml' echo 'evidence_bundle: 1' echo 'spec:' echo " link: $issue_url" echo " delta: [\"$SERVICE behaves as it did before ${SHA:0:12}\"]" echo ' unrequested: []' echo 'acceptance:' echo " - criterion: every file ${SHA:0:12} changed matches its parent commit" echo " check: git diff --quiet ${SHA:0:12}^ HEAD -- <the ${#files[@]} files>" echo " result: $restored" echo 'checks:' echo " - command: git diff --quiet ${SHA:0:12}^ HEAD -- <the ${#files[@]} files>" echo " exit_code: $code" echo 'risk:' echo " class: $class # the reverted pull request's class; CI fails anything lower" echo ' touches: [auth, money, schema, migrations, infra] # over-declared on purpose' echo ' oracle_changes:' for f in "${files[@]}"; do printf ' - { path: "%s", direction: neutral, reason: "revert of %s" }\n' "$f" "${SHA:0:12}" done echo " rollback: revert this revert once the incident eval passes" echo 'provenance:' echo ' agent: slo-rollback workflow (no agent)' echo ' model: none' echo " session: $RUN_URL" echo " task: $issue_url" echo " human_owner: \"$HUMAN_OWNER\"" echo '~~~'} > /tmp/revert.md
git push -u origin "$branch"gh pr create --head "$branch" --title "Revert ${pr:+#$pr }(${SHA:0:12}): SLO breach on $SERVICE" \ --body-file /tmp/revert.md --label revertTwo choices in the generated bundle need a word. It declares every sensitive class and lists every reverted file as a neutral oracle change, because the checker fails a bundle that omits a touched class or oracle file, and the floor still comes from the diff: a reverted test file lifts it to at least standard. And it declares the class of the pull request it reverts, which is at least the floor the checker computes for the same files, so a docs revert stays low and can auto-merge, while a billing revert is high and waits for its code owner. If the reverted pull request carries no risk:* label, the script declares high. That wait is safe, because traffic is already back on the previous release and deploy.yml refuses to ship main while the revert is open.
Step 6: turn the incident into an eval through the same pipeline
Section titled “Step 6: turn the incident into an eval through the same pipeline”The incident issue closes only when its lesson exists as a check. The eval is an agent task that goes through stages 1 to 4 like any other; the difference is the proof: it must fail on the bad commit and pass on the revert. Postmortems, containment and loop demotion are on when an agent causes an incident.
Add /evals/ @acme/tech-leads to CODEOWNERS, so a person signs every eval. Then turn the incident into a task:
A person reviews the issue, applies agent:ready, and the pipeline builds the eval. Before approving that pull request, the code owner runs the proof, because the bundle can only claim it:
# Terminal, repository root. BAD_SHA is the reverted commit from the incident issue.git worktree add /tmp/inc-431 "$BAD_SHA"mkdir -p /tmp/inc-431/evals/incidentscp evals/incidents/INC-431.test.ts /tmp/inc-431/evals/incidents/(cd /tmp/inc-431 && npm ci && npm test -- evals/incidents/INC-431.test.ts) # must failnpm test -- evals/incidents/INC-431.test.ts # must passgit worktree remove --force /tmp/inc-431Only then does the change return: revert the revert, fix the defect so the eval passes, and send it through the pipeline again. The eval stays in the suite and runs on every harness change, as continuous evals describes.
The tools differ only at the trigger. For local runs, save the prompt as .github/agent/incident-to-task.md with the incident number filled in.
Label the eval issue agent:ready, and agent-issue.yml runs it through anthropics/claude-code-action@v1 like any other task. To draft the task locally, run the prompt headless; claude -p starts in Manual permission mode, so tools outside the allowlist are denied:
# Terminal, repository root (Claude Code 2.1.283)claude -p "$(cat .github/agent/incident-to-task.md)" \ --allowedTools "Read,Grep,Glob,Bash(gh issue view *),Bash(gh pr view *)" \ --max-budget-usd 2Label the eval issue agent:ready, and the Codex version of agent-issue.yml runs it through openai/codex-action@v1 with the :workspace permission profile. To draft the task locally, use the read-only profile, which cannot write to the checkout:
# Terminal, repository root (Codex CLI 0.157.1)codex exec -c default_permissions=":read-only" \ -o incident-task.md "$(cat .github/agent/incident-to-task.md)"Gather the incident and pull request text with gh before the run and paste it into the prompt, because a read-only run cannot fetch it.
Label the eval issue agent:ready, and the Cursor version of agent-issue.yml starts a Cloud Agent through @cursor/sdk. Cursor Automations can also start a Cloud Agent from a Sentry or PagerDuty trigger (checked 2026-08-28), so the eval task can be drafted the moment an alert fires; keep its output an issue for a person to label, not a pull request. See Cursor Cloud Agents and Automations.
How do you prove the whole pipeline works before it deploys anything?
Section titled “How do you prove the whole pipeline works before it deploys anything?”Run these five drills against a staging service with the same ruleset and workflows. Each proves one link in the chain.
| Drill | What you do | Pass when |
|---|---|---|
| Low path | An agent fixes a typo in docs/setup.md | Labelled risk:low, approved, merged automatically, deploy.yml runs a normal rollout |
| Class cannot be lowered | An agent changes one line in src/billing/refund.ts and declares class: low | CI labels risk:high, auto-merge stays off, review is requested from the code owner, approve-high waits |
| Canary abort | Ship a standard build that returns HTTP 500 on 5% of GET /orders requests | The analysis aborts, the dispatch arrives with stage: canary, an incident issue and a revert pull request open, restore is skipped |
| Full rollback | Ship a low change and fire the SLO alert by hand | restore runs rollback.sh with no build, the revert opens, and the next push to main fails the plan job with the freeze message |
| Eval proof | Run the incident-to-eval task for the canary drill | The eval fails on the bad commit and passes on main, and a code owner approves it |
Time the canary and full-rollback drills from bad deploy to zero exposure: that, not your budget table, is how long a defect reaches customers.
Which numbers show intent-to-production is working?
Section titled “Which numbers show intent-to-production is working?”These measures come from labels, deployments and incident issues you already have. The tech lead reviews them monthly; the CTO reads the trend beside metrics frameworks.
| Measure | Definition | What to do with it |
|---|---|---|
| Intent-to-production lead time | From agent:ready to 100% exposure, median, per class | Compare with human-authored changes of the same class; the gap shows the queue |
| Change failure rate by class | Merged changes that caused a rollback, a revert or an incident, divided by merged changes, per class | low above standard means the low rules are too wide |
| Caught before 100% | Rollbacks that happened at canary or staged exposure, divided by all rollbacks | Low values mean customers find defects first; add guardrails |
| Time to zero exposure | From the first guardrail breach or alert to 0% traffic on the change, median | Hours mean the rollback waits for a build or a person |
| Incident-to-eval closure | Incident issues closed with a proven eval or an enforced policy change, divided by incident issues | Below 100% means lessons still live in chat threads |
DORA’s 2025 report (Google Cloud, 23 September 2025) found “a positive relationship between AI adoption on both software delivery throughput and product performance”, and also that “AI adoption does continue to have a negative relationship with software delivery stability.” Faros AI’s AI Engineering Report 2026 (April 2026, vendor telemetry from 22,000 developers) measured task throughput per developer up 33.7% alongside incidents per pull request up 242.7%. These measures show which side of those findings your pipeline is on.
Sign-off stays with named people. The issue owner signs the intent; the tech lead owns the policy file, caps and risk table; a code owner reads every high change and every eval. The CTO owns which classes reach production without a person reading code, and changes that one class at a time, on these numbers.
What breaks between merge and production?
Section titled “What breaks between merge and production?”Every agent pull request fails evidence with “No evidence bundle found”. The body was built without the bundle. Recovery: apply step 2 and edit the open pull requests’ bodies; the edited trigger re-runs the check.
Auto-merge happens but no deploy starts. The merge ran with GITHUB_TOKEN, so the push event started no workflow. Recovery: arm auto-merge with the release app’s token, as merge-by-class.yml does.
A low pull request never merges. Nobody approved it, or a rule adds one. GitHub’s rulesets require one extra approval by default for Copilot cloud agent pull requests that are not attributed to a person (public preview, GitHub documentation checked 2026-09-26), so those need two. Recovery: decide who approves low (step 3 tabs), and recheck the ruleset when a new agent starts opening pull requests.
A merge queue stalls every pull request. The required checks never ran on merge_group, the event GitHub’s merge queue documentation says you must use to trigger Actions workflows for queued pull requests. The evidence checker reads the pull request body, which that event does not carry. Recovery: add the merge_group trigger to each required workflow, with a bundle job that takes the pull request number from the queue branch name (gh-readonly-queue/<base>/pr-<number>-<sha>), reads that pull request’s head SHA, and passes only if gh api repos/$GH_REPO/commits/$SHA/check-runs reports a successful evidence run for it. Or turn the queue off.
The rollback waited for an approver. The restore job used the production environment, whose reviewers apply to rollbacks too. Recovery: give rollback its own production-rollback environment with a branch rule and no reviewers, and rehearse it with the full-rollback drill.
The revert pull request fails its own bundle check. The reverted change touched UI paths that need runtime evidence the script cannot produce, or later commits changed the same files. Recovery: production is already restored, so a person finishes the revert by hand and records why in the incident issue.
The dispatch arrives with the wrong SHA. The service reports a build number rather than a commit, or the relay reads the stable pods instead of the canary. Recovery: export the commit SHA as a version label on every pod and read it from the canary’s series; the script rejects anything that is not a commit in the repository.
Incidents close without an eval. On-call ticks the boxes to clear the queue. Recovery: someone other than the incident owner checks the links, and the closure measure goes on the monthly review.