Skip to content

From intent to production: a worked end-to-end pipeline

An intent-to-production pipeline carries one agent change from a written intent to full production traffic with no manual step except the approvals its risk class demands. CI computes the class, a GitHub ruleset and auto-merge merge by that class, the class sets the rollout budget, an SLO breach rolls back automatically, and each incident becomes an eval.

Your issue-to-PR pipeline works, and its pull requests arrive with green gates and an evidence bundle. Then they wait: a person merges on Thursday, someone deploys on Friday, a regression pages on-call on Saturday, and the lesson ends up in a chat thread no agent will read. The code was automated; the path to production and back was not.

This tutorial is for the developer who wires the pipeline, the tech lead who sets the classes and caps, and the CTO who signs off on what may reach production without a person reading it. It follows one change, cursor pagination on GET /orders, through every stage, and every file is meant to be copied.

What the intent-to-production pipeline gives you

Section titled “What the intent-to-production pipeline gives you”
  • A stage map from intent to production, with the file and the human role at each stage.
  • The integration fixes that make the issue-to-PR pipeline and the evidence bundle work as one.
  • merge-by-class.yml, which arms auto-merge for low pull requests only, from a job that never touches pull request code.
  • A deploy.yml plan job that picks the blast-radius budget from the merged pull request’s class and waits for a named approver on high.
  • slo-rollback.yml and open-revert.sh: an SLO alert restores the previous release, opens an incident issue and a revert pull request, and freezes deploys.
  • An incident-to-eval loop, three copy-paste prompts, five acceptance drills and the numbers that show the chain works.

What does the pipeline look like from intent to production?

Section titled “What does the pipeline look like from intent to production?”

The pipeline has eight stages. Read the last column twice: people stay in the loop where their judgment adds something, not at every diff.

StageWhat happensRepository fileWho decidesHuman role
1. IntentA product need becomes an agent-ready issue with acceptance criteria, allowed paths and a guardrail metric.github/ISSUE_TEMPLATE/agent-task.ymlThe issue ownerWrites or approves the intent, applies agent:ready
2. BuildThe agent implements it in an isolated run and a draft pull request opensagent-issue.yml, agent-loop.ymlThe workflowNone while the loop is inside its caps
3. ClassifyCI recomputes the risk class from the diff and labels itevidence-bundle.yml, check-evidence.mjsCI, from the base branch’s policyOwns the policy file
4. Mergelow auto-merges after an approval; standard waits for a reviewer on the bundle; high waits for a code ownermerge-by-class.yml, the ruleset, CODEOWNERSThe rulesetApproves on the bundle; code owners read high code
5. ReleaseThe class picks the exposure: full deploy, canary or staged rolloutdeploy.yml, scripts/deploy.shThe deploy planNamed approver promotes high
6. WatchGuardrails and SLO alerts judge the change on real trafficRollout analysis, alert rulesThe controller or the alertNone
7. Roll backTraffic returns to the previous release; a revert pull request and an incident issue openslo-rollback.yml, scripts/open-revert.shThe alert, through a dispatchOn-call reads the incident issue
8. LearnThe incident becomes an eval and, if needed, a policy changeevals/incidents/, evidence-policy.ymlThe pipeline, then a code ownerSigns off the eval and the policy change

The worked example moves through those stages in one morning. The times are illustrative, not a measurement.

TimeEvent
09:02Issue #418, “Paginate GET /orders with a cursor”, is labelled agent:ready; pull request #423 opens with its bundle at 09:31.
10:07CI classified it standard (behaviour change in covered code, new tests), so a reviewer approved on the bundle and merged. deploy.yml starts a 5% canary.
10:24The canary’s 5xx ratio breaches its guardrail. The controller aborts, the alert dispatches slo-breach, incident issue #431 and revert pull request #432 open, and deploys freeze.
11:40An eval task from #431 runs through the same pipeline; the eval fails on the bad commit, passes on the revert, and a code owner approves it.

Step 1: capture intent the pipeline can route

Section titled “Step 1: capture intent the pipeline can route”

The issue form from the issue-to-PR pipeline already asks for the outcome, acceptance criteria and allowed paths. Add two fields that the later stages need. The guardrail comes from the service’s SLO document and is written before the change exists, because an agent that picks its threshold after seeing canary numbers is grading its own work.

# Append to the body of .github/ISSUE_TEMPLATE/agent-task.yml
- type: dropdown
id: expected-risk
attributes:
label: Expected risk class
description: CI computes the real class from the diff. This can only raise it.
options: [low, standard, high]
validations: { required: true }
- type: textarea
id: guardrail
attributes:
label: Guardrail
description: "Service, SLO document and the metric that triggers rollback, e.g. orders-api, docs/slo/orders-api.md, 5xx ratio < 0.01 on the canary"
validations: { required: true }

The first prompt turns a product note into an issue that fills those fields. It works unchanged in Claude Code, Codex and Cursor.

Step 2: join the issue-to-PR pipeline and the evidence bundle

Section titled “Step 2: join the issue-to-PR pipeline and the evidence bundle”

The two pages were written to stand alone, so running them together needs three changes. Without them, every agent pull request fails the bundle check, the ruleset cannot tell two checks apart, or a label step fails.

  1. Put the bundle in the agent’s pull request. The issue-to-PR workflow builds the pull request body from the agent’s self-report, and the evidence checker looks for a YAML block that starts with evidence_bundle:. Add a step to .github/agent/implement.md so the agent writes one:

    8. Write .agent/bundle.yml in the format of the evidence_bundle block in
    .github/pull_request_template.md, starting with the line
    "evidence_bundle: 1". Copy the guardrail from the issue into
    rollout.guardrails. Declare the risk class from the issue or higher.

    Add .agent/bundle.yml to the artifact paths of the agent job, then append it in the open-pr job, after the self-report and before gh pr create:

    Terminal window
    if [ -s /tmp/agent/.agent/bundle.yml ]; then
    { echo; echo '## Evidence bundle'; echo; echo '~~~yaml'
    cat /tmp/agent/.agent/bundle.yml
    echo '~~~'; } >> body.md
    fi

    Do the same in push-fix with gh pr edit "$PR" --body-file, so a fix that changes the tests also updates the bundle. The edited trigger on evidence-bundle.yml then re-runs the check.

  2. Give each check its own name. The gates job in agent-loop.yml reports as evidence, and so does the evidence job in evidence-bundle.yml. Rename the first one to name: agent-gates, keep its job ID gates so the needs: references still work, and require both names in the ruleset. GitHub’s documentation does not say which result a ruleset reads when two jobs report the same check name, so do not leave it to chance.

  3. Create the new labels. revert, incident and incident:agent, next to the risk:* labels the evidence bundle already uses. gh does not create missing labels.

Step 3: merge by risk class with a ruleset and auto-merge

Section titled “Step 3: merge by risk class with a ruleset and auto-merge”

The ruleset decides whether a merge is allowed, CODEOWNERS decides who must read, and a small workflow asks GitHub to merge as soon as the ruleset allows it, for low only. Configure the ruleset on the default branch (Settings, then Rulesets):

RuleSettingThe hole it closes
Require a pull request before merging1 approval, dismiss stale approvals, require review from code ownersGitHub dismisses an approval “as stale” when the diff changes after it, so a fix pushed after approval cannot ride on it
Require status checks to passevidence, agent-gates and your test job; pin each to the GitHub Actions app as its expected sourceAny integration with write access can set a status; pinning the source means only your workflows can turn a check green
Allowed merge methodsSquash onlyA squash merge has one parent, so the git revert in step 5 needs no -m and cannot fail on a merge commit
Block force pushesOnHistory cannot be rewritten under an approved head
Bypass listEmpty, and never the release or agent appAn app on the bypass list merges without any of the above

Then add the workflow. It runs on workflow_run, after evidence-bundle finishes, so it executes from the default branch and never checks out the pull request. That is what lets it hold the release app’s token safely.

.github/workflows/merge-by-class.yml
name: merge-by-class
on:
workflow_run:
workflows: [evidence-bundle]
types: [completed]
permissions: {}
jobs:
route:
if: github.event.workflow_run.event == 'pull_request'
runs-on: ubuntu-latest
steps:
- uses: actions/create-github-app-token@v3
id: app
with:
client-id: ${{ vars.RELEASE_APP_CLIENT_ID }}
private-key: ${{ secrets.RELEASE_APP_PRIVATE_KEY }}
- name: Arm auto-merge for low risk only, pinned to the checked commit
env:
GH_TOKEN: ${{ steps.app.outputs.token }}
GH_REPO: ${{ github.repository }}
HEAD_SHA: ${{ github.event.workflow_run.head_sha }}
CONCLUSION: ${{ github.event.workflow_run.conclusion }}
run: |
pr=$(gh api "repos/$GH_REPO/commits/$HEAD_SHA/pulls" \
--jq '[.[] | select(.state == "open")] | first // {} | .number // empty')
[ -n "$pr" ] || { echo "No open pull request for $HEAD_SHA"; exit 0; }
classes=$(gh pr view "$pr" --json labels \
--jq '[.labels[].name | select(startswith("risk:"))] | join(",")')
if [ "$CONCLUSION" = "success" ] && [ "$classes" = "risk:low" ]; then
gh pr merge "$pr" --auto --squash --match-head-commit "$HEAD_SHA"
else
gh pr merge "$pr" --disable-auto || true
fi

Three details carry the safety. --match-head-commit merges only the commit that was classified. The else branch disarms auto-merge when a later push raises the class. And the merge uses a GitHub App token because, per GitHub’s documentation, events created with the repository’s GITHUB_TOKEN do not start new workflow runs, so a merge made with it would never start deploy.yml. Create the release app with Contents, Pull requests and Issues write access on this repository only, turn on Allow auto-merge in the repository settings, and keep the app out of CODEOWNERS.

The ruleset still requires one approval, even for low. Who gives it is the one place the three tools differ.

Claude Code does not approve pull requests. Its managed Code Review posts findings, and its check run “always completes with a neutral conclusion so it never blocks merging through branch protection rules” (Claude Code documentation, checked 2026-09-26). Feed those findings to the standard reviewer, and have a person approve low pull requests on the bundle. If you build your own approver identity, confirm on a test branch that the ruleset counts its approval. Setup is on review automation with Claude Code.

Step 4: release each class with its blast-radius budget

Section titled “Step 4: release each class with its blast-radius budget”

The deploy workflow reads the class of the pull request that produced the commit, not a class the agent declared. The budgets (first exposure, promotion steps, bake time and rollback trigger per class) are defined on progressive delivery for agent-written changes; this workflow only chooses between them.

# .github/workflows/deploy.yml: plan and approval jobs. Build and push your image before "deploy".
name: deploy
on:
push:
branches: [main]
concurrency:
group: deploy-production
cancel-in-progress: false
permissions: {}
jobs:
plan:
runs-on: ubuntu-latest
permissions:
contents: read
pull-requests: read
outputs:
risk: ${{ steps.plan.outputs.risk }}
steps:
- id: plan
env:
GH_TOKEN: ${{ github.token }}
GH_REPO: ${{ github.repository }}
run: |
# An open revert means production runs an older release on purpose.
if [ "$(gh pr list --label revert --state open --json number --jq length)" != "0" ]; then
echo "An open revert pull request freezes deploys. Merge or close it first."; exit 1
fi
# No merged pull request, or no class label, means high.
read -r pr risk < <(gh api "repos/$GH_REPO/commits/$GITHUB_SHA/pulls" \
--jq '[.[] | select(.merged_at != null)] | first // {}
| "\(.number // "none") \(([.labels[]?.name | select(startswith("risk:"))] | first) // "risk:high")"')
echo "Pull request $pr, class ${risk#risk:}"
echo "risk=${risk#risk:}" >> "$GITHUB_OUTPUT"
approve-high:
needs: plan
if: needs.plan.outputs.risk == 'high'
runs-on: ubuntu-latest
environment: production-approval # required reviewers, no secrets
steps:
- run: echo "Staged rollout of $GITHUB_SHA approved"
deploy:
needs: [plan, approve-high]
if: >-
always() && needs.plan.result == 'success' &&
(needs.approve-high.result == 'success' || needs.approve-high.result == 'skipped')
runs-on: ubuntu-latest
environment: production
permissions:
contents: read
steps:
- uses: actions/checkout@v7
with: { persist-credentials: false }
- name: Deploy with this class's budget
env:
RISK: ${{ needs.plan.outputs.risk }}
run: scripts/deploy.sh "$RISK" "$GITHUB_SHA"

scripts/deploy.sh is yours: it maps low to a normal rollout, standard to the analysed canary and high to the staged rollout. Keep the mapping in that one script so a budget change is one reviewed diff. The prompt below drafts it with the step 5 rollback script, in any of the three tools.

The production-approval environment holds the production approval gate for high: up to six required reviewers, with self-review prevented, so the person who merged cannot also release. On the Free, Pro and Team plans, required reviewers for environments are available only in public repositories (GitHub documentation, checked 2026-09-26), so check your plan before you design around them.

Step 5: roll back on an SLO breach without waiting for a person

Section titled “Step 5: roll back on an SLO breach without waiting for a person”

A canary controller catches a bad release below full exposure and has already moved traffic back. An SLO alert catches a low change that went straight to 100%, and nothing has moved yet. Both send the same event to GitHub, and one workflow handles the rest.

Your alerting system (Alertmanager and Grafana alerting can both send a webhook) calls a small relay, which sends a repository dispatch carrying the service, the commit SHA the service reports as its version, the alert name, and the exposure, canary or full:

Terminal window
# In the relay. GITHUB_DISPATCH_TOKEN comes from the relay's secret store.
curl -sS -X POST "https://api.github.com/repos/acme/shop/dispatches" \
-H "Accept: application/vnd.github+json" \
-H "Authorization: Bearer $GITHUB_DISPATCH_TOKEN" \
-d '{"event_type":"slo-breach","client_payload":{"service":"orders-api","sha":"'"$VERSION_SHA"'","alert":"OrdersErrorBudgetBurn","stage":"canary"}}'

Issue that token from a GitHub App installed on this repository alone, with the minimum permission GitHub’s REST reference lists for creating a repository dispatch event. A repository_dispatch workflow always runs from the default branch, so the payload can choose values but never code.

.github/workflows/slo-rollback.yml
name: slo-rollback
on:
repository_dispatch:
types: [slo-breach]
concurrency:
group: slo-rollback-${{ github.event.client_payload.service }}
cancel-in-progress: false
permissions: {}
jobs:
restore:
# A canary abort has already restored traffic; a full deploy has not.
if: github.event.client_payload.stage == 'full'
runs-on: ubuntu-latest
environment: production-rollback # branch rule only, no reviewers: a rollback must not wait
permissions:
contents: read
steps:
- uses: actions/checkout@v7
with: { persist-credentials: false }
- name: Return the service to its previous release, with no build
env:
SERVICE: ${{ github.event.client_payload.service }}
run: |
[[ "$SERVICE" =~ ^[a-z0-9-]{1,40}$ ]] || { echo "Invalid service"; exit 1; }
scripts/rollback.sh "$SERVICE"
record:
needs: restore
if: always() # open the incident and the revert even if the restore failed
runs-on: ubuntu-latest
permissions:
contents: read
steps:
- uses: actions/create-github-app-token@v3
id: app
with:
client-id: ${{ vars.RELEASE_APP_CLIENT_ID }}
private-key: ${{ secrets.RELEASE_APP_PRIVATE_KEY }}
- uses: actions/checkout@v7
with:
fetch-depth: 0
token: ${{ steps.app.outputs.token }}
- name: Open the incident issue and the revert pull request
env:
GH_TOKEN: ${{ steps.app.outputs.token }}
SHA: ${{ github.event.client_payload.sha }}
SERVICE: ${{ github.event.client_payload.service }}
ALERT: ${{ github.event.client_payload.alert }}
STAGE: ${{ github.event.client_payload.stage }}
RESTORE: ${{ needs.restore.result }}
HUMAN_OWNER: ${{ vars.ONCALL_OWNER }}
RUN_URL: ${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }}
run: scripts/open-revert.sh

scripts/rollback.sh wraps your platform’s own undo, for example kubectl rollout undo deployment/"$SERVICE" for a plain Kubernetes Deployment. It must not build anything: a rollback that waits for CI is a second incident. The production-rollback environment holds the same deploy credentials as production, with a deployment-branch rule and no required reviewers.

The script opens the incident issue first, so the revert’s bundle can link to it, then commits the revert and opens its pull request with a bundle generated from the diff:

#!/usr/bin/env bash
# scripts/open-revert.sh: run by slo-rollback.yml on the default branch with the release
# app's token in GH_TOKEN. Opens the incident issue and a revert pull request for the
# merge that breached an SLO. It never builds or runs the reverted code.
set -euo pipefail
[[ "$SHA" =~ ^[0-9a-f]{7,40}$ ]] || { echo "Invalid sha: $SHA"; exit 1; }
[[ "$SERVICE" =~ ^[a-z0-9-]{1,40}$ ]] || { echo "Invalid service: $SERVICE"; exit 1; }
SHA=$(git rev-parse --verify "$SHA^{commit}")
# Provenance: the pull request that merged the commit, and whether an agent wrote it.
pr_json=$(gh api "repos/$GITHUB_REPOSITORY/commits/$SHA/pulls" \
--jq '[.[] | select(.merged_at != null)] | first // {}')
pr=$(jq -r '.number // empty' <<<"$pr_json")
labels=$(jq -r '[.labels[]?.name] | join(",")' <<<"$pr_json")
# Same files, same computed floor; the reverted class is at least that floor, so declare it.
class=$(jq -r '[.labels[]?.name | select(startswith("risk:"))] | first // "risk:high"' <<<"$pr_json")
class=${class#risk:}
kind=incident
[[ ",$labels," == *",agent-pr,"* ]] && kind=incident:agent
cat > /tmp/incident.md <<MD
**Alert:** $ALERT on \`$SERVICE\` · **Stage:** $STAGE · **Restore job:** $RESTORE
**Pull request:** ${pr:+#$pr} (\`${SHA:0:12}\`) · **Labels:** ${labels:-none} · **Run:** $RUN_URL
Close this issue only when each box is ticked with a link:
- [ ] Eval or test in \`evals/incidents/\` that fails on \`${SHA:0:12}\` and passes on the revert
- [ ] Policy change, if the risk class or the guardrail let this through
- [ ] Loop level decision, if an agent loop produced the change
MD
issue_url=$(gh issue create --title "SLO breach on $SERVICE after ${pr:+#$pr }${SHA:0:12}" \
--label "$kind" --body-file /tmp/incident.md)
# The revert, and one check a reviewer can trust: every file the commit touched is back to its parent.
branch="revert/${SHA:0:12}"
git switch -c "$branch"
git -c user.name="release-bot" -c user.email="release-bot@users.noreply.github.com" \
revert --no-edit "$SHA"
mapfile -t files < <(git diff --name-only "$SHA^" "$SHA")
if git diff --quiet "$SHA^" HEAD -- "${files[@]}"; then restored=pass; code=0; else restored=fail; code=1; fi
{
echo "Reverts ${pr:+#$pr }(\`${SHA:0:12}\`) after an SLO breach. Incident: $issue_url"
echo
echo '~~~yaml'
echo 'evidence_bundle: 1'
echo 'spec:'
echo " link: $issue_url"
echo " delta: [\"$SERVICE behaves as it did before ${SHA:0:12}\"]"
echo ' unrequested: []'
echo 'acceptance:'
echo " - criterion: every file ${SHA:0:12} changed matches its parent commit"
echo " check: git diff --quiet ${SHA:0:12}^ HEAD -- <the ${#files[@]} files>"
echo " result: $restored"
echo 'checks:'
echo " - command: git diff --quiet ${SHA:0:12}^ HEAD -- <the ${#files[@]} files>"
echo " exit_code: $code"
echo 'risk:'
echo " class: $class # the reverted pull request's class; CI fails anything lower"
echo ' touches: [auth, money, schema, migrations, infra] # over-declared on purpose'
echo ' oracle_changes:'
for f in "${files[@]}"; do
printf ' - { path: "%s", direction: neutral, reason: "revert of %s" }\n' "$f" "${SHA:0:12}"
done
echo " rollback: revert this revert once the incident eval passes"
echo 'provenance:'
echo ' agent: slo-rollback workflow (no agent)'
echo ' model: none'
echo " session: $RUN_URL"
echo " task: $issue_url"
echo " human_owner: \"$HUMAN_OWNER\""
echo '~~~'
} > /tmp/revert.md
git push -u origin "$branch"
gh pr create --head "$branch" --title "Revert ${pr:+#$pr }(${SHA:0:12}): SLO breach on $SERVICE" \
--body-file /tmp/revert.md --label revert

Two choices in the generated bundle need a word. It declares every sensitive class and lists every reverted file as a neutral oracle change, because the checker fails a bundle that omits a touched class or oracle file, and the floor still comes from the diff: a reverted test file lifts it to at least standard. And it declares the class of the pull request it reverts, which is at least the floor the checker computes for the same files, so a docs revert stays low and can auto-merge, while a billing revert is high and waits for its code owner. If the reverted pull request carries no risk:* label, the script declares high. That wait is safe, because traffic is already back on the previous release and deploy.yml refuses to ship main while the revert is open.

Step 6: turn the incident into an eval through the same pipeline

Section titled “Step 6: turn the incident into an eval through the same pipeline”

The incident issue closes only when its lesson exists as a check. The eval is an agent task that goes through stages 1 to 4 like any other; the difference is the proof: it must fail on the bad commit and pass on the revert. Postmortems, containment and loop demotion are on when an agent causes an incident.

Add /evals/ @acme/tech-leads to CODEOWNERS, so a person signs every eval. Then turn the incident into a task:

A person reviews the issue, applies agent:ready, and the pipeline builds the eval. Before approving that pull request, the code owner runs the proof, because the bundle can only claim it:

Terminal window
# Terminal, repository root. BAD_SHA is the reverted commit from the incident issue.
git worktree add /tmp/inc-431 "$BAD_SHA"
mkdir -p /tmp/inc-431/evals/incidents
cp evals/incidents/INC-431.test.ts /tmp/inc-431/evals/incidents/
(cd /tmp/inc-431 && npm ci && npm test -- evals/incidents/INC-431.test.ts) # must fail
npm test -- evals/incidents/INC-431.test.ts # must pass
git worktree remove --force /tmp/inc-431

Only then does the change return: revert the revert, fix the defect so the eval passes, and send it through the pipeline again. The eval stays in the suite and runs on every harness change, as continuous evals describes.

The tools differ only at the trigger. For local runs, save the prompt as .github/agent/incident-to-task.md with the incident number filled in.

Label the eval issue agent:ready, and agent-issue.yml runs it through anthropics/claude-code-action@v1 like any other task. To draft the task locally, run the prompt headless; claude -p starts in Manual permission mode, so tools outside the allowlist are denied:

Terminal window
# Terminal, repository root (Claude Code 2.1.283)
claude -p "$(cat .github/agent/incident-to-task.md)" \
--allowedTools "Read,Grep,Glob,Bash(gh issue view *),Bash(gh pr view *)" \
--max-budget-usd 2

How do you prove the whole pipeline works before it deploys anything?

Section titled “How do you prove the whole pipeline works before it deploys anything?”

Run these five drills against a staging service with the same ruleset and workflows. Each proves one link in the chain.

DrillWhat you doPass when
Low pathAn agent fixes a typo in docs/setup.mdLabelled risk:low, approved, merged automatically, deploy.yml runs a normal rollout
Class cannot be loweredAn agent changes one line in src/billing/refund.ts and declares class: lowCI labels risk:high, auto-merge stays off, review is requested from the code owner, approve-high waits
Canary abortShip a standard build that returns HTTP 500 on 5% of GET /orders requestsThe analysis aborts, the dispatch arrives with stage: canary, an incident issue and a revert pull request open, restore is skipped
Full rollbackShip a low change and fire the SLO alert by handrestore runs rollback.sh with no build, the revert opens, and the next push to main fails the plan job with the freeze message
Eval proofRun the incident-to-eval task for the canary drillThe eval fails on the bad commit and passes on main, and a code owner approves it

Time the canary and full-rollback drills from bad deploy to zero exposure: that, not your budget table, is how long a defect reaches customers.

Which numbers show intent-to-production is working?

Section titled “Which numbers show intent-to-production is working?”

These measures come from labels, deployments and incident issues you already have. The tech lead reviews them monthly; the CTO reads the trend beside metrics frameworks.

MeasureDefinitionWhat to do with it
Intent-to-production lead timeFrom agent:ready to 100% exposure, median, per classCompare with human-authored changes of the same class; the gap shows the queue
Change failure rate by classMerged changes that caused a rollback, a revert or an incident, divided by merged changes, per classlow above standard means the low rules are too wide
Caught before 100%Rollbacks that happened at canary or staged exposure, divided by all rollbacksLow values mean customers find defects first; add guardrails
Time to zero exposureFrom the first guardrail breach or alert to 0% traffic on the change, medianHours mean the rollback waits for a build or a person
Incident-to-eval closureIncident issues closed with a proven eval or an enforced policy change, divided by incident issuesBelow 100% means lessons still live in chat threads

DORA’s 2025 report (Google Cloud, 23 September 2025) found “a positive relationship between AI adoption on both software delivery throughput and product performance”, and also that “AI adoption does continue to have a negative relationship with software delivery stability.” Faros AI’s AI Engineering Report 2026 (April 2026, vendor telemetry from 22,000 developers) measured task throughput per developer up 33.7% alongside incidents per pull request up 242.7%. These measures show which side of those findings your pipeline is on.

Sign-off stays with named people. The issue owner signs the intent; the tech lead owns the policy file, caps and risk table; a code owner reads every high change and every eval. The CTO owns which classes reach production without a person reading code, and changes that one class at a time, on these numbers.

Every agent pull request fails evidence with “No evidence bundle found”. The body was built without the bundle. Recovery: apply step 2 and edit the open pull requests’ bodies; the edited trigger re-runs the check.

Auto-merge happens but no deploy starts. The merge ran with GITHUB_TOKEN, so the push event started no workflow. Recovery: arm auto-merge with the release app’s token, as merge-by-class.yml does.

A low pull request never merges. Nobody approved it, or a rule adds one. GitHub’s rulesets require one extra approval by default for Copilot cloud agent pull requests that are not attributed to a person (public preview, GitHub documentation checked 2026-09-26), so those need two. Recovery: decide who approves low (step 3 tabs), and recheck the ruleset when a new agent starts opening pull requests.

A merge queue stalls every pull request. The required checks never ran on merge_group, the event GitHub’s merge queue documentation says you must use to trigger Actions workflows for queued pull requests. The evidence checker reads the pull request body, which that event does not carry. Recovery: add the merge_group trigger to each required workflow, with a bundle job that takes the pull request number from the queue branch name (gh-readonly-queue/<base>/pr-<number>-<sha>), reads that pull request’s head SHA, and passes only if gh api repos/$GH_REPO/commits/$SHA/check-runs reports a successful evidence run for it. Or turn the queue off.

The rollback waited for an approver. The restore job used the production environment, whose reviewers apply to rollbacks too. Recovery: give rollback its own production-rollback environment with a branch rule and no reviewers, and rehearse it with the full-rollback drill.

The revert pull request fails its own bundle check. The reverted change touched UI paths that need runtime evidence the script cannot produce, or later commits changed the same files. Recovery: production is already restored, so a person finishes the revert by hand and records why in the incident issue.

The dispatch arrives with the wrong SHA. The service reports a build number rather than a commit, or the relay reads the stable pods instead of the canary. Recovery: export the commit SHA as a version label on every pod and read it from the canary’s series; the script rejects anything that is not a commit in the repository.

Incidents close without an eval. On-call ticks the boxes to clear the queue. Recovery: someone other than the incident owner checks the links, and the closure measure goes on the monthly review.

Where to go next from intent to production

Section titled “Where to go next from intent to production”