Skip to content

The evidence bundle: what an agent's pull request must prove

An evidence bundle is a structured block in an agent’s pull request that proves the change: the spec link and behavior delta, each acceptance criterion mapped to a check that ran, command results, runtime or eval evidence, the risk class, sensitive paths, oracle changes, and provenance. A CI check parses the bundle and blocks the merge when it is incomplete.

Your team’s agents open a dozen pull requests a day, and every description reads the same: “Implemented pagination, added tests, all checks pass.” Three of them quietly edited a test fixture, one touched the billing webhook, and none says which acceptance criterion each test proves. You cannot tell the safe ones from the risky ones without reading every diff, which is the work the agents were supposed to save.

This page is for developers who configure the agents, tech leads who own the review policy, and CTOs who need one auditable record per change. It is the site’s canonical change-manifest schema: change provenance and risk routing routes on these fields, and reading evidence instead of code explains how a reviewer reads them.

  • A field-by-field schema for the bundle, with who fills each field and how CI verifies it.
  • A pull request template your agents and your people fill in the same way.
  • A tested CI check (about 80 lines of Node) that fails an incomplete bundle and recomputes the risk class from the diff.
  • A routing table that turns the risk class into who approves and who must read code.
  • Three copy-paste prompts: produce the bundle, audit a bundle against the diff, and repair a failing check.
  • The failure modes of a bundle gate, and how to recover from each.

Why a pull request description is not evidence

Section titled “Why a pull request description is not evidence”

A description is prose the agent wrote about its own work. It can be wrong in ways nobody sees until production. The load that makes this matter is measured: Faros AI’s Acceleration Whiplash report (April 2026; telemetry from 22,000 developers and more than 4,000 teams; vendor telemetry from Faros customers) found median time in review up 441.5% and 31.3% more pull requests merging with no review at all. Generation scaled; reading did not.

The fix is to require output from checks instead of claims about them. As one practitioner put it, “A model can argue that its work is complete. A deterministic validator can prove that a required field is missing.” (NARESH, Graph Engineering for AI Coding Agents, DEV Community, 30 July 2025.) The bundle is what that validator reads.

The bundle has seven sections. The Verified by CI column is the point of the design: every field the agent could get wrong in its own favor is either recomputed from the diff or checked against the branch.

SectionFieldsFilled byVerified by CIRequired when
Specspec.link, spec.delta (plain-sentence behavior changes), spec.unrequestedAgent, from the ticket or spec.mdLink is a URL or a file that exists in the branch; delta is not emptyAlways
AcceptanceOne entry per criterion: criterion, check (file:line or a command), resultAgent, after running the checksEvery result is pass; every cited file existsAlways
Checkscommand, exit_code, optional summary (for example a mutation score)AgentEvery exit_code is 0. CI also runs its own required jobs; the bundle never replaces themAlways
Runtimekind (screenshot, trace, video, preview-url), ref, covers (criterion IDs)Agent, from a preview or browser runPresent when UI paths changedPolicy globs match
Evalssuite, score, baselineAgent or CI eval jobscore ≥ baselinePrompt, model-config or agent-harness paths changed
Riskclass (low, standard, high), touches (sensitive classes), oracle_changes (path, direction, reason), rollbackAgent proposesCI recomputes the sensitive classes and oracle files from the diff, sets a floor, and fails a lower declarationAlways
Provenanceagent (tool and version), model, session, task, human_ownerAgent, plus the human who owns the changeagent, model and human_owner presentAlways

Three rules keep the schema honest:

  1. The declared risk class can only raise the computed floor. An agent that writes class: low on a billing change fails the check; a person who marks a harmless change high gets a stricter review, which is allowed.
  2. An oracle change is never neutral by default. Any edited test, snapshot, CI workflow, lint or type config must be listed with a direction of stricter, looser or neutral. A looser entry forces high, because it changes what “green” means for every other change.
  3. Provenance is recorded, not routed on. Human, AI-assisted and agent changes use the same bundle and the same routing. The provenance and routing page explains why authorship is a poor risk proxy.

The setup is five files. Add them in one pull request, merge it, then turn the check on as required. The order matters, as the failure-modes section explains.

  1. Add the pull request template. GitHub pre-fills the description of pull requests opened in the web UI from .github/pull_request_template.md. Agents that open pull requests from the command line get the same structure from the bundle prompt further down, which points them at this file.

    .github/pull_request_template.md
    ## What changed and why
    <!-- Two or three sentences for a human. The bundle below is the evidence. -->
    ## Evidence bundle
    ~~~yaml
    evidence_bundle: 1
    spec:
    link: docs/specs/orders-pagination.md # or the issue URL
    delta:
    - GET /orders returns 50 items per page and a next_cursor
    - The old page parameter returns 400
    unrequested: [] # behavior nobody asked for
    acceptance:
    - criterion: "AC1: first page has 50 items and a next_cursor"
    check: tests/orders.contract.test.ts:42
    result: pass # pass | fail | unverified
    - criterion: "AC2: the page parameter returns 400"
    check: tests/orders.contract.test.ts:88
    result: pass
    checks:
    - command: npm test
    exit_code: 0
    - command: npx stryker run --mutate "src/orders/**/*.ts"
    exit_code: 0
    summary: 81% mutation score on changed files
    evals: [] # suite, score, baseline
    runtime: [] # kind, ref, covers
    risk:
    class: standard # low | standard | high
    touches: [] # auth, money, schema, migrations, infra
    oracle_changes:
    - path: tests/orders.contract.test.ts
    direction: stricter # stricter | looser | neutral
    reason: new contract cases for AC1 and AC2
    rollback: revert the merge commit; no migration, no flag
    provenance:
    agent: Claude Code 2.1.283
    model: claude-opus-5-5
    session: https://claude.ai/code/session_EXAMPLE
    task: https://github.com/acme/shop/issues/412
    human_owner: "@anna"
    ~~~
  2. Write the policy file. It names the sensitive classes with their minimum risk, the oracle paths, and the paths that need runtime evidence or evals. Adjust the globs to your repository; these match the escalation classes the rest of the site uses.

    # .github/evidence-policy.yml (owned by CODEOWNERS; editing it is an oracle change)
    sensitive:
    auth: { globs: ['**/auth/**', '**/middleware/**', '**/*permission*'], min_risk: high }
    money: { globs: ['**/billing/**', '**/payments/**', '**/pricing/**'], min_risk: high }
    schema: { globs: ['**/*.sql', '**/openapi*', '**/*.proto'], min_risk: high }
    migrations: { globs: ['**/migrations/**'], min_risk: high }
    infra: { globs: ['infra/**', '**/*.tf', '.github/workflows/**'], min_risk: standard }
    oracle:
    - '**/*.test.*'
    - '**/__snapshots__/**'
    - '.github/workflows/**'
    - '.github/evidence-policy.yml'
    - '.github/CODEOWNERS'
    - 'scripts/check-evidence.mjs'
    - 'tsconfig*.json'
    - 'eslint.config.*'
    runtime_required:
    - 'src/components/**'
    - 'src/pages/**'
    evals_required:
    - 'prompts/**'
  3. Add the checker. It reads the pull request body from the GitHub event, parses the first yaml block that starts with evidence_bundle:, and compares it with the diff. It depends on yaml (2.9.1 on npm) and minimatch (10.2.6 on npm), both checked on 2026-09-26.

    // scripts/check-evidence.mjs: fails a pull request whose evidence bundle is
    // missing, incomplete, or contradicted by the diff. Usage:
    // node check-evidence.mjs <policy.yml>
    import { appendFileSync, existsSync, readFileSync } from 'node:fs';
    import { execFileSync } from 'node:child_process';
    import { parse } from 'yaml';
    import { minimatch } from 'minimatch';
    const RANK = { low: 0, standard: 1, high: 2 };
    const policy = parse(readFileSync(process.argv[2], 'utf8'));
    const pr = JSON.parse(readFileSync(process.env.GITHUB_EVENT_PATH, 'utf8')).pull_request;
    const errors = [];
    const fail = (msg) => errors.push(msg);
    // 1. What CI knows without asking the agent: the changed files.
    const changed = (process.env.EVIDENCE_CHANGED_FILES ??
    execFileSync('git', ['diff', '--name-only', `${pr.base.sha}...HEAD`], { encoding: 'utf8' }))
    .split('\n').filter(Boolean);
    const touching = (globs) => changed.filter((f) => globs.some((g) => minimatch(f, g, { dot: true })));
    // 2. The bundle: the first yaml fence (backticks or tildes) that starts with evidence_bundle:
    const block = (pr.body ?? '').match(/(`{3}|~{3})ya?ml\r?\n(evidence_bundle:[\s\S]*?)\1/);
    if (!block) {
    console.error('No evidence bundle found. Fill in the template from .github/pull_request_template.md.');
    process.exit(1);
    }
    const b = parse(block[2]) ?? {};
    // 3. Spec link and delta
    if (!b.spec?.link) fail('spec.link is empty');
    else if (!/^https?:\/\//.test(b.spec.link) && !existsSync(b.spec.link)) fail(`spec.link ${b.spec.link} does not exist`);
    if (!b.spec?.delta?.length) fail('spec.delta lists no behavior change');
    // 4. Acceptance: every criterion names a check that exists and passed
    if (!b.acceptance?.length) fail('acceptance is empty');
    for (const a of b.acceptance ?? []) {
    if (a.result !== 'pass') fail(`acceptance "${a.criterion}" is ${a.result ?? 'missing a result'}`);
    const ref = /^([^\s:]+):(\d+)$/.exec(a.check ?? '');
    if (!a.check) fail(`acceptance "${a.criterion}" names no check`);
    else if (ref && !existsSync(ref[1])) fail(`acceptance "${a.criterion}" cites ${ref[1]}, which is not in the branch`);
    }
    for (const c of b.checks ?? []) if (c.exit_code !== 0) fail(`check "${c.command}" exited ${c.exit_code}`);
    if (!b.checks?.length) fail('checks lists no command that ran');
    // 5. Sensitive paths: CI recomputes them and sets the risk floor
    let floor = 'low';
    const declared = new Set(b.risk?.touches ?? []);
    for (const [name, rule] of Object.entries(policy.sensitive)) {
    const hits = touching(rule.globs);
    if (!hits.length) continue;
    if (!declared.has(name)) fail(`diff touches ${name} (${hits.join(', ')}) but risk.touches omits it`);
    if (RANK[rule.min_risk] > RANK[floor]) floor = rule.min_risk;
    }
    // 6. Oracle changes: every edited test, CI or lint file is declared with a direction
    const listed = new Map((b.risk?.oracle_changes ?? []).map((o) => [o.path, o.direction]));
    for (const f of touching(policy.oracle)) {
    const dir = listed.get(f);
    if (!dir) fail(`oracle file ${f} changed but is not in risk.oracle_changes`);
    else if (!['stricter', 'looser', 'neutral'].includes(dir)) fail(`oracle change ${f} has direction ${dir}; use stricter, looser or neutral`);
    if (RANK[floor] < RANK.standard) floor = 'standard';
    if (dir === 'looser') floor = 'high';
    }
    // 7. Runtime evidence and evals, where the policy requires them
    if (touching(policy.runtime_required).length && !b.runtime?.length) fail('UI paths changed but runtime lists no screenshot or trace');
    if (touching(policy.evals_required).length) {
    if (!b.evals?.length) fail('prompt or model paths changed but evals is empty');
    for (const e of b.evals ?? []) if (!(e.score >= e.baseline)) fail(`eval ${e.suite} scored ${e.score} against baseline ${e.baseline}`);
    }
    // 8. Risk class and provenance
    if (!Object.hasOwn(RANK, b.risk?.class ?? '')) fail('risk.class must be low, standard or high');
    else if (RANK[b.risk.class] < RANK[floor]) fail(`risk.class is ${b.risk.class}, but the diff requires at least ${floor}`);
    if (!b.risk?.rollback) fail('risk.rollback is empty');
    for (const k of ['agent', 'model', 'human_owner']) if (!b.provenance?.[k]) fail(`provenance.${k} is empty`);
    const risk = RANK[b.risk?.class] > RANK[floor] ? b.risk.class : floor;
    if (process.env.GITHUB_OUTPUT) appendFileSync(process.env.GITHUB_OUTPUT, `risk=${risk}\n`);
    console.log(`Changed files: ${changed.length}. Risk class: ${risk}.`);
    if (errors.length) {
    console.error(`Evidence bundle incomplete (${errors.length}):\n- ${errors.join('\n- ')}`);
    process.exit(1);
    }
    console.log('Evidence bundle complete.');
  4. Add the workflow. Two details make it trustworthy. It loads the checker and the policy from the base commit, so a pull request cannot weaken the checker or the policy that judge it. The workflow file itself runs from the pull request, which is why step 5 gives .github/workflows/ a code owner and requires code-owner review (or you make it a ruleset-required workflow). Checkout sets persist-credentials: false, so the token never lands in .git/config; gh pr edit reads GH_TOKEN from its own step. And it runs on edited, so fixing the description re-runs the check without a new push.

    .github/workflows/evidence-bundle.yml
    name: evidence-bundle
    on:
    pull_request:
    types: [opened, edited, synchronize, reopened, ready_for_review]
    permissions:
    contents: read
    pull-requests: write
    jobs:
    evidence:
    runs-on: ubuntu-latest
    steps:
    - uses: actions/checkout@v7
    with:
    fetch-depth: 0
    persist-credentials: false
    - uses: actions/setup-node@v7
    with:
    node-version: 24
    - name: Load the checker and policy from the base branch
    run: |
    mkdir -p /tmp/evidence
    git show "${{ github.event.pull_request.base.sha }}:scripts/check-evidence.mjs" > /tmp/evidence/check.mjs
    git show "${{ github.event.pull_request.base.sha }}:.github/evidence-policy.yml" > /tmp/evidence/policy.yml
    npm install --prefix /tmp/evidence --no-save yaml@2.9.1 minimatch@10.2.6
    - name: Check the evidence bundle
    id: check
    run: node /tmp/evidence/check.mjs /tmp/evidence/policy.yml
    - name: Label the risk class
    if: always() && steps.check.outputs.risk != ''
    env:
    GH_TOKEN: ${{ github.token }}
    run: |
    RISK="${{ steps.check.outputs.risk }}"
    STALE=$(printf 'risk:%s\n' low standard high | grep -vx "risk:$RISK" | paste -sd, -)
    gh pr edit "${{ github.event.pull_request.number }}" --remove-label "$STALE" --add-label "risk:$RISK"

    Create the labels risk:low, risk:standard and risk:high once; gh pr edit --add-label does not create missing labels. The step also removes the other two risk:* labels, because --add-label alone never takes one off: a pull request that moves from standard to high after a push would otherwise carry both, and anything that routes on labels would read the stale class.

  5. Make the check required and protect the gate. In a branch ruleset for main, require the evidence status check and require review from code owners. Then give an owner who is not an agent to three groups of paths: the gate’s own files (including CODEOWNERS itself, the prompts and the pull request template, because GitHub applies the base branch’s CODEOWNERS, so a pull request that edits it would otherwise drop its own owner unreviewed), every min_risk: high glob, and the protected oracles: the acceptance or contract tests that encode the spec, plus the repository-wide type and lint config. That is what enforces the high row of the routing table for those paths: the required code-owner review applies to any pull request touching them, whatever label it carries.

    Do not give every **/*.test.* or snapshot path an owner. Ordinary tests change in most standard pull requests, and owning them would put every one of those in front of a code owner, which removes the one-reviewer lane. Those edits are still declared in risk.oracle_changes and read by the standard reviewer. One limit to know: a looser edit to an unowned test raises the class to high and the label to risk:high, but CODEOWNERS cannot see labels, so nothing forces a code owner onto it. The reviewer escalates it by hand, and a test whose loosening would be expensive belongs under a protected path.

    # .github/CODEOWNERS
    .github/evidence-policy.yml @acme/platform
    .github/workflows/ @acme/platform
    scripts/check-evidence.mjs @acme/platform
    .github/CODEOWNERS @acme/platform
    .github/prompts/ @acme/platform
    .github/pull_request_template.md @acme/platform
    # Every min_risk: high glob in the policy
    **/auth/** @acme/security
    **/middleware/** @acme/security
    **/*permission* @acme/security
    **/billing/** @acme/payments
    **/payments/** @acme/payments
    **/pricing/** @acme/payments
    **/*.sql @acme/data
    **/openapi* @acme/data
    **/*.proto @acme/data
    **/migrations/** @acme/data
    # Protected oracles only: spec-level tests and repo-wide config
    **/*.contract.test.* @acme/platform
    tsconfig*.json @acme/platform
    eslint.config.* @acme/platform

Test the gate the way you test any oracle: with fixtures that must pass and fixtures that must fail. The EVIDENCE_CHANGED_FILES variable replaces git diff, so a fixture needs only a fake event file and a list of paths:

Terminal window
# Terminal, repository root. ev-billing.json holds {"pull_request":{"base":{"sha":"x"},"body":"..."}}
GITHUB_EVENT_PATH=fixtures/ev-billing.json \
EVIDENCE_CHANGED_FILES=$'src/billing/refund.ts\nsrc/pages/orders.astro' \
node scripts/check-evidence.mjs .github/evidence-policy.yml

Run against the template above with a billing file and a UI page changed, and with the cited spec and test files (docs/specs/orders-pagination.md, tests/orders.contract.test.ts) present in the branch, the checker exits 1 with three errors: risk.touches omits money, runtime is empty, and risk.class is standard, but the diff requires at least high. Keep one fixture per rule (missing bundle, unverified criterion, cited test file missing, undeclared oracle change, looser oracle change, missing eval) and run them in your normal test job. A gate change that breaks a fixture is caught before it reaches main.

The bundle decides who reviews, and what they read. This table is a starting policy; the organization-wide version belongs in the autonomy and risk-class policy.

Risk classTypical triggerApprovalWho reads codeAfter merge
lowDocs, copy, isolated UI with runtime evidence, no oracle changesOne reviewer on the bundle, or an approval rule (see the Cursor tab)Nobody by default; sampled per the trust logNormal deploy
standardBehavior change in covered code; stricter or neutral oracle changes; infraOne reviewer on the bundle, plus a review agent’s findingsThe reviewer reads the oracle changes and any unrequested behaviorCanary or flag where available
highAuth, money, schema, migrations, a looser oracle changeCode owner from CODEOWNERSA named code owner reads the code, in addition to the bundleStaged rollout and a production approval gate

Paths listed in CODEOWNERS (the gate’s own files, the sensitive globs and the protected oracles) need their code owner whatever the class. A stricter or neutral edit to an ordinary test or snapshot stays in the one-reviewer lane; a looser one is labelled risk:high, and the reviewer has to escalate it to a code owner by hand, because CODEOWNERS does not see labels.

standard corresponds to medium in the organization policy; critical changes are high here plus the production approval gate.

A reviewer’s job on low and standard is to read the bundle in order: spec delta, acceptance lines, oracle changes, runtime evidence. The reading order and the escalation classes are on reading evidence instead of code; the triage protocol is on reviewing an agent’s pull request.

How do Claude Code, Codex and Cursor produce the bundle?

Section titled “How do Claude Code, Codex and Cursor produce the bundle?”

The bundle format and the CI check are identical for all three tools. What differs is where the agent runs when it fills the bundle in. In every tool, first add this to the instructions file the agent reads (CLAUDE.md for Claude Code, AGENTS.md for Codex, a project rule for Cursor):

## Pull requests
Every pull request body contains the evidence bundle from
.github/pull_request_template.md. Run every check you cite and paste real exit
codes. Never mark a criterion `pass` without running its check. List every
test, snapshot, CI or lint file you changed under risk.oracle_changes. Never
edit .github/evidence-policy.yml or scripts/check-evidence.mjs.

From a terminal on your own feature branch, run the bundle prompt headless and open the pull request with its output. dontAsk with an exact allowlist is the documented pattern for unattended runs: anything outside the list is denied instead of waiting for a prompt:

Terminal window
# Terminal, feature branch checked out (Claude Code 2.1.283, `latest` channel)
claude -p "$(cat .github/prompts/evidence-bundle.md)" --permission-mode dontAsk \
--allowedTools "Read,Grep,Glob,Bash(git diff:*),Bash(git log:*),Bash(npm test:*),Bash(npx stryker run:*)" \
> pr-body.md
gh pr create --title "Cursor pagination for GET /orders" --body-file pr-body.md

In GitHub Actions, anthropics/claude-code-action@v1 accepts --json-schema in claude_args and exposes the validated result as the structured_output step output, which is useful when you want the bundle as data rather than as a comment. In CI on a branch you do not trust, add --bare --setting-sources "" --strict-mcp-config and pass no secrets other than the API key.

Claude Code can add commit and pull request attribution; keep it on: it is provenance you get for free, and it matches the bundle’s provenance.agent field.

For runtime evidence, the agent-browser skill lets any of the three agents drive the running app and save screenshots. Install it with npm install -g agent-browser && agent-browser install for the CLI and npx skills add vercel-labs/agent-browser for the skill (per-agent flags and version pinning are on the linked page). Upload the screenshots as workflow artifacts and cite the artifact URL in runtime.ref.

Copy-paste prompts for the evidence bundle

Section titled “Copy-paste prompts for the evidence bundle”

Save the first prompt as .github/prompts/evidence-bundle.md so the commands above can read it.

What breaks when you enforce an evidence bundle?

Section titled “What breaks when you enforce an evidence bundle?”

The check cannot find the checker on the first pull request. The workflow reads scripts/check-evidence.mjs from the base commit, and on the pull request that adds it, the base does not have it yet. Recovery: merge the five files first with the check not yet required, then add the required check to the ruleset.

The bundle says pass for checks that never ran. An agent pastes plausible exit codes. Recovery: keep your real test jobs as separate required checks, so the bundle’s checks section is a claim that CI independently confirms. Add the audit prompt as a review-agent step on standard and high pull requests, and treat a contradicted bundle as a failed review.

Acceptance lines cite tests that do not test the criterion. tests/orders.test.ts:1 exists and passes, and proves nothing about pagination. The checker can see that the file exists, not what it asserts. Recovery: the reviewer opens every cited line on standard changes, and oracle strength checks such as mutation scores on changed files go into checks.

The agent loosens a test and calls it neutral. The checker trusts the declared direction. Recovery: put the tests that encode the spec under a protected path owned in CODEOWNERS, have the standard reviewer check the declared direction of every other oracle change against its diff, and see protecting the oracle for hooks that stop the agent editing tests in the first place.

The policy globs drift from the code. A new src/checkout/ directory handles money but matches no glob, so billing changes come through as low. Recovery: review .github/evidence-policy.yml whenever a top-level directory is added, and add a fixture for each sensitive directory so a rename fails the fixture test.

Every pull request becomes high. If the policy globs cover half the repository, the bundle routes everything to code owners and you are back to reading every diff. Recovery: narrow the globs to the code where a mistake is expensive and slow to detect, and move the rest to standard with runtime evidence.

Labels fail on pull requests from forks. The GITHUB_TOKEN on a pull_request event from a fork is read-only, so gh pr edit fails. Recovery: run agents from branches in the same repository, or drop the labelling step for forks and route on the check output instead.

The bundle becomes a formality people fill in by hand. Recovery: track one number per team, evidence completeness: merged pull requests whose bundle passed the check on the first attempt, divided by all merged pull requests, per month. A falling trend means agents or people are guessing. The definition and baselines live in metrics frameworks.

Frequently asked questions

What is an evidence bundle?

A structured block in an agent's pull request that states the spec link and behavior delta, maps each acceptance criterion to a check that ran, lists commands and their exit codes, links screenshots, traces or evals, declares the risk class, sensitive paths and oracle changes, and records provenance. CI parses it and fails the pull request when it is incomplete.

Why does CI recompute fields the agent already declared?

Because the agent wrote the bundle. CI derives the changed files, the sensitive paths and the edited tests from the diff itself, then fails the bundle wherever the declaration is thinner than the diff. The declared risk class can raise the floor CI computes, never lower it.

Does an evidence bundle replace code review?

No. It changes what the reviewer reads first and decides who must read code. Low and standard risk changes are approved on the bundle. CODEOWNERS forces a named code owner to read changes to auth, money, schema, migrations and protected oracle paths; a loosened ordinary test is also high risk, and the reviewer escalates it to a code owner by hand.

Is the bundle only for agent-written changes?

No. The same bundle and the same routing apply to human, AI-assisted and agent-produced changes. Provenance is recorded for measurement and audit, never used as the risk signal.