Reading evidence instead of code
Reading evidence instead of code means approving an agent’s change on four artifacts rather than on its diff: the spec delta, the acceptance results, oracle strength, and runtime signals from a running build. A named human still reads code for authentication, money, schema and migrations, changes to the tests themselves, and anything no check covers.
The agents opened 14 pull requests overnight. Each one is green, each one is 300 lines, and you have a planning meeting at ten. Reading all of them properly takes the whole day. Skimming them takes an hour and proves nothing, because a skimmed diff is an approval with a worse audit trail. This page gives that hour a better use.
What evidence-based review gives you
Section titled “What evidence-based review gives you”- A reading order for any agent pull request: spec delta, acceptance results, oracle strength, runtime signals, with the question each one answers.
- An escalation table of the change classes where someone still reads the code, with the reason each class is on it and file globs to start from.
- Three copy-paste prompts that make the agent produce the evidence, audit its own oracle, and classify the change.
- A personal trust log that moves one change class at a time from “I read everything” to “I read the evidence”.
- The warning signs that you stopped reading too early, and how to recover from each one.
- For tech leads, CTOs and executives: three metric definitions that separate an evidence review from no review at all.
Why reading every diff stops working
Section titled “Why reading every diff stops working”The load is measured. Faros AI’s Acceleration Whiplash report (April 2026; two years of telemetry across 22,000 developers and 4,000 teams) found median time in review up 441.5% and 31.3% more pull requests merging with no review at all. In the same data, epics completed per developer rose 66.2% while incidents per pull request rose 242.7%. Generation scaled; reading did not.
Developers already distrust what they are not checking. In Sonar’s developer survey (January 2026, more than 1,100 respondents), “96% of developers do not fully trust AI-generated code, and only 48% always verify it before committing.”
So the choice is not between reading every line and reading nothing. Level 3 of the ladder ends when review bandwidth becomes the ceiling. The way past it is to change what you read: output from checks that ran, instead of text that looks plausible.
What do you read instead of the diff?
Section titled “What do you read instead of the diff?”Read four artifacts, in this order. Each one answers a question the diff answers poorly, and each one can fail in a way you can see.
| # | Artifact | The question it answers | Where it comes from | What a failing one looks like |
|---|---|---|---|---|
| 1 | Spec delta | Did the agent change the behaviour we asked for, and only that? | The spec or ticket, plus a short, plain-language list of behaviour changes | Behaviour changes nobody asked for; “refactored X while I was there” |
| 2 | Acceptance results | Does each acceptance criterion have a check that ran and passed? | Test, e2e and eval output, mapped one criterion to one check | A criterion with no check, or a check that was skipped |
| 3 | Oracle strength | Could those checks have failed, and did the agent touch them? | The test diff, mutation results on changed code, and whether the decisive check existed before the change | Tests edited in the same pull request, assertions loosened, snapshots regenerated |
| 4 | Runtime signals | Does the running software behave as specified? | Preview or staging runs, browser screenshots, traces, then canary metrics and error rates after deploy | No run at all, or a screenshot of the happy path only |
Start with the spec delta
Section titled “Start with the spec delta”The spec delta is the shortest artifact and the one most worth your time. It states, in sentences, which behaviour changed: “GET /orders now returns 50 items per page and a next_cursor; the old page parameter returns 400.” You compare it with what you asked for. Scope creep, a misread requirement, and a silent contract change all show up here in seconds, and none of them needs code to spot. If nobody wrote down what was asked, write that first; see writing acceptance criteria an agent cannot misread.
Then map acceptance criteria to checks
Section titled “Then map acceptance criteria to checks”Every criterion in the spec gets exactly one line: the criterion, the check that proves it, and the result. “Old page parameter returns 400 → orders.contract.test.ts:88 → pass.” A criterion without a line is unverified, however green the pipeline looks. An agent can produce this mapping cheaply because it knows both the criteria and the tests it wrote; that is also why the next check, oracle strength, matters.
Then ask whether the oracle could have failed
Section titled “Then ask whether the oracle could have failed”The oracle is the set of checks that decides “done”. Because the agent may have written both the feature and its tests, green means little until you know two things. First, did the pull request change any test, fixture, snapshot, CI step, or lint rule? If so, read that part of the diff; it is the only part that can make every other check lie. Second, would the tests fail if the code were wrong? Mutation testing answers that mechanically: it breaks the changed code on purpose and counts how many mutants the tests catch. Scope it to the changed files, because a whole-project run is slow in CI. For JavaScript and TypeScript (@stryker-mutator/core 10.0.0 on npm), pass the changed source files to --mutate: npx stryker run --mutate "$(git diff --name-only --diff-filter=d main -- 'src/*.ts' ':!*.test.ts' | paste -sd, -)", and skip the step when the list is empty. For Python (mutmut 3.8.0 on PyPI), mutmut run takes no path arguments, so list the changed files under only_mutate in [tool.mutmut] in pyproject.toml (or [mutmut] in setup.cfg) before you run it; source_paths stays your package root (it replaced paths_to_mutate in mutmut 3.6.0). The full method is on how strong your oracle is, and keeping the agent’s hands off it is on protecting the oracle.
Stripe shows why a pre-existing oracle matters. Its Minions agents produce “over 1,300 Stripe pull requests” a week that are “human-reviewed, but containing no human-written code” (Stripe engineering blog, 19 February 2026). They run against an existing suite of “over three million” tests, bounded to “at most two rounds of CI” (Part 1 of the same series, 9 February 2026). The oracle predates the agent, and that is why the review can be short.
Finish with runtime signals
Section titled “Finish with runtime signals”Tests prove what someone thought to test. A running build shows what the change does. Before merge, that means a preview or staging run with a screenshot or trace of each acceptance path. The agent-browser skill lets the agent drive the running app and capture that evidence itself. After merge, it means canary metrics, error rates, and an automatic rollback. Progressive delivery is the safety net that makes pre-merge reading optional for low-risk classes.
The four artifacts together form the evidence bundle, the pull request contract that CI enforces. This page covers the reader’s side; the bundle page covers the template and the CI check that fails an incomplete one.
Which changes still need a human to read the code?
Section titled “Which changes still need a human to read the code?”Evidence works when the oracle is strong, feedback is fast, and a mistake is cheap to undo. The classes below fail at least one of those conditions, so a named person reads the code every time, in addition to the evidence. Keep this list short; a list that covers half the repository puts you back at Level 3.
| Change class | Why evidence is not enough | Globs to start from (adjust to your repository) |
|---|---|---|
| Authentication and authorization | Tests prove the allowed paths; the failure is a path nobody tested. The damage is a breach, not a bug. | **/auth/**, **/middleware/**, **/*permission*, **/*policy* |
| Money | Rounding, currency, idempotency and refund edge cases surface weeks later, in customer accounts. | **/billing/**, **/payments/**, **/pricing/**, webhook handlers |
| Schema | A contract other services and old clients depend on; the break shows up outside this repository’s tests. | **/schema.*, **/*.sql, **/openapi*, **/*.proto |
| Data migrations | Often irreversible, run once against production data the test fixtures do not resemble. | **/migrations/** |
| The oracle itself | A change to tests, CI, lint or type config changes what “green” means for every other change. | **/*.test.*, **/__snapshots__/**, .github/workflows/**, lint and tsconfig files |
| Anything no check covers | With no oracle, the only evidence is the code. | Decided per pull request: an acceptance line with no check |
The first four classes match the ones the human’s job keeps lit on every loop. The last two are conditions, not directories: they apply wherever they occur. Route all six through CODEOWNERS so the required reader is enforced by the forge, not remembered by a person. The autonomy and risk-class policy is where an organization writes this list down once.
How do you get the evidence from Claude Code, Codex and Cursor?
Section titled “How do you get the evidence from Claude Code, Codex and Cursor?”The prompts in the next section work in all three tools. What differs is where the evidence is produced and who can block the merge.
Run the evidence prompt headless in CI or from your terminal. -p sessions start in Manual permission mode, so any Bash command outside the allowlist is denied. The prompt asks the agent to run the acceptance checks, so the allowlist has to include your test commands. Below, git diff, git log and the test runners are the only shell commands the run may execute; replace them with your repository’s own:
# Terminal or CI, from the repository root (Claude Code 2.1.283)claude -p "$(cat .github/prompts/evidence.md)" \ --allowedTools "Read,Grep,Glob,Bash(git diff:*),Bash(git log:*),Bash(npm test:*),Bash(npx vitest:*),Bash(npx stryker run:*)" \ --output-format json > evidence.jsonThe run executes the pull request’s own test commands with your API key in the environment. In CI, run it on pull_request from branches in this repository, never on pull_request_target with the contributor’s checkout.
For a second opinion on the code itself, /code-review reviews the local diff inside a session, and claude ultrareview runs the cloud-hosted multi-agent review from the shell. The managed Code Review (research preview, Team and Enterprise plans) check run “always completes with a neutral conclusion so it never blocks merging”: treat it as a reviewer’s comments, not a gate.
Run the evidence prompt with plain codex exec, which takes it as the task and writes the final message to a file. The checks write caches and Stryker’s sandbox directory, so give the run a writable workspace with the :workspace permission profile (beta); on a setup without profiles, --sandbox workspace-write does the same. Use one or the other, not both:
# Terminal or CI, from the repository root (Codex CLI 0.157.1)codex exec -c default_permissions=":workspace" "$(cat .github/prompts/evidence.md)" -o evidence.md
# Separate code review of the branch; review rules come from AGENTS.mdcodex exec review --base main -o review.mdcodex exec review --base main reviews the branch against main. A custom review prompt works only without --base, --commit or --uncommitted: combining them fails with error: the argument '--base <BRANCH>' cannot be used with '[PROMPT]' (checked on 0.157.1). That is why the evidence prompt runs through codex exec and the review rules live in AGENTS.md.
On GitHub, @codex review on a pull request runs the same kind of review, and it also reads its custom review rules from AGENTS.md. Put the escalation table’s globs there so the review names the class a change falls into.
Paste the evidence prompt into the agent in the editor before you open the pull request, and use its output as the pull request description.
On the pull request, Bugbot “reviews pull requests and identifies bugs, security issues, and code quality problems”. PR Routing & Approval “assigns reviewers based on code ownership and commit history, and can approve low-risk PRs when your criteria are met” (both checked on cursor.com on 2026-08-28). Encode the escalation table as the criteria: approval is allowed only when no escalation class is touched. See PR Routing & Approval in Cursor.
Runtime evidence is identical across the three tools. The agent-browser CLI is installed once and shared; the skill that teaches an agent to use it is installed per agent, so name each agent you use. The -a and -y flags keep skills add non-interactive:
# Terminal. agent-browser 0.38.1 on npm, checked 2026-09-26npm install -g agent-browser && agent-browser installnpx skills add vercel-labs/agent-browser -a claude-code -a codex -a cursor -yCopy-paste prompts for evidence-based review
Section titled “Copy-paste prompts for evidence-based review”Keep a personal trust log
Section titled “Keep a personal trust log”Trust in evidence is earned per change class, not per tool or per agent. A trust log is a short file where you record, for each class, what you read and whether reading the code found anything the evidence missed. Keep it in your notes or in the repository under docs/ if the team shares it.
# Trust log: Anna, orders-service
| Date | PR | Class | Stage | Evidence said | Code read found | Escaped to prod? || ---------- | ---- | --------------- | ----------- | ------------- | -------------------------- | ---------------- || 2026-09-22 | #412 | API pagination | read-all | pass | nothing extra | no || 2026-09-23 | #418 | API pagination | read-all | pass | unasked rename of a field | no || 2026-09-24 | #421 | UI copy | sampled | pass | not read (sample skipped) | no || 2026-09-25 | #425 | billing webhook | always-read | pass | missing idempotency check | no |-
Start every class at
read-all. You read the evidence first, then the code, and write down whether the code showed you anything the evidence did not. -
Promote a class to
sampledafter a clean run. Our starting rule is 20 consecutive pull requests in which the code read found nothing that mattered and the spec delta caught what did. The number is a team policy, not a research finding; lower it for trivial classes and raise it for anything user-facing. Insampled, you read the code of one pull request in five, chosen before you look at the evidence. -
Promote to
evidence-onlyafter a second clean run: another 20 consecutive pull requests atsampledin which the sampled reads found nothing that mattered and no defect escaped to production in the class. You read the four artifacts and approve. -
Demote one stage on the first miss. A defect that reached production, or a code read that found a real problem the evidence missed, sends the class back one stage. Fix the evidence gap first (a missing acceptance check, a weak test), then earn the stage again.
-
Never promote the escalation classes. Auth, money, schema, migrations and oracle changes stay at
always-read. What improves there is the evidence you read alongside the code, not whether you read the code.
The log is also your answer when someone asks why you approved a change without reading it. “This class has 40 clean pull requests and no escapes in my log” is an accountable statement. “It looked fine” is not. For rolling the same protocol out across a team, see helping the team stop reading every diff.
What should tech leads, CTOs and executives measure?
Section titled “What should tech leads, CTOs and executives measure?”The failure the Faros data describes is not “reviews got shorter”. It is 31.3% more pull requests merging with no review. The fix is not to demand more reading, which the same data shows does not scale. The fix is to make the lighter review a real one, and to measure that it is. Three definitions to adopt as-is:
| Metric | Definition | What a bad trend means |
|---|---|---|
| Evidence completeness | Share of merged agent pull requests with all four artifacts present (spec delta, acceptance mapping, oracle-change statement, runtime evidence) | Approvals are running on trust, not evidence |
| Escape rate by stage | Production defects traced to pull requests approved at evidence-only, divided by all pull requests approved at that stage, per change class, per month | A class was promoted too early; demote it |
| Escalation coverage | Share of merged pull requests touching an escalation class that record a named human code reader | The escalation list is being bypassed |
Report them per change class, never as one blended number. The canonical definitions and baselines live in metrics frameworks for agentic engineering. For executives, the question to ask your CTO is short: “Which classes of change merge without a human reading the code, who decided that, and what is the escape rate in those classes?”
Signs you stopped reading too early
Section titled “Signs you stopped reading too early”Tests change in the same pull request as the code they judge, and nobody notices. This is the most common way an agent’s green run becomes meaningless. Recovery: make oracle changes an escalation class in CODEOWNERS, and add a CI step that lists every modified test file in the pull request summary.
The evidence summary paraphrases the diff instead of reporting results. “Updated the handler to validate input” is a description of code, not evidence. Recovery: reject any acceptance line without a file:line or a command and its exit status, and require the agent to run the checks it cites.
A defect escapes in a class you approve on evidence alone. Recovery: demote the class one stage in your trust log, write the check that would have caught the defect, and add it to the oracle before you promote again. Tag the incident with the failed control, as described in the failure taxonomy.
You can no longer explain a module you approved last month. Anthropic’s internal study (Saffron Huang and colleagues, 2 December 2025; 132 engineers and researchers surveyed, 53 interviewed) calls this the “paradox of supervision”: overseeing agents needs exactly the skills that over-delegation erodes. Recovery: read one sampled pull request a week in full, in a class you own, and ask the agent to explain the design before you read it. Your craft and career covers which skills to keep deliberately.
The escalation list quietly grows or quietly shrinks. A list that grows to cover half the repository sends you back to Level 3; a list that shrinks without a written decision is how auth code merges unread. Recovery: change the list only in a reviewed pull request to CODEOWNERS, with a reason in the commit message.
Runtime evidence is always the happy path. One screenshot of a successful checkout proves one path. Recovery: require runtime evidence for each acceptance criterion, including the error cases the spec names.
Where to go next with evidence-based review
Section titled “Where to go next with evidence-based review”Frequently asked questions
What do you read instead of an agent's diff?
Four things, in order: the spec delta (what behaviour the change claims to alter), the acceptance results (each criterion mapped to a check that ran), the oracle's strength (whether the tests could have failed and whether the agent touched them), and runtime signals from a preview, staging or canary.
Which changes still need a human to read the code?
Authentication and authorization, money, schema, and data migrations, plus any change that edits its own oracle (tests, CI, lint or type config) and any change no check covers. In those classes the oracle is weak, feedback is slow, or the damage is hard to undo, so a named person reads the code.
How do I know I stopped reading code too early?
Defects reach production in a class you approve on evidence alone, tests change in the same pull request as the code they judge, you can no longer explain a module you approved last month, or the evidence summary is a paraphrase of the diff instead of results from checks that ran. Each sign sends that class back one trust stage.
Is approving on evidence the same as merging without review?
No. A merge with no review has nobody accountable and no record. An evidence review has a named approver, a written spec delta, checks that ran, and a risk class that decided whether code reading was required.