How strong is your oracle? Trusting tests you did not read
Oracle strength is how likely a test suite is to fail when the code under test is wrong. It is measured on the code a change touches, with four numbers: mutation score, behaviour coverage of the acceptance criteria, test provenance, and flake rate. A loop runs unattended only after all four clear a bar the team set in advance.
An agent opens a pull request for the invoice export. It adds 38 tests, line coverage on the module is 96%, and CI is green. You flip one >= to > in the date-range filter by hand and run the suite again. It is still green. Nobody read those 38 tests, and now you know they do not check the boundary the ticket was about. This page is for the developer who approves agent pull requests on evidence, and for the tech lead who decides which loops may run overnight.
What you get from measuring oracle strength
Section titled “What you get from measuring oracle strength”- Four metric definitions (mutation score, behaviour coverage, provenance, flake rate) with the exact command that produces each one.
- A bar per lane (interactive, background, overnight) that a loop must clear before it runs with less supervision.
- A CI job that mutation-tests only the files a pull request changed and fails below your threshold.
- Four copy-paste prompts that audit an oracle, kill surviving mutants without touching production code, add negative controls, and hunt flaky tests.
- The tool differences for running the audit in Claude Code, Codex, and Cursor.
Why is a green suite not proof?
Section titled “Why is a green suite not proof?”A test is an oracle only if it can fail. Green tells you that the checks ran and passed; it says nothing about whether they would have failed against a wrong implementation. Agent-written suites make the gap wider for three reasons: the same agent often writes the code and the tests in one pull request, it optimizes for green because green is the stop condition, and it writes many tests quickly, so volume looks like rigour.
Line coverage does not close the gap. This test executes every line of inRange and asserts nothing about the result:
// 100% line coverage, zero oracle strengthit('filters invoices by date range', () => { const result = inRange(invoices, '2026-01-01', '2026-01-31'); expect(result).toBeDefined();});Any implementation that returns an array passes it. Coverage counts lines that ran; oracle strength counts behaviour that was checked. The four measures below are how you tell the two apart without reading every test.
This is the control that the DORA research team names. The 2025 DORA report found “a positive relationship between AI adoption on both software delivery throughput and product performance”, and that “AI adoption does continue to have a negative relationship with software delivery stability”. Its explanation: “Without robust control systems, like strong automated testing, mature version control practices, and fast feedback loops, an increase in change volume leads to instability” (Google Cloud blog, Nathen Harvey and Derek DeBellis, 2025-09-23).
What are the four measures of oracle strength?
Section titled “What are the four measures of oracle strength?”| Measure | Definition | Produced by | What a weak value hides |
|---|---|---|---|
| Mutation score on changed code | Detected mutants (killed or timed out) divided by valid mutants (detected, survived, or not covered), for the files the change touched | Stryker (JS/TS), mutmut (Python), cargo-mutants (Rust), PIT (JVM) | Assertions that never compare the result, missing boundary cases, dead branches |
| Behaviour coverage | Acceptance criteria with a named check that fails when the criterion is violated, divided by all acceptance criteria | The criterion-to-check map in the pull request, plus one recorded negative control per check | Behaviour nobody specified as a check; tests that cover code, not requirements |
| Test provenance | For each decisive check: did it exist before the run, who approved it, and did this change edit it | git log and git diff on the oracle paths | An agent grading its own homework: tests written or loosened in the same change |
| Flake rate | Gating tests that both passed and failed on the same commit during the sweep window, divided by gating tests | Repeated runs with retries off | Checks that go red at random, which trains everyone, human and agent, to rerun instead of read |
Line and branch coverage stay as a floor. Keep a threshold so that untested code shows up, but never read a high number as strength.
Measure mutation score on the code the change touched
Section titled “Measure mutation score on the code the change touched”A mutation tool makes small deliberate bugs (a flipped comparison, a removed call, a constant changed) and runs your tests against each one. A mutant that makes a test fail is killed; one that passes every test survived. Stryker computes the score as detected over valid mutants, where detected is killed plus timed out, and valid adds survived and not-covered mutants; mutants that fail to compile are excluded (checked in mutation-testing-metrics, bundled with @stryker-mutator/core 10.0.0, 2026-09-26).
Score the changed code, not the whole repository. A whole-repository score dilutes a weak new module with years of well-tested code, and a full run is too slow to gate every pull request. Stryker’s --mutate option takes a comma-separated list of files, and even line ranges (src/index.js:1:3-1:5), so you can mutate exactly what the pull request changed:
// stryker.config.json (Stryker 10.0.0){ "testRunner": "vitest", "coverageAnalysis": "perTest", "reporters": ["clear-text", "progress", "json"], "thresholds": { "high": 80, "low": 60, "break": 60 }, "incremental": true}thresholds.break is the gate: when the final score falls below it, Stryker logs the failure and sets exit code 1. With break unset (its default is null), Stryker never fails the build, whatever the score. The high and low values above are Stryker’s own defaults for colouring the report, not a research finding; the break value is your team’s policy. incremental stores results in reports/stryker-incremental.json and reuses them on the next run, and the json reporter writes reports/mutation/mutation.json, which your evidence bundle can read. The runner plugin is a separate package: npm i -D @stryker-mutator/core @stryker-mutator/vitest-runner.
The equivalents in other stacks, checked on 2026-09-26:
- Python: mutmut 3.8.0. Set
source_paths(and optionallyonly_mutate) under[tool.mutmut]inpyproject.toml, runmutmut run, inspect withmutmut resultsandmutmut show <name>, and write CI numbers withmutmut export-cicd-stats(tomutants/mutmut-cicd-stats.json). - Rust: cargo-mutants 27.1.0.
git diff origin/main... > pr.diff && cargo mutants --in-diff pr.difftests only mutants that overlap the diff. Its own documentation warns that a diff which changes only test code runs no mutants, so a pull request that weakens tests needs the provenance check below. - JVM: PIT (
pitest) is the established tool; use its Maven or Gradle plugin.
Not every survivor is a gap. An equivalent mutant changes the code without changing behaviour (for example, a mutated log message nobody asserts on). Triage each survivor as one of three things: a missing test, an equivalent mutant, or behaviour you decide not to pin. Record the decision; an untriaged survivor counts against the loop.
Measure behaviour coverage against the acceptance criteria
Section titled “Measure behaviour coverage against the acceptance criteria”Mutation score tells you whether the tests notice change. Behaviour coverage tells you whether they check the right things. Start from the ticket’s acceptance criteria (see executable acceptance criteria) and require one line per criterion in the pull request: the criterion, the check that proves it, and a negative control, which is the result of breaking that behaviour on purpose and watching the named check go red.
AC-1 reversed range returns 400 -> tests/invoices/export.test.ts "rejects reversed range" negative control: swapped the guard to `from < to`; check failed as expectedAC-2 amounts in invoice currency -> MISSINGA negative control turns “a test exists” into “this test fails when the behaviour breaks”. It is a targeted mutation that you choose, where the mutation tool chooses at random. MISSING is an acceptable answer; it moves the criterion to a human, and it is better than a check the agent invented to fill the row.
Check the provenance of every decisive test
Section titled “Check the provenance of every decisive test”A check proves the most when it existed before the agent started and the agent could not edit it. Two commands answer the provenance questions for a pull request:
# Terminal or CI: which oracle files did this change touch?git diff --name-only origin/main...HEAD -- \ '*.test.*' '*.spec.*' '*__snapshots__*' '.github/workflows/*' \ 'tsconfig*.json' 'vitest.config.*' 'stryker.config.*'
# When was a decisive test first added, and by whom?git log --diff-filter=A --format='%h %an %ad' --date=short -- tests/invoices/export.test.tsClassify each decisive check as pre-existing (added before the run started), human-approved (written as its own slice and approved before implementation began), or same-change (written or edited in this pull request). Same-change checks can support the evidence, but they do not count toward the bar for background or overnight runs. Enforcing the paths the agent may not edit, through hooks, CODEOWNERS and CI-owned checks, is covered in protecting the oracle.
Measure flake rate with retries off
Section titled “Measure flake rate with retries off”A flaky test weakens the oracle in both directions: it fails good changes, and it teaches people and agents that red means “run it again”. Measure it on a schedule, not in the pull request, and with retries disabled so that nothing is hidden:
# Nightly, from the repository rootnpx playwright test --repeat-each 20 --retries 0 --reporter=json > flake-e2e.jsonnpx vitest run --repeats 10 --reporter=json --outputFile=flake-unit.jsonAny test that fails in a repeated run on a commit where it also passes is flaky. Quarantine it out of the gating set for every loop that depends on it, fix or delete it, and restore it only after it passes a sweep. If you keep retries in pull-request CI, run Playwright with --fail-on-flaky-tests so that a test that passes on retry still fails the run (flags checked in @playwright/test 1.63.0 and Vitest 5.0.2, 2026-09-26).
What bar must a loop clear before it runs unattended?
Section titled “What bar must a loop clear before it runs unattended?”A loop is a repeatable class of change with its own trigger, oracle and stop condition, as defined in the one map. The bar belongs to the loop, not to the repository: dependency bumps can clear it while checkout feature work does not. The lanes below match those in shaping a backlog for agents.
| Measure | Interactive | Background | Overnight (unattended) |
|---|---|---|---|
| Mutation score on changed code | Reported | At or above the team’s break threshold | At or above the break threshold, and every survivor triaged |
| Behaviour coverage | Every criterion has a check or is marked MISSING | 100% mapped, each with a negative control | 100% mapped, each with a negative control |
| Provenance | Oracle files touched are read by a human | Decisive checks pre-existing or human-approved; oracle edits escalate | Decisive checks pre-existing; zero oracle files touched, enforced outside the prompt |
| Flake rate | Known | Zero flaky tests in the loop’s gating set | Zero flaky results across the sweep window, retries off |
| Line and branch coverage | Floor | Floor | Floor |
The thresholds are policy you set and write down, not benchmarks. Start with Stryker’s default low band (60) as the background break value and its high band (80) for overnight, then move them on your own data: an escaped defect in a loop is a reason to raise the bar or demote the loop, and a long run of clean pull requests with no survivors is a reason to keep it. Organisations that run agents at volume keep the oracle outside the agent’s reach and bound the run. Stripe’s agents run against “over three million” existing tests and stop after “at most two rounds of CI” (Stripe engineering blog, Alistair Gray, February 2026).
Run an oracle audit on one loop
Section titled “Run an oracle audit on one loop”The audit takes an afternoon for one loop. Run it before you promote a loop, and again after any escaped defect.
-
Name the loop and its oracle paths. Write down the trigger, the stop condition, and the test, snapshot, CI and config globs that make up its oracle. Put the globs in
CODEOWNERSso that any change to them needs a named reviewer. -
Map the acceptance criteria to checks. For the last five pull requests in the loop, list every criterion and its check. Mark gaps
MISSING. Use the negative-control prompt below to record one control per check. -
Run a baseline mutation score. Run the mutation tool on the loop’s modules with
--incrementalso later runs are fast. Triage every survivor as missing test, equivalent, or accepted. -
Close the gaps in a test-only task. Give the agent the survivors and the
MISSINGcriteria, forbid edits to production code, and have a human approve the new tests. They become pre-existing checks for every later run. -
Turn on the per-change gate. Add the CI job below so every pull request in the loop mutation-tests its own changed files and fails below the break threshold.
-
Start the nightly flake sweep. Run the repeat commands on a schedule, quarantine anything flaky, and keep the loop out of the overnight lane until a full window passes clean.
-
Record the result in the evidence bundle. Each pull request carries the mutation score, the criterion map, the list of oracle files touched, and the flake status, in the format in the evidence bundle.
This job mutation-tests only the production files a pull request changed. It relies on thresholds.break in stryker.config.json for the pass or fail decision:
name: oracle-strengthon: pull_requestjobs: mutation: runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 with: fetch-depth: 0 - uses: actions/setup-node@v4 with: node-version-file: .node-version - run: npm ci - name: Mutation-test the files this pull request changed run: | FILES=$(git diff --name-only --diff-filter=AM "origin/${{ github.base_ref }}...HEAD" -- 'src/*.ts' \ | grep -vE '\.(test|spec)\.ts$' | paste -sd, -) if [ -z "$FILES" ]; then echo "No production files changed"; exit 0; fi npx stryker run --mutate "$FILES" - uses: actions/upload-artifact@v4 if: always() with: name: mutation-report path: reports/mutation/In a Git pathspec, * also matches /, so 'src/*.ts' selects TypeScript files at any depth under src/.
Copy-paste prompts for measuring your oracle
Section titled “Copy-paste prompts for measuring your oracle”How do you run the audit in Claude Code, Codex, and Cursor?
Section titled “How do you run the audit in Claude Code, Codex, and Cursor?”The measures, commands and prompts are identical in all three tools, because the mutation and test tools do the measuring. What differs is how you run the audit headless, how you keep the agent from editing the oracle while it works, and where a test-only improvement task runs. Commands were checked against Claude Code 2.1.283 and Codex CLI 0.157.1 on 2026-09-26; Cursor features were checked on cursor.com on 2026-08-28.
Save the first prompt above, “audit the oracle of this pull request”, as prompts/oracle-audit.md; the headless commands below read it from there. Step 2 of that prompt reads the linked issue, so a headless run needs either gh access (the Claude Code allow-list below includes gh issue view) or the acceptance criteria pasted into the prompt file.
Headless audit in CI. -p sessions start in Manual permission mode, so grant only what the audit needs. This allows reads, Git history, Stryker and reading the linked issue, and nothing that edits files:
# CI or terminal, from the repository rootclaude -p "$(cat prompts/oracle-audit.md)" \ --allowedTools "Read,Grep,Glob,Bash(git diff:*),Bash(git log:*),Bash(npx stryker run:*),Bash(gh issue view:*)" \ --max-budget-usd 3 --output-format json > oracle-audit.jsonAdd --json-schema with a schema for the PASS or FAIL table when a later CI step parses the result.
Test-only improvement. In an interactive session, set a completion condition with /goal, for example /goal every survivor in reports/mutation/mutation.json under src/invoices is killed or listed as EQUIVALENT, and no file under src/ is modified. Claude keeps working until the condition is met, a model judges it impossible, or an error clears the goal. Deny edits to src/** in your permission settings rather than trusting the prompt, and check git diff --name-only -- src/ is empty before you accept the result.
Mutation skills. Trail of Bits publishes a mutation-testing plugin (1.9.2) that sets up mewt or muton campaigns and analyzes surviving mutants: claude plugin marketplace add trailofbits/skills, then claude plugin install mutation-testing@trailofbits.
Headless audit in CI. codex exec runs non-interactively. A read-only sandbox cannot run Stryker, which writes a sandbox directory and reports, so use workspace-write and have a later CI step fail the job if anything outside reports/ changed:
# CI or terminal, from the repository root (Codex CLI 0.157.1)codex exec --sandbox workspace-write --output-schema oracle-audit.schema.json \ -o oracle-audit.json "$(cat prompts/oracle-audit.md)"--output-schema constrains the final message to your JSON Schema, and -o writes that message to a file.
Test-only improvement. Use /goal with the same completion condition as in the Claude Code tab. Do not rely on the prompt to keep the agent out of src/: fail the task if git diff --name-only -- src/ is not empty, and see protecting the oracle for enforcement that runs before the edit.
Mutation skills. The same Trail of Bits marketplace installs in Codex: codex plugin marketplace add trailofbits/skills, then codex plugin add mutation-testing@trailofbits.
Audit in the editor. Paste the audit prompt into the agent with the pull request branch checked out. Start in Plan Mode so that it reports before it touches anything.
Test-only improvement. /goal (rolling out since 2026-08-19) gives the agent a long-lived objective; use the same completion condition as in the other tabs, and reject the result if git diff --name-only -- src/ is not empty.
Nightly flake sweep. A Scheduled Automation can start a Cloud Agent that runs the repeat commands and opens an issue listing flaky tests. Keep that agent’s instructions to “report, do not fix”, so that the sweep never edits the tests it measures. Setup is on Cursor Cloud Agents and Automations.
Who signs off on oracle strength?
Section titled “Who signs off on oracle strength?”Nobody reads every test, so the measurement itself must be checkable:
- CI computes the numbers. The mutation score, the list of oracle files touched, and the flake status come from tool output attached to the pull request, never from the agent’s summary of it.
- A human approves every new decisive check before it counts as pre-existing. That approval is the one place where a person reads test code on purpose.
- The tech lead owns the thresholds and the oracle globs, reviews the survivor triage for promoted loops, and demotes a loop when a defect escapes it. The promotion rules are in helping the team stop reading every diff.
- Every escaped defect becomes a check. Write the test that would have caught it, confirm it with a negative control, and add it to the oracle before the loop runs unattended again.
What breaks when you measure oracle strength?
Section titled “What breaks when you measure oracle strength?”The agent kills mutants with change-detector tests. It asserts on call counts, private helpers or exact log strings, the score rises, and the tests break on every refactor while catching nothing new. Recovery: require assertions on observable behaviour (the second prompt does), and reject any new test whose only failure mode is a harmless refactor.
The score is gamed through scope. The mutate list shrinks, files move into an ignore pattern, or thresholds.break drops to make CI pass. Recovery: treat the Stryker config as an oracle file with a code owner, and compute the mutated file list in CI from the diff, as the workflow above does, not from a list the pull request can edit.
Mutation runs are too slow to gate on. A whole-repository run on every pull request takes too long, so the team turns it off. Recovery: mutate only changed files per pull request, use --incremental, and run the full score weekly on a schedule.
Timeouts count as kills. Stryker counts a timed-out mutant as detected, which is correct for an infinite loop but can hide a slow, flaky test. Recovery: if many mutants time out in one file, look for tests with long waits before trusting the score.
Retries hide the flake rate. A retry turns a flaky failure into a pass and the rate reads zero. Recovery: measure flakiness only in the nightly sweep with retries off, and use --fail-on-flaky-tests wherever retries stay on.
Same-change tests are counted as proof. The agent writes the code and the tests that bless it, and every number looks healthy. Recovery: the provenance classification decides what counts; same-change checks never clear the background or overnight bar on their own.
Where to go next with oracle strength
Section titled “Where to go next with oracle strength”Frequently asked questions
What is oracle strength?
Oracle strength is how likely a test suite is to fail when the code is wrong. It is measured with four numbers on the code a change touches: mutation score, behaviour coverage of the acceptance criteria, test provenance, and flake rate. Line coverage is a floor, not a measure of strength.
Why is line coverage not enough for agent-written tests?
Line coverage counts code that ran, not code that was checked. A test with no assertion executes every line and passes against any implementation. Mutation testing changes the code on purpose and counts how many of those changes the tests notice, which is the property you actually rely on.
What bar should a loop clear before it runs unattended?
Every acceptance criterion maps to a check with a recorded negative control, the decisive checks existed before the run and are outside the paths the agent may edit, the mutation score on changed code clears the team's break threshold with every survivor triaged, and the gating tests had no flaky result in the sweep window with retries off.
Can the agent improve the mutation score itself?
Yes, in a separate, test-only task. The agent writes tests to kill surviving mutants, may not edit production code, and a human approves the new tests before they count toward the oracle. Tests written in the same pull request as the code they judge do not count.