Skip to content

Protecting the oracle: checks the agent cannot edit

Protecting the oracle means making the checks that decide “done” (tests, snapshots, test config and CI workflows) impossible for a coding agent to change in the task they judge. The protection is layered: session deny rules and sandboxing, CODEOWNERS, a required CI check outside the diff, holdout scenarios outside the repository, and a test-weakening audit.

You asked the agent to fix a failing checkout test. Twenty minutes later the suite is green and the summary says “fixed the rounding bug”, and the diff is 40 lines. Four of them are in the test: toBe(19.99) became toBeCloseTo(20, 0). Nobody lied on purpose. The agent was asked to make the test pass, and editing the test was the shortest path. This page makes that path unavailable, and makes it loud when a test change is genuinely needed.

What you’ll walk away with from locking the oracle

Section titled “What you’ll walk away with from locking the oracle”
  • A map of your oracle, as globs you reuse in every layer.
  • Session locks you can paste: Claude Code deny rules, sandbox and hook; a Codex permission profile tested on Codex CLI 0.157.1; and Cursor’s limits.
  • Forge and CI locks: CODEOWNERS, a required workflow the pull request cannot edit, and a job that runs the base branch’s tests against the new code.
  • A holdout job that runs scenarios the agent has never seen and reports only IDs and counts.
  • A weakening audit script, tested against real weakened changes.
  • Four copy-paste prompts and a red-team drill.

Why an agent edits the test instead of the code

Section titled “Why an agent edits the test instead of the code”

“Make the tests pass” is satisfied equally by fixing the code or by changing the test, and the test is often the smaller edit. “Never modify tests” in CLAUDE.md or AGENTS.md is text the model weighs against the task, not a rule anything enforces. As the autonomy ledger at Level 5 puts it: “If the agent can edit the oracle, the oracle is a suggestion.” This page assumes a feedback loop already exists (the test stage) and leaves its strength to how strong your oracle is; it only keeps that loop out of the agent’s reach.

The oracle is everything whose change can turn a red run green without the code getting better. It is wider than the tests/ folder:

Oracle componentExamplesHow it gets weakened
Test codetests/**, **/*.test.ts, **/*_test.go, test_*.pyLoosened assertion, deleted case, .skip, xfail
Recorded expectations**/__snapshots__/**, golden files, fixturesSnapshot regenerated with -u, fixture edited to match the bug
Test and coverage configvitest.config.ts, jest.config.*, pytest.ini, .coveragercFiles excluded from the run, coverage threshold lowered
Static gatestsconfig.json, lint config, // @ts-ignore, # noqastrict turned off, rule disabled, suppression comment added
The pipeline.github/workflows/**Test step removed or marked continue-on-error
The locks themselvesCODEOWNERS, .claude/, .codex/The protection rule deleted along with the change

Suppression comments inside production files cannot be path-locked; the weakening audit later on this page catches them.

No single control covers every agent and every bypass, so you stack five.

LayerWhere it livesWhat it stopsWhat gets past itWhich agents it binds
1. Session lockAgent settings on the machine that runs the agentThe edit, at the moment the agent tries it, with a message it can act onAnything the settings do not cover; cloud agents that run with other settingsOne tool, one machine
2. Forge lockCODEOWNERS plus the “Require review from Code Owners” ruleMerging an oracle change without a named owner’s approvalNothing, if the rule is on and the file is validEvery agent and every human
3. CI outside the diffA required workflow stored in another repositoryA pull request that rewrites its own pipelineWeak tests: CI runs what existsEvery pull request
4. HoldoutsA separate repository the agent cannot readCode tuned to the visible testsBehaviour no scenario describesEvery pull request
5. Weakening auditA CI job over the diffLegitimate-looking test edits that loosen the checkSubtle semantic weakening; that is what the code owner reviewsEvery pull request

Layer 1 gives the fastest feedback. Layers 2 to 5 hold when the agent runs in the cloud or under a tool you have not configured.

  1. Map the oracle. Run the first prompt below and review the list of globs it returns once.

  2. Lock it in the session. Apply and commit the settings for your tool from the tabs in the next section.

  3. Give the oracle owners. Add the CODEOWNERS block and turn on “Require review from Code Owners” for the default branch.

  4. Move the decisive check out of the diff. Put the test, base-oracle and audit jobs in a workflow the pull request cannot edit, and require it through an organization ruleset.

  5. Add holdouts. Create a private scenario repository and the holdout job. Start with 10 to 20 scenarios for the flows that make you money.

  6. Split test changes from code changes. A test-authoring session writes new tests, a code owner approves them, and a separate session makes them pass with the tests locked.

  7. Red-team the lock with the drill at the end of this page, after every agent or CI upgrade.

How do you lock test files in Claude Code, Codex and Cursor?

Section titled “How do you lock test files in Claude Code, Codex and Cursor?”

The three tools enforce the session lock differently, and one of them cannot be fully verified today. Use the tab for each tool your team runs.

Use three mechanisms in .claude/settings.json: deny rules for the built-in file tools, the sandbox for everything a shell command writes, and a PreToolUse hook that tells the agent what to do instead. Checked against Claude Code 2.1.283 and its permissions, sandboxing and hooks documentation on 26 September 2026.

{
"permissions": {
"deny": [
"Edit(tests/**)",
"Edit(**/*.test.ts)",
"Edit(**/__snapshots__/**)",
"Edit(/vitest.config.ts)",
"Edit(/.github/**)",
"Edit(/CODEOWNERS)"
]
},
"sandbox": {
"enabled": true,
"allowUnsandboxedCommands": false,
"filesystem": {
"denyWrite": ["./tests", "./.github", "./CODEOWNERS", "./vitest.config.ts"]
}
},
"hooks": {
"PreToolUse": [
{
"matcher": "Edit|Write|NotebookEdit",
"hooks": [
{ "type": "command", "command": "${CLAUDE_PROJECT_DIR}/.claude/hooks/protect-oracle.sh" }
]
}
]
}
}

Why all three:

  • Deny rules block in every permission mode, including bypassPermissions, and an allow rule cannot carve an exception out of them. Edit(...) rules cover Write too; a path rule written as Write(...) is accepted but never consulted. As deny rules, Edit(tests/**) matches a tests directory at any depth, while the leading / in Edit(/vitest.config.ts) anchors to the project root. Edit deny rules also apply to shell commands Claude Code recognises, such as sed, tee and > redirects.
  • The sandbox closes the gap the deny rules leave open: the documentation says they “don’t apply to … arbitrary subprocesses that read or write files indirectly, like a Python or Node script that opens files itself”. sandbox.filesystem.denyWrite is enforced by the operating system on every shell command and its children. Note the path syntax: ./tests here, /tests in a permission rule. The sandbox runs on macOS, Linux and WSL2, not native Windows. allowUnsandboxedCommands: false stops a blocked command from being retried outside the sandbox.
  • The hook turns a bare denial into guidance. The agent reads the stderr of an exit-2 block:
#!/usr/bin/env bash
# .claude/hooks/protect-oracle.sh: refuse edits to the files that decide "done"
command -v jq >/dev/null || { echo "protect-oracle: jq missing" >&2; exit 2; }
file=$(jq -r '.tool_input.file_path // .tool_input.notebook_path // empty')
rel=${file#"$CLAUDE_PROJECT_DIR"/}
case "$rel" in
tests/*|*/tests/*|*.test.ts|*.spec.ts|*/__snapshots__/*|.github/*|CODEOWNERS|vitest.config.*|.claude/*)
echo "BLOCKED: $rel is part of the oracle for this task. Change the code under test instead." \
"If you believe the test itself is wrong, stop and write your reasoning to TEST_DISPUTE.md." >&2
exit 2 ;;
esac
exit 0

The script needs jq on PATH; without it, the guard line blocks every edit with “jq missing” rather than silently allowing them all (the deny rules still block either way). The */tests/* pattern mirrors the deny rule’s any-depth match, so packages/api/tests/x.ts in a monorepo gets the same guidance message. Make it executable (chmod +x .claude/hooks/protect-oracle.sh). Exit code 2 is the code that blocks; exit 1 is a non-blocking error and the edit goes ahead. The .claude/ directory is also a protected path: in Manual, acceptEdits and auto mode, a write to it prompts you or goes to the classifier instead of being auto-approved, so the agent cannot quietly delete its own lock.

The forge lock is the one layer that binds every agent surface: a local CLI, a cloud agent, a GitHub-triggered agent and a human. It makes an oracle change mergeable only with a named team’s approval.

# .github/CODEOWNERS: changes to the oracle need a named owner's approval
/tests/ @acme/test-owners
*.test.ts @acme/test-owners
**/__snapshots__/ @acme/test-owners
/vitest.config.ts @acme/test-owners
/.github/ @acme/platform
/.claude/ @acme/platform
/.codex/ @acme/platform

Then enable Require review from Code Owners in the branch protection rule or ruleset for your default branch. Four behaviours from GitHub’s documentation decide whether this works:

  • The base branch’s CODEOWNERS applies. A pull request that edits CODEOWNERS cannot change who reviews that same pull request, but it can weaken the file for every later one. The /.github/ line makes CODEOWNERS changes need platform approval too.
  • Invalid lines are skipped silently. A typo leaves a path unowned. Check the file’s error view on GitHub, or the “list CODEOWNERS errors” REST endpoint, after every edit.
  • Any one owner is enough. If a path has several owners, one approval satisfies the rule. Keep oracle owners to a team that actually reviews tests.
  • Owners need write access to the repository to be requested.

If CODEOWNERS requests nobody, one of those is the cause: a skipped line, an owner without write access, or a target branch whose CODEOWNERS differs. Prove the setup with a throwaway pull request that touches tests/. GitHub recommends the same pattern, CODEOWNERS plus the code-owner review rule, for the configuration files its Copilot cloud agent could change.

Run the decisive check where the pull request cannot edit it

Section titled “Run the decisive check where the pull request cannot edit it”

A pull_request workflow runs from the pull request’s merge commit. So the workflow YAML that runs is the YAML the pull request contains, and an agent that removes the test step from .github/workflows/ci.yml still gets a green check named test. CODEOWNERS on .github/ makes that edit visible. A required workflow removes the possibility.

In an organization ruleset, the rule Require workflows to pass before merging enforces a workflow file stored in a different repository. The pull request cannot edit it, because it is not in the diff. Put the oracle jobs there:

# acme/ci-policy/.github/workflows/oracle.yml, required by an org ruleset
name: oracle
on: pull_request
permissions:
contents: read
pull-requests: read # lets oracle-approved.sh list the pull request's reviews
jobs:
base-oracle:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
with: { fetch-depth: 0, persist-credentials: false }
- uses: actions/checkout@v7
with:
repository: acme/ci-policy
path: .policy
token: ${{ secrets.POLICY_READ_TOKEN }}
persist-credentials: false
- id: owner
name: Check for an oracle owner's approval of this exact commit, before any pull request code runs
env:
GH_TOKEN: ${{ github.token }}
REPO: ${{ github.repository }}
PR: ${{ github.event.pull_request.number }}
HEAD_SHA: ${{ github.event.pull_request.head.sha }}
run: bash .policy/oracle-approved.sh >> "$GITHUB_OUTPUT"
- name: Run the base branch's tests and runner config against this pull request's code
env:
OWNER_APPROVED: ${{ steps.owner.outputs.approved }}
run: |
git checkout ${{ github.event.pull_request.base.sha }} -- tests/ vitest.config.ts
npm ci
npx vitest run tests/ || {
[ "$OWNER_APPROVED" = true ] || exit 1
echo "::warning::Old expectations fail; accepted because an oracle owner approved this commit."
}
oracle-audit:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
with: { fetch-depth: 0, persist-credentials: false }
- uses: actions/checkout@v7
with:
repository: acme/ci-policy
path: .policy
token: ${{ secrets.POLICY_READ_TOKEN }}
persist-credentials: false
- id: owner
env:
GH_TOKEN: ${{ github.token }}
REPO: ${{ github.repository }}
PR: ${{ github.event.pull_request.number }}
HEAD_SHA: ${{ github.event.pull_request.head.sha }}
run: bash .policy/oracle-approved.sh >> "$GITHUB_OUTPUT"
- env:
OWNER_APPROVED: ${{ steps.owner.outputs.approved }}
run: |
bash .policy/oracle-audit.sh ${{ github.event.pull_request.base.sha }} || {
[ "$OWNER_APPROVED" = true ] || exit 1
echo "::warning::Weakening signal accepted because an oracle owner approved this commit."
}

Both jobs run code the pull request controls (npm ci lifecycle scripts, the test run), so the top-level permissions block limits their GITHUB_TOKEN to reading the repository and its pull requests. The base-oracle job restores every test file the pull request modified or deleted, and the runner config, to their base-branch versions, keeps the tests it added, and runs the lot against the new code. A failure lists exactly which old expectations the change breaks. Each must appear in the pull request’s spec delta as an intended behaviour change; one with no matching line is a regression or a weakened test. Adjust tests/ and the runner command to your layout.

GitHub’s ruleset documentation adds two constraints. The rule blocks direct pushes, so apply it only to branches that change through pull requests. A private workflow repository can serve only private repositories and an internal one only internal and private ones; a public one serves all. If the policy repository is private or internal, also allow access to its workflows from other repositories under Settings → Actions → General → Access in that repository, or the required workflow never runs.

A legitimate behaviour change breaks old expectations by design, so both jobs need a pass path, or every intended test change would be blocked forever. The pass path is an approval, not a label: anyone with triage access can add a label, including an agent’s token. oracle-approved.sh lives in the policy repository next to the audit script, and runs before any pull request code:

#!/usr/bin/env bash
# oracle-approved.sh: print approved=true only if a listed oracle owner's latest review
# approves this exact head commit. Any API or parse error fails the step (fail closed).
set -euo pipefail
: "${REPO:?}" "${PR:?}" "${HEAD_SHA:?}"
owners=$(grep -vE '^\s*(#|$)' .policy/oracle-approvers.txt | jq -R . | jq -sc .)
reviews=$(gh api --paginate "repos/$REPO/pulls/$PR/reviews" | jq -s 'add // []')
ok=$(jq -r --argjson owners "$owners" --arg sha "$HEAD_SHA" '
[ .[] | select(.state != "COMMENTED" and .state != "PENDING") ]
| group_by(.user.login) | map(max_by(.submitted_at))
| any(.[]; .state == "APPROVED" and .commit_id == $sha
and (.user.login as $u | $owners | index($u)))' <<<"$reviews")
if [ "$ok" = true ]; then echo "approved=true"; else echo "approved=false"; fi

oracle-approvers.txt holds one GitHub login per line, the people in the oracle team from CODEOWNERS. It is a file in the policy repository rather than a team lookup because the workflow’s GITHUB_TOKEN cannot read organization team membership. Keep the two lists in step.

The flow for an intended change: the jobs fail, a listed owner reviews the spec delta and approves, then clicks Re-run failed jobs. A re-run uses the same event and commit but queries the reviews again, so the check now passes with a warning annotation. Three properties keep it closed. The approval must be on the exact head commit, so a push after approval re-arms the check. An owner’s later “request changes” or a dismissal cancels the approval, while a plain comment does not. An API error, a missing approvers file or an empty output fails the job instead of passing it. We ran the script against mocked review lists: an owner’s approval of the head commit gave approved=true; a non-owner’s approval, an approval of an older commit, an approval followed by “request changes”, a dismissed approval and no reviews all gave approved=false; an API error exited non-zero. Keep “Require review from Code Owners” on as well: the script decides whether the check passes, and the branch rule is what stops a merge without an owner.

Keep holdout scenarios out of the repository

Section titled “Keep holdout scenarios out of the repository”

Even a locked suite teaches the agent what “done” looks like, because the agent can read the tests. Over many iterations, code drifts towards passing exactly those cases. A holdout set is machine learning’s answer to the same problem: scenarios the agent has never seen, run only in CI, reported only as pass or fail.

# Two jobs in the required oracle.yml (same read-only top-level permissions)
build:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
with: { persist-credentials: false }
- run: npm ci && npm run build
- uses: actions/upload-artifact@v7
with: { name: app-build, path: dist/ }
holdout:
needs: build
runs-on: ubuntu-latest
steps:
- uses: actions/download-artifact@v8
with: { name: app-build, path: dist/ }
- uses: actions/checkout@v7
with:
repository: acme/checkout-holdouts
path: .holdout
token: ${{ secrets.HOLDOUT_READ_TOKEN }}
persist-credentials: false
- name: Run holdouts and print IDs and counts only
run: |
npm ci --prefix .holdout
.holdout/node_modules/.bin/playwright install --with-deps chromium
.holdout/node_modules/.bin/playwright test --config .holdout/playwright.config.ts --reporter=json > "$RUNNER_TEMP/holdout.json" || true
node .holdout/summarize.mjs "$RUNNER_TEMP/holdout.json"

The rules that keep a holdout a holdout:

  • The agent has no read access to the scenario repository. Not on the developer’s machine, not through an MCP server, not in its prompt. The token is a CI secret with read access to that one repository.
  • The build artifact must be self-contained. The holdout job restores only dist/, which works for a static or fully bundled app. If your server needs runtime dependencies, have the build job upload its production node_modules too, or point the scenarios at a deployed preview URL.
  • Build in one job, judge in another. The pull request’s npm ci and build run in the build job, which never sees the scenarios. The holdout job has no pull request checkout and no pull request node_modules: Playwright comes from the holdout repository’s own lockfile, so a dependency the pull request added cannot read .holdout/ and print it to the log. persist-credentials: false keeps the token out of the checkout’s git config.
  • Name the residual risk. The application under test is still pull request code, and its server runs on the same runner as .holdout/. For a high-stakes set, run the app in a container that does not mount .holdout/, or point the scenarios at a deployed preview URL.
  • Logs carry IDs, not content. summarize.mjs prints something like holdout: 41 passed, 2 failed (HX-07, HX-19) and exits 1 on any failure. It also exits 1 when the JSON report is missing or holds zero results, so a Playwright crash (hidden by || true) fails the job instead of passing it. If an agent later reads the CI log to fix its pull request, it learns which scenario failed, not what the scenario checks.
  • A failed holdout goes to a human, who turns it into a spec change worded as behaviour, not as the scenario.
  • Secrets reach same-repository branches only. GitHub does not pass secrets to pull_request runs from forks, so the holdout job does not run for fork pull requests.

Scenarios written from executable acceptance criteria make good holdouts. Invariant checks from property-based testing are harder still to overfit, because they test a rule rather than an example.

Detect test weakening when tests legitimately change

Section titled “Detect test weakening when tests legitimately change”

Tests have to change when behaviour changes, so layers 1 and 2 cannot mean “tests never change”. They mean “tests never change silently”. The audit script flags the mechanical forms of weakening and fails the job so a code owner has to look:

#!/usr/bin/env bash
# oracle-audit.sh BASE_SHA: flag changes that can make a check pass without fixing the code
set -euo pipefail
base="$1"
oracle='(^|/)(tests?|__snapshots__)/|\.(test|spec)\.[jt]sx?$|_test\.(py|go)$|^\.github/|^CODEOWNERS$|(vitest|jest|playwright)\.config\.|^\.coveragerc$|^pytest\.ini$'
changed=$(git diff --name-only "$base"...HEAD | grep -E "$oracle" || true)
[ -z "$changed" ] && { echo "Oracle untouched."; exit 0; }
echo "Oracle files changed:"; echo "$changed" | sed 's/^/ /'
diff=$(git diff -U0 "$base"...HEAD -- $changed)
skips=$(grep -cE '^\+.*(\.skip\(|\.only\(|\bxit\(|\bxdescribe\(|@pytest\.mark\.(skip|xfail)|t\.Skip\()' <<<"$diff" || true)
removed=$(grep -cE '^-.*\b(expect|assert)' <<<"$diff" || true)
added=$(grep -cE '^\+.*\b(expect|assert)' <<<"$diff" || true)
snaps=$(grep -cE '__snapshots__/|\.snap$' <<<"$changed" || true)
configs=$(grep -cE '(vitest|jest|playwright)\.config\.|^\.coveragerc$|^pytest\.ini$' <<<"$changed" || true)
echo "Added skip/only/xfail markers: $skips"
echo "Assertion lines removed: $removed, added: $added"
echo "Snapshot files changed: $snaps"
echo "Runner or coverage config files changed: $configs"
if [ "$skips" -gt 0 ] || [ "$removed" -gt "$added" ] || [ "$snaps" -gt 0 ] || [ "$configs" -gt 0 ]; then
echo "WEAKENING SIGNAL: a code owner must approve this oracle change." >&2
exit 1
fi

We ran it against a test changed from two exact assertions to one toBeTruthy() inside test.skip: one skip marker, two assertion lines removed, one added, exit 1. An exclude glob added to vitest.config.ts counted one config file and exited 1. A pull request that only added a test passed. The script counts lines; it does not understand them. A swap from toBe(19.99) to toBeCloseTo(20, 0) keeps the count level, which is why every oracle change also goes to a code owner, and why the review prompt below asks about semantics. Extend the patterns with your stack’s suppressions (@ts-ignore, # noqa, eslint-disable) and threshold keys.

The cleanest defence is never to let one session hold both pens. This is test-driven development with the provenance made explicit:

  1. Test-authoring session. The agent writes or changes tests from the acceptance criteria, with production code locked (in Claude Code, Edit(src/**) in deny). A code owner approves the tests. Reading tests is reading the spec, which is far shorter than reading the implementation.
  2. Implementation session. A new session makes the approved tests pass with the oracle locked. If it concludes a test is wrong, it stops and writes TEST_DISPUTE.md, and a human decides.

Prove the lock with a red-team drill in a scratch branch, and keep the result as evidence.

Read the table against the layers. Routes 1 to 4 should stop in the session in Claude Code and Codex. Route 5 should stop in the session if the runner config is locked; if it gets through, the audit fails on the config change and base-oracle runs with the base branch’s config. Route 6 should stop in the session, and would still fail to matter because the required workflow lives elsewhere. Push the branch anyway and confirm that CODEOWNERS requests the owners and that the audit job fails. Any route that got through becomes a fix in the layer that should have stopped it.

The code owners of the oracle paths approve every oracle change. The tech lead owns the lock configuration and runs the drill on adoption, after each agent CLI or CI runner upgrade, and when the layout changes. Track one number per week: the share of agent pull requests that modify oracle files, and how many of those the audit flagged. A rising share means test and code changes are mixing again.

git checkout or git merge fails with “unable to unlink old”. The Claude Code sandbox refuses to replace a file under a denyWrite path, and branch switches need to do exactly that. Recovery: switch branches outside the agent session, or give each task its own worktree (see running parallel agents in worktrees) so the agent never switches branches.

The agent is stuck because a test really is wrong. A lock with no exit makes the agent thrash. Recovery: the TEST_DISPUTE.md instruction in the hook and the implementation prompt gives it a legal way to stop; a human decides, and any test change happens in a test-authoring session.

Holdout content leaked into an agent’s context. Someone pasted a scenario into a prompt, or the log printed assertion text. Recovery: move that scenario into the visible suite, write a replacement, and fix summarize.mjs.

The oracle was never strong enough to protect. A locked suite that does not catch bugs is locked theatre. Recovery: measure it with mutation testing, as covered in how strong your oracle is, before you trust green from an unattended loop.

Instructions and the lock disagree. CLAUDE.md or AGENTS.md tells the agent to “update tests as needed”, and the lock blocks it every time. Recovery: replace that line with the two-phase protocol, so instructions and enforcement say the same thing.