Skip to content

Test: give the session a feedback loop

The test stage gives every agent session a fast feedback loop that fails closed: one-command checks, a measurable definition of done, and a rule that the session cannot report success while a check is red. A fresh-context verifier and a CI re-run confirm the result, and evals regression-test the configuration that steers the agent.

This page is for the developer who runs agent sessions and the tech lead who owns the team’s verification contract. The agent says “done, all tests pass”. You pull the branch, run make test, and two tests fail. The session had never run them; it read the code and decided it worked. Now you are the test runner, pasting stack traces back into the chat.

Traditional: the signal that code works arrives late. CI reports minutes later, a tester days later, production weeks later. With an agent producing the code, a late signal means a person checks all of its output.

AI-native: the session checks its own work before a person sees it. It runs the tests, the build, and the screenshot comparison, reads the failures, and iterates until the checks pass. The person reads evidence, not every line.

The feedback loop and the verifier are two different things. The loop runs throughout the task. The verifier is one final check in a fresh context window, after the session believes it is finished.

  • A verification block for CLAUDE.md or AGENTS.md that the agent runs before it reports a task complete.
  • Four copy-paste prompts: a definition of done, a bug reproduction, a visual check, and a fresh-context verification.
  • A Claude Code Stop hook that stops a session from finishing while the gate is red, and the equivalents in Codex and Cursor.
  • A headless verifier run for CI with claude -p or codex exec.
  • An eval runner that tells you whether a change to instructions, skills, or hooks made the agent worse.

Before you give the session a feedback loop

Section titled “Before you give the session a feedback loop”
  • A test suite and a build that each run locally with one command.
  • For UI work, a browser tool the agent can drive, such as Playwright MCP.
  • An accepted plan.md with proof commands, from Build: plan.md, then implement. You can start the loop without one, but the verifier needs something to check the diff against.

Put a check in the loop when the agent can run it in seconds to a few minutes and read the result. Leave slow or shared checks to CI, where the agent cannot edit them.

CheckRunsCatchesOwner
Lint and type check on changed filesAfter every edit (hook)Syntax, types, style driftThe session
Unit tests (fast project)Before the session reports doneBroken behaviour in the changed moduleThe session
BuildBefore the session reports doneMissing imports, bundling errorsThe session
Screenshot or browser checkUI tasks, each iterationLayout and interaction bugsThe session
Verifier subagentOnce, at the endUnmet plan items, regressions in neighbouring codeA fresh context
Full suite, E2E, security scansCI on every pushEverything the session skipped or weakenedCI, outside the diff

Put the verification commands in the instruction file

Section titled “Put the verification commands in the instruction file”
  1. Wrap each check in a single target that exits non-zero on failure: make test, npm test, pytest. An agent cannot reliably tell a pass from a failure in a command that prints errors and exits 0.

  2. Add a verification block to the project instruction file, with the healthy output of each command. Claude Code reads CLAUDE.md, Codex reads AGENTS.md, and Cursor reads project Rules.

    ## Verifying your work
    - Build: make build (must finish with "Build succeeded")
    - Test: make test (all green; never skip or delete a failing test)
    - Lint: make lint (zero warnings)
    Run all three before reporting any task complete, and paste the
    output with each exit code. If a test fails, fix the code, not the test.
  3. Break the gate on purpose once. Change one assertion so it must fail, run make test, and confirm it goes red. A gate you have never seen fail proves nothing. How strong is your oracle? covers mutation testing for the same question at scale.

Give the agent a definition of done it can check

Section titled “Give the agent a definition of done it can check”

A subjective target (“make sure it works”) lets the agent stop when it feels finished. A measurable target lets it stop only when a command says so. Put the target in the prompt, next to the task.

Subjective targetMeasurable target
“Make sure the code looks good”“npm run lint exits 0 with zero warnings”
“Handle the edge cases”“Three new tests for signature failure, duplicate events and malformed payloads are green”
“Match the design”“The screenshot at 1280 and 375 pixels matches mocks/checkout.png; list every difference”
“Don’t break anything”“All existing tests in tests/billing/ still pass”

Make “done” impossible while a check is red

Section titled “Make “done” impossible while a check is red”

An instruction in CLAUDE.md is advisory; the agent can forget it in a long session. A hook runs every time. Wire the fast gate into the moment the agent tries to finish.

A Stop hook fires when Claude finishes responding. Exit code 2 blocks the stop and feeds stderr back to Claude, so it keeps working on the failure. Add this to .claude/settings.json:

{
"hooks": {
"Stop": [
{
"hooks": [
{
"type": "command",
"command": "cd \"$CLAUDE_PROJECT_DIR\" && make verify-fast 1>&2 || exit 2"
}
]
}
]
}
}

verify-fast is a make target you add that runs only lint, the type check and the unit tests, for example verify-fast: lint typecheck test-unit; the full make test stays in CI.

Stop takes no matcher. In Claude Code 2.1.283, a Stop hook that blocks eight times in a row is overridden and the turn ends; CLAUDE_CODE_STOP_HOOK_BLOCK_CAP changes that cap. A custom script can read stop_hook_active from the hook input and return success while it is true. Test generation and TDD with Claude Code adds a PostToolUse hook that runs only the tests related to the edited file.

A bug fix starts with a test that reproduces the bug and fails for the reason you expect. You commit that test before any application code changes. The fix run may not edit it. That order stops the most common way an agent “fixes” a bug: by weakening the check.

After you commit the test, start the fix with: “Make the test in tests/refunds/ pass without editing any file under tests/.” Then block test edits mechanically for the fix run. Protecting the oracle has the deny rules, hooks, and CODEOWNERS setup, and Test-driven development with AI assistance covers the red-green loop for new features.

Unit tests do not see a misaligned button. For UI work, give the agent a browser, the mock, and a measurable comparison, and let it iterate: implement, screenshot, compare, adjust. Playwright MCP (@playwright/mcp 0.0.82 on npm, 2026-09-26) is Microsoft’s official browser server.

Terminal window
claude mcp add playwright -- npx @playwright/mcp@latest --isolated

--isolated keeps the browser profile in memory. A persistent profile can be used by only one browser at a time, so parallel sessions in separate worktrees need it. Playwright’s README also recommends @playwright/cli with skills for coding agents, because MCP tool schemas and page snapshots cost context. Install it with npm install -g @playwright/cli@latest, then run playwright-cli install --skills to copy its skill into the workspace. Playwright MCP compares the two, and agent-browser is a separate browser CLI for agents if you want a second option.

The session that wrote the code is bound by its own assumptions; its “all tests pass” can mean “all the tests I thought of pass”. A verifier subagent starts with a clean context, reads the diff and plan.md, runs the checks itself, and reports without fixing anything. Define it once and commit it, so every session uses the same verifier:

A subagent is a Markdown file in .claude/agents/. The Build page has a complete verifier.md.

Give the verifier read and run tools only (Bash, Read, Grep in Claude Code), and make its instruction adversarial: look for regressions in the neighbouring modules, not only the changed lines.

Re-run the checks where the agent cannot edit them

Section titled “Re-run the checks where the agent cannot edit them”

The session’s pasted output is a claim. CI re-running the same commands is the proof, because CI’s configuration lives outside the diff the agent controls. Run the same make targets in CI, so the local gate and the merge gate cannot drift.

For a second opinion in CI, run the verifier headless with read-and-run tools only and a spend cap:

Terminal window
claude -p "Run make test and make build. Compare git diff origin/main...HEAD \
with plan.md and list any plan item with no passing proof. Do not edit files." \
--allowedTools "Read" "Grep" "Bash(make test)" "Bash(make build)" "Bash(git diff *)" \
--output-format json --max-budget-usd 2 > verifier.json

Name each make target: Bash(make *) would also allow make deploy or make clean. claude -p starts in Manual permission mode, so a tool call that needs approval and is not in --allowedTools is denied instead of prompted. --max-budget-usd works only with --print.

Regression-test the harness with continuous evals

Section titled “Regression-test the harness with continuous evals”

The code has tests; the configuration that steers the agent needs them too. A change to CLAUDE.md, AGENTS.md, a rule, a skill, a hook, or the model can make the agent worse at work it did well last week, and nothing in the product test suite notices. An eval suite is that test.

  1. Collect 20 to 50 real tasks from recent work, each with the base commit and the accepted outcome. Include the ones the agent got wrong.

  2. Write each task as a case: prompt.md, base-commit, and a check.sh that decides pass or fail with tests kept outside the worktree the agent runs in, so the agent under test cannot see or edit them. Keep the cases out of the base commits, or store them outside the repository.

  3. Run the suite on every change to the harness files, and on a schedule to catch model and tool updates.

  4. Gate the change on the pass rate. A skill edit that drops the rate gets reviewed before it merges.

  5. Turn each production incident into a case, written by the team that owned the incident, and keep it as a regression test.

A minimal runner, evals/run.sh, run from the repository root:

#!/usr/bin/env bash
set -u
root=$(pwd); pass=0; total=0; mkdir -p evals/out
for case in evals/cases/*/; do
id=$(basename "$case"); total=$((total + 1)); dir=$(mktemp -d)
git worktree add --detach "$dir" "$(cat "$case/base-commit")" >/dev/null \
|| { echo "SKIP $id (bad base commit)"; total=$((total - 1)); continue; }
# The base commit has the OLD harness: copy the version under test into it.
cp -R CLAUDE.md .claude "$dir"/ 2>/dev/null
(cd "$dir" && claude -p "$(cat "$root/$case/prompt.md")" \
--permission-mode acceptEdits \
--allowedTools "Bash(make test)" "Bash(make build)" "Bash(make verify-fast)" \
--max-budget-usd 2 --output-format json > "$root/evals/out/$id.json")
if bash "$case/check.sh" "$dir"; then pass=$((pass + 1)); echo "PASS $id"; else echo "FAIL $id"; fi
git worktree remove --force "$dir"
done
echo "pass rate: $pass/$total"
[ "$total" -gt 0 ] || { echo "no cases ran"; exit 1; }
[ $((pass * 100 / total)) -ge "${EVAL_THRESHOLD:-90}" ]

For Codex, replace the agent line with codex exec -C "$dir" -c default_permissions=":workspace" --ephemeral "$(cat "$root/$case/prompt.md")" (permission profiles are beta in Codex 0.157.1) and copy AGENTS.md and .codex/ instead. Run the suite several times per change: agents are not deterministic, and one run of one case is a sample, not a verdict. Cases that every configuration passes no longer discriminate; replace them with new ones from monitoring. Continuous evals covers scheduling and reporting, and Maintain covers turning incidents into cases.

How do you know the feedback loop works without reading every line?

Section titled “How do you know the feedback loop works without reading every line?”

DORA’s 2025 report (Google Cloud, 2025-09-23) found that “without robust control systems, like strong automated testing, mature version control practices, and fast feedback loops, an increase in change volume leads to instability”. The loop is that control system for each session. Check it by its evidence:

  • Evidence arrives with the claim. Every “done” message contains the pasted output and exit code of each verification command. A message without them is not done.
  • CI agrees with the session. The first CI run on an agent’s pull request is green. When the session said green and CI says red, the loop is broken; find out why before the next task.
  • Bug fixes keep their test. The failing test was committed before the fix, and the fix diff touches nothing under the protected test paths.
  • The verifier’s table is attached. The pull request carries the verifier’s plan-item table. The evidence bundle defines the full set.
  • Harness changes carry an eval result. A pull request that edits instructions, skills, or hooks shows the pass rate before and after.

Who signs off. The developer accepts the evidence for routine changes and reads the diff for intent and risk, not for compile errors. The tech lead owns the verification block, the stop gate, and the eval threshold, and reviews every change to them, because an agent that can edit the gate can weaken it.

Leading indicators: first-pass CI success rate for agent-written changes, and eval pass rate over time. Lagging indicators: review time per pull request, change failure rate, and regressions caught in CI versus in production. Metrics for an AI-native SDLC defines each one.

What breaks in a session feedback loop, and how to recover

Section titled “What breaks in a session feedback loop, and how to recover”

The agent reports done without running anything. The verification block is in the instruction file, but the session skipped it late in a long context. Add the Stop hook, so finishing is impossible while the gate is red, and treat any “done” without pasted exit codes as not done.

The agent edits the test instead of the code. It deletes an assertion, adds a skip, or special-cases the test input. Commit the reproducing test first, block edits to test paths during fix runs, and have CI flag any fix diff that touches tests. Protecting the oracle has the mechanics.

The loop spins on something the code cannot fix. A database is down, a secret is missing, or an external API rate-limits the tests. The agent retries or starts mocking production code. Put mock boundaries for external services in the instruction file, tell the agent to stop and report environment failures, and rely on the Stop hook cap to end the turn.

The gate is green but proves nothing. Tests assert nothing, or a health route returns 200 without touching the database. Break an assertion on purpose whenever the gate changes, and measure test strength with mutation testing.

Flaky tests teach the agent to rerun until green. The session reruns a failing test three times, it passes once, and the session reports success. Quarantine known flaky tests out of the session gate, and tell the agent to report a test that fails and then passes as flaky instead of calling it fixed.

The verifier agrees with everything. It received the implementing session’s summary and inherited its assumptions. Give it only the diff, the plan, and the commands, and require it to run each command itself.

Parallel sessions fight over the browser. Two worktrees share one Playwright profile and one dev-server port. Use --isolated and assign each worktree its own port.

The eval suite passes forever. Every case is solved by every configuration, so the suite no longer tells configurations apart. Retire saturated cases and add new ones from recent failures and incidents.