Test: give the session a feedback loop
The test stage gives every agent session a fast feedback loop that fails closed: one-command checks, a measurable definition of done, and a rule that the session cannot report success while a check is red. A fresh-context verifier and a CI re-run confirm the result, and evals regression-test the configuration that steers the agent.
This page is for the developer who runs agent sessions and the tech lead who owns the team’s verification contract. The agent says “done, all tests pass”. You pull the branch, run make test, and two tests fail. The session had never run them; it read the code and decided it worked. Now you are the test runner, pasting stack traces back into the chat.
Traditional: the signal that code works arrives late. CI reports minutes later, a tester days later, production weeks later. With an agent producing the code, a late signal means a person checks all of its output.
AI-native: the session checks its own work before a person sees it. It runs the tests, the build, and the screenshot comparison, reads the failures, and iterates until the checks pass. The person reads evidence, not every line.
The feedback loop and the verifier are two different things. The loop runs throughout the task. The verifier is one final check in a fresh context window, after the session believes it is finished.
What a session feedback loop gives you
Section titled “What a session feedback loop gives you”- A verification block for
CLAUDE.mdorAGENTS.mdthat the agent runs before it reports a task complete. - Four copy-paste prompts: a definition of done, a bug reproduction, a visual check, and a fresh-context verification.
- A Claude Code
Stophook that stops a session from finishing while the gate is red, and the equivalents in Codex and Cursor. - A headless verifier run for CI with
claude -porcodex exec. - An eval runner that tells you whether a change to instructions, skills, or hooks made the agent worse.
Before you give the session a feedback loop
Section titled “Before you give the session a feedback loop”- A test suite and a build that each run locally with one command.
- For UI work, a browser tool the agent can drive, such as Playwright MCP.
- An accepted
plan.mdwith proof commands, from Build: plan.md, then implement. You can start the loop without one, but the verifier needs something to check the diff against.
Which checks belong in the session loop?
Section titled “Which checks belong in the session loop?”Put a check in the loop when the agent can run it in seconds to a few minutes and read the result. Leave slow or shared checks to CI, where the agent cannot edit them.
| Check | Runs | Catches | Owner |
|---|---|---|---|
| Lint and type check on changed files | After every edit (hook) | Syntax, types, style drift | The session |
| Unit tests (fast project) | Before the session reports done | Broken behaviour in the changed module | The session |
| Build | Before the session reports done | Missing imports, bundling errors | The session |
| Screenshot or browser check | UI tasks, each iteration | Layout and interaction bugs | The session |
| Verifier subagent | Once, at the end | Unmet plan items, regressions in neighbouring code | A fresh context |
| Full suite, E2E, security scans | CI on every push | Everything the session skipped or weakened | CI, outside the diff |
Put the verification commands in the instruction file
Section titled “Put the verification commands in the instruction file”-
Wrap each check in a single target that exits non-zero on failure:
make test,npm test,pytest. An agent cannot reliably tell a pass from a failure in a command that prints errors and exits 0. -
Add a verification block to the project instruction file, with the healthy output of each command. Claude Code reads
CLAUDE.md, Codex readsAGENTS.md, and Cursor reads project Rules.## Verifying your work- Build: make build (must finish with "Build succeeded")- Test: make test (all green; never skip or delete a failing test)- Lint: make lint (zero warnings)Run all three before reporting any task complete, and paste theoutput with each exit code. If a test fails, fix the code, not the test. -
Break the gate on purpose once. Change one assertion so it must fail, run
make test, and confirm it goes red. A gate you have never seen fail proves nothing. How strong is your oracle? covers mutation testing for the same question at scale.
Give the agent a definition of done it can check
Section titled “Give the agent a definition of done it can check”A subjective target (“make sure it works”) lets the agent stop when it feels finished. A measurable target lets it stop only when a command says so. Put the target in the prompt, next to the task.
| Subjective target | Measurable target |
|---|---|
| “Make sure the code looks good” | “npm run lint exits 0 with zero warnings” |
| “Handle the edge cases” | “Three new tests for signature failure, duplicate events and malformed payloads are green” |
| “Match the design” | “The screenshot at 1280 and 375 pixels matches mocks/checkout.png; list every difference” |
| “Don’t break anything” | “All existing tests in tests/billing/ still pass” |
Make “done” impossible while a check is red
Section titled “Make “done” impossible while a check is red”An instruction in CLAUDE.md is advisory; the agent can forget it in a long session. A hook runs every time. Wire the fast gate into the moment the agent tries to finish.
A Stop hook fires when Claude finishes responding. Exit code 2 blocks the stop and feeds stderr back to Claude, so it keeps working on the failure. Add this to .claude/settings.json:
{ "hooks": { "Stop": [ { "hooks": [ { "type": "command", "command": "cd \"$CLAUDE_PROJECT_DIR\" && make verify-fast 1>&2 || exit 2" } ] } ] }}verify-fast is a make target you add that runs only lint, the type check and the unit tests, for example verify-fast: lint typecheck test-unit; the full make test stays in CI.
Stop takes no matcher. In Claude Code 2.1.283, a Stop hook that blocks eight times in a row is overridden and the turn ends; CLAUDE_CODE_STOP_HOOK_BLOCK_CAP changes that cap. A custom script can read stop_hook_active from the hook input and return success while it is true. Test generation and TDD with Claude Code adds a PostToolUse hook that runs only the tests related to the edited file.
Codex CLI 0.157.1 has hooks with 12 events, including Stop, PostToolUse and SubagentStop. Codex reads .codex/hooks.json in a trusted project, in the same JSON shape as Claude Code, and runs the command in the session’s working directory, so resolve the repository root yourself:
{ "hooks": { "Stop": [ { "hooks": [ { "type": "command", "command": "cd \"$(git rev-parse --show-toplevel)\" && make verify-fast 1>&2 || exit 2", "timeout": 600 } ] } ] }}Exit code 2 continues the turn with stderr as the next prompt. Project hooks run only after you trust them in /hooks, and an administrator can restrict a fleet to managed hooks with allow_managed_hooks_only = true in requirements.toml. For a longer task, /goal keeps Codex working toward a stated completion condition. Use hooks as deterministic guardrails has a complete stop-gate.sh for both tools.
Cursor Hooks are processes that exchange JSON over stdio and run before or after stages of the agent loop. Register a stop hook that runs the fast gate when the agent finishes; loop_limit caps how many times it can send the agent back.
{ "version": 1, "hooks": { "stop": [ { "command": "node .cursor/hooks/verify-on-stop.mjs", "timeout": 600, "loop_limit": 3 } ] }}The script runs your gates and, on failure, returns the tail of the output as a followup_message; Advanced Cursor techniques has the complete verify-on-stop.mjs. Confirm the field names in Cursor’s Hooks docs for your version before you copy them. Put the verification block in a project rule, because Cursor reads Rules, not CLAUDE.md.
Start a bug fix from a failing test
Section titled “Start a bug fix from a failing test”A bug fix starts with a test that reproduces the bug and fails for the reason you expect. You commit that test before any application code changes. The fix run may not edit it. That order stops the most common way an agent “fixes” a bug: by weakening the check.
After you commit the test, start the fix with: “Make the test in tests/refunds/ pass without editing any file under tests/.” Then block test edits mechanically for the fix run. Protecting the oracle has the deny rules, hooks, and CODEOWNERS setup, and Test-driven development with AI assistance covers the red-green loop for new features.
Close the loop on UI with a browser
Section titled “Close the loop on UI with a browser”Unit tests do not see a misaligned button. For UI work, give the agent a browser, the mock, and a measurable comparison, and let it iterate: implement, screenshot, compare, adjust. Playwright MCP (@playwright/mcp 0.0.82 on npm, 2026-09-26) is Microsoft’s official browser server.
claude mcp add playwright -- npx @playwright/mcp@latest --isolatedcodex mcp add playwright -- npx @playwright/mcp@latest --isolatedAdd the server to .cursor/mcp.json:
{ "mcpServers": { "playwright": { "command": "npx", "args": ["@playwright/mcp@latest", "--isolated"] } } }--isolated keeps the browser profile in memory. A persistent profile can be used by only one browser at a time, so parallel sessions in separate worktrees need it. Playwright’s README also recommends @playwright/cli with skills for coding agents, because MCP tool schemas and page snapshots cost context. Install it with npm install -g @playwright/cli@latest, then run playwright-cli install --skills to copy its skill into the workspace. Playwright MCP compares the two, and agent-browser is a separate browser CLI for agents if you want a second option.
Finish with a verifier in a fresh context
Section titled “Finish with a verifier in a fresh context”The session that wrote the code is bound by its own assumptions; its “all tests pass” can mean “all the tests I thought of pass”. A verifier subagent starts with a clean context, reads the diff and plan.md, runs the checks itself, and reports without fixing anything. Define it once and commit it, so every session uses the same verifier:
A subagent is a Markdown file in .claude/agents/. The Build page has a complete verifier.md.
Ask the session to spawn a subagent; multi-agent is on by default in 0.157.1, and /subagents switches between them. Keep the verifier’s brief in AGENTS.md or a skill so every session sends the same one. Codex multi-agent workflows covers the settings.
Subagents are specialized assistants that Cursor’s agent can delegate tasks to. Create one with the verifier’s brief and keep the hand-off instruction in a project rule, so every session delegates the final check the same way.
Give the verifier read and run tools only (Bash, Read, Grep in Claude Code), and make its instruction adversarial: look for regressions in the neighbouring modules, not only the changed lines.
Re-run the checks where the agent cannot edit them
Section titled “Re-run the checks where the agent cannot edit them”The session’s pasted output is a claim. CI re-running the same commands is the proof, because CI’s configuration lives outside the diff the agent controls. Run the same make targets in CI, so the local gate and the merge gate cannot drift.
For a second opinion in CI, run the verifier headless with read-and-run tools only and a spend cap:
claude -p "Run make test and make build. Compare git diff origin/main...HEAD \with plan.md and list any plan item with no passing proof. Do not edit files." \ --allowedTools "Read" "Grep" "Bash(make test)" "Bash(make build)" "Bash(git diff *)" \ --output-format json --max-budget-usd 2 > verifier.jsonName each make target: Bash(make *) would also allow make deploy or make clean. claude -p starts in Manual permission mode, so a tool call that needs approval and is not in --allowedTools is denied instead of prompted. --max-budget-usd works only with --print.
# The gate: a red check fails the job here, before the agent starts.make test > test.log 2>&1 && make build > build.log 2>&1codex exec -c default_permissions=":read-only" --ephemeral -o verifier.md \ "Read test.log and build.log. Compare git diff origin/main...HEAD with plan.md \and list any plan item with no passing proof. Do not edit files."The :read-only permission profile (beta in Codex 0.157.1) blocks writes, and a build writes output by definition, so CI runs make test and make build as ordinary steps and the read-only verifier only reads their logs, the diff and plan.md. In GitHub Actions, openai/codex-action@v1 runs codex exec for you.
# The gate: a red check fails the job here, before the agent starts.make test > test.log 2>&1 && make build > build.log 2>&1agent -p "Read test.log and build.log. Compare git diff origin/main...HEAD with plan.md \and list any plan item with no passing proof. Do not edit files." \ --mode ask --trust --output-format text > verifier.md--mode ask is the read-only mode that answers without editing files, so CI runs the checks as ordinary steps, as in the Codex tab. --trust answers the headless workspace-trust prompt; use it only on a checkout you trust. These flags come from Cursor’s cookbook and public workflows, so confirm them in agent --help on your runner. Scripting the Cursor CLI has the full GitHub Actions job, and Cloud Agents can run the same checks in an isolated VM.
Regression-test the harness with continuous evals
Section titled “Regression-test the harness with continuous evals”The code has tests; the configuration that steers the agent needs them too. A change to CLAUDE.md, AGENTS.md, a rule, a skill, a hook, or the model can make the agent worse at work it did well last week, and nothing in the product test suite notices. An eval suite is that test.
-
Collect 20 to 50 real tasks from recent work, each with the base commit and the accepted outcome. Include the ones the agent got wrong.
-
Write each task as a case:
prompt.md,base-commit, and acheck.shthat decides pass or fail with tests kept outside the worktree the agent runs in, so the agent under test cannot see or edit them. Keep the cases out of the base commits, or store them outside the repository. -
Run the suite on every change to the harness files, and on a schedule to catch model and tool updates.
-
Gate the change on the pass rate. A skill edit that drops the rate gets reviewed before it merges.
-
Turn each production incident into a case, written by the team that owned the incident, and keep it as a regression test.
A minimal runner, evals/run.sh, run from the repository root:
#!/usr/bin/env bashset -uroot=$(pwd); pass=0; total=0; mkdir -p evals/outfor case in evals/cases/*/; do id=$(basename "$case"); total=$((total + 1)); dir=$(mktemp -d) git worktree add --detach "$dir" "$(cat "$case/base-commit")" >/dev/null \ || { echo "SKIP $id (bad base commit)"; total=$((total - 1)); continue; } # The base commit has the OLD harness: copy the version under test into it. cp -R CLAUDE.md .claude "$dir"/ 2>/dev/null (cd "$dir" && claude -p "$(cat "$root/$case/prompt.md")" \ --permission-mode acceptEdits \ --allowedTools "Bash(make test)" "Bash(make build)" "Bash(make verify-fast)" \ --max-budget-usd 2 --output-format json > "$root/evals/out/$id.json") if bash "$case/check.sh" "$dir"; then pass=$((pass + 1)); echo "PASS $id"; else echo "FAIL $id"; fi git worktree remove --force "$dir"doneecho "pass rate: $pass/$total"[ "$total" -gt 0 ] || { echo "no cases ran"; exit 1; }[ $((pass * 100 / total)) -ge "${EVAL_THRESHOLD:-90}" ]For Codex, replace the agent line with codex exec -C "$dir" -c default_permissions=":workspace" --ephemeral "$(cat "$root/$case/prompt.md")" (permission profiles are beta in Codex 0.157.1) and copy AGENTS.md and .codex/ instead. Run the suite several times per change: agents are not deterministic, and one run of one case is a sample, not a verdict. Cases that every configuration passes no longer discriminate; replace them with new ones from monitoring. Continuous evals covers scheduling and reporting, and Maintain covers turning incidents into cases.
How do you know the feedback loop works without reading every line?
Section titled “How do you know the feedback loop works without reading every line?”DORA’s 2025 report (Google Cloud, 2025-09-23) found that “without robust control systems, like strong automated testing, mature version control practices, and fast feedback loops, an increase in change volume leads to instability”. The loop is that control system for each session. Check it by its evidence:
- Evidence arrives with the claim. Every “done” message contains the pasted output and exit code of each verification command. A message without them is not done.
- CI agrees with the session. The first CI run on an agent’s pull request is green. When the session said green and CI says red, the loop is broken; find out why before the next task.
- Bug fixes keep their test. The failing test was committed before the fix, and the fix diff touches nothing under the protected test paths.
- The verifier’s table is attached. The pull request carries the verifier’s plan-item table. The evidence bundle defines the full set.
- Harness changes carry an eval result. A pull request that edits instructions, skills, or hooks shows the pass rate before and after.
Who signs off. The developer accepts the evidence for routine changes and reads the diff for intent and risk, not for compile errors. The tech lead owns the verification block, the stop gate, and the eval threshold, and reviews every change to them, because an agent that can edit the gate can weaken it.
Leading indicators: first-pass CI success rate for agent-written changes, and eval pass rate over time. Lagging indicators: review time per pull request, change failure rate, and regressions caught in CI versus in production. Metrics for an AI-native SDLC defines each one.
What breaks in a session feedback loop, and how to recover
Section titled “What breaks in a session feedback loop, and how to recover”The agent reports done without running anything. The verification block is in the instruction file, but the session skipped it late in a long context. Add the Stop hook, so finishing is impossible while the gate is red, and treat any “done” without pasted exit codes as not done.
The agent edits the test instead of the code. It deletes an assertion, adds a skip, or special-cases the test input. Commit the reproducing test first, block edits to test paths during fix runs, and have CI flag any fix diff that touches tests. Protecting the oracle has the mechanics.
The loop spins on something the code cannot fix. A database is down, a secret is missing, or an external API rate-limits the tests. The agent retries or starts mocking production code. Put mock boundaries for external services in the instruction file, tell the agent to stop and report environment failures, and rely on the Stop hook cap to end the turn.
The gate is green but proves nothing. Tests assert nothing, or a health route returns 200 without touching the database. Break an assertion on purpose whenever the gate changes, and measure test strength with mutation testing.
Flaky tests teach the agent to rerun until green. The session reruns a failing test three times, it passes once, and the session reports success. Quarantine known flaky tests out of the session gate, and tell the agent to report a test that fails and then passes as flaky instead of calling it fixed.
The verifier agrees with everything. It received the implementing session’s summary and inherited its assumptions. Give it only the diff, the plan, and the commands, and require it to run each command itself.
Parallel sessions fight over the browser. Two worktrees share one Playwright profile and one dev-server port. Use --isolated and assign each worktree its own port.
The eval suite passes forever. Every case is solved by every configuration, so the suite no longer tells configurations apart. Retire saturated cases and add new ones from recent failures and incidents.