Skip to content

Error-Driven Development

Error-driven development is a way of working with AI coding agents in which each compiler error, failing test, or runtime exception becomes the agent’s next specification. The agent fixes the first error’s root cause, re-runs the same check, and repeats. The loop is trustworthy only when a final, untouched gate proves it converged.

This page is for developers who run agents day to day and tech leads who want the loop to behave the same way for the whole team. It assumes a check command that runs in seconds; if you do not have one yet, start with giving the session a feedback loop.

CI is red: 14 TypeScript errors, three failing tests, and a stack trace pointing at payment.ts:212. You have spent twenty minutes guessing which of the last six commits broke it, and your last fix turned one failing test into three. Pasting the whole log with “fix everything” makes it worse, because the agent patches symptoms in parallel and each patch creates new errors.

  • A six-step loop you can run the same way in Claude Code, Codex, and Cursor, interactively or headless
  • Copy-paste prompts for a production stack trace, a compiler cascade, a failing-test loop, and a stuck loop
  • A four-check gate that proves the loop converged honestly, so nobody has to read every changed line
  • A way out when fixing one error keeps producing three more
  • The Sentry MCP setup that replaces copy-pasted stack traces with an issue ID
  • A habit that turns a hard-won fix into a check or a rule, so the same error never returns
  1. Make a change, or let the agent make one.
  2. Run the verification command: type-check, lint, the relevant tests, or the build.
  3. Take the first error only.
  4. Give it to the agent with the files it names and one sentence on what the code is meant to do.
  5. Ask for the root cause and the fix, not a suppression.
  6. Re-run the exact command from step 2. Repeat until it passes, then run the final gate described in how to prove the loop converged.

The discipline that makes the loop work is one error at a time. Errors cascade: one wrong type produces dozens of downstream errors, and fixing the first often removes many of the rest. Fixing error 14 while error 1 is still present wastes a cycle and often introduces a conflicting fix.

The second discipline is a bound. An unbounded loop against a check it cannot satisfy burns context and money. Cap the number of attempts in the prompt, and in headless runs add an outer bound too: a spend flag where the tool has one (--max-budget-usd in Claude Code), or a timeout around the job where it does not.

How to hand an error to Claude Code, Codex, and Cursor

Section titled “How to hand an error to Claude Code, Codex, and Cursor”

The cycle is the same in every tool: surface the error, add context, fix, and re-run the command that failed. The tools differ in who runs the command, how you run the loop without a person watching, and how you back out a bad fix.

In an interactive session, Claude Code runs the check through its Bash tool, reads the output, and iterates without you copying anything. Paste this into the prompt:

Run npm run type-check. Fix the FIRST error only: read the file and line it
names, find the root cause, fix it, and re-run npm run type-check. Repeat.
Never add @ts-ignore, @ts-expect-error or `as any`. After 6 attempts, or if
the same error comes back twice, stop and report what you tried.

To run the same loop headless from a terminal or a CI job, use print mode with an explicit bound and a narrow tool allowance:

Terminal window
claude -p "Run npm run type-check and fix the errors one at a time, root cause first. Never suppress an error. Stop when it passes." \
--max-budget-usd 2 \
--permission-mode acceptEdits \
--allowedTools "Bash(npm run type-check)"

--max-budget-usd stops a loop that cannot converge by capping what it can spend. acceptEdits lets the agent edit files without a prompt, and --allowedTools limits it to the one command it needs. Keep the prompt directly after -p: --allowedTools accepts several values, so a prompt placed after it is read as another tool name and the run stops with no input. Print mode skips the workspace trust dialog, so run it only in a directory you trust. In an interactive session, undo a bad fix with /rewind (or Esc Esc).

A user hit a crash and your error tracker captured it. Give the agent the full exception and stack, not the top line: the top frame is where the value was read, rarely where it went wrong.

The right fix addresses where total became undefined (for example, a cart created with no items), not an optional-chaining ?. added at line 212. The regression test is what makes the fix provable: it fails before the change and passes after it.

With the Sentry MCP server connected, the agent fetches the issue, stack trace, and breadcrumbs itself, and the prompt shrinks to “fetch Sentry issue PROJ-1234 and fix the root cause”. Sentry’s official server is remote at https://mcp.sentry.dev/mcp and signs in with OAuth:

Terminal window
claude mcp add --transport http sentry https://mcp.sentry.dev/mcp

If you authenticate with a token instead of OAuth, the header is Authorization: Sentry-Bearer <token>, not Bearer. Debugging production from the editor covers Sentry alongside Grafana, Datadog, and PostHog.

Work through a compiler cascade after a refactor

Section titled “Work through a compiler cascade after a refactor”

You changed the signature of a core function and the type-checker reports 30 errors across the codebase. This is the loop’s strongest case, because the errors form an exact, machine-generated worklist. Do not fix anything by hand first; let the full list appear, then hand it over.

The instruction to stop on call sites with no currency matters: it turns a guess the agent would otherwise make silently into a short list you decide on.

The strongest form of the loop is a failing test written first, because the test is an oracle the agent can run but must not change. Test-driven development covers writing that test; the prompt below is the error-driven half.

The last sentence matters. Without it, an agent sometimes makes a failure pass by loosening the assertion. A prompt is a request, not a control, so protect the oracle with checks the agent cannot edit when the stakes are higher.

A green re-run of the command that failed proves less than it seems. The agent may have suppressed the error, skipped a test, weakened an assertion, or fixed one check while breaking another. Run this gate before you accept the result. Each check is mechanical, so nobody has to read every changed line.

  1. Run every check fresh, not only the one that failed. For example, npm run type-check && npm run lint && npm test. A type fix that breaks a test elsewhere shows up here.

  2. Scan the diff for suppressions. Any hit needs a written reason or a revert:

    Terminal window
    git diff main -U0 | grep -nE '^\+.*(@ts-ignore|@ts-expect-error|as any|eslint-disable|\.skip\(|\.only\(|catch( \([^)]*\))? *\{ *\})'
  3. Confirm the tests were not edited to pass. This command lists only removed or changed lines in test files and snapshots. It must print nothing: added tests are expected, removed or changed lines are not.

    Terminal window
    git diff main -U0 -- '*.test.*' '*.spec.*' '*__snapshots__*' | grep -E '^-[^-]'
  4. Prove each regression test catches the bug. Put the old source back, keep the new test, and watch it fail. Commit the fix first, then run, for example: git checkout main -- src/services/payment.ts && npx vitest run src/services/payment.test.ts; git checkout HEAD -- src/services/payment.ts. The test must fail here; the last command then restores the fixed file. A regression test that passes against the old code tests nothing.

The developer who ran the loop signs off on the root-cause explanation and on any hit from steps 2 and 3. CI running the same checks is the merge gate. To move steps 1 to 3 out of the prompt and into the harness, a Claude Code Stop hook that runs them and exits with code 2 on failure keeps the agent working instead of declaring victory. A hook that always fails keeps the agent going indefinitely, so have it check stop_hook_active (or cap its retries) and keep the --max-budget-usd bound on headless runs; see hooks as deterministic guardrails. Codex also has a Stop hook event.

How to break a cascade that will not converge

Section titled “How to break a cascade that will not converge”

A compiler worklist shrinks as you work through it. The dangerous cascade does the opposite: fixing error 1 introduces error 2, fixing that introduces error 3, and the context fills with failed attempts. The sign is the same file appearing in every round. At that point, stop fixing errors and look at the structure.

Use /rewind to return to the last good state, or /clear to drop the polluted context, then restart with the analysis prompt below. After two failed correction cycles on the same error, a fresh context usually does better than a longer one, because every failed attempt stays in the window and steers the next one.

How to make sure the same error never returns

Section titled “How to make sure the same error never returns”

The most valuable errors are the ones you never see again. When an error was hard to diagnose, capture the lesson in the strongest form available:

  1. A check that fails deterministically: a lint rule, a type, a test, or a CI step. This works for every agent and every person, whether or not they read instructions.
  2. A project instruction, when the lesson cannot be expressed as a check. Keep it short and specific.

Add the lesson to the project’s CLAUDE.md so the whole team gets it. Claude Code’s auto memory also stores learnings in ~/.claude/projects/<project>/memory/, but that directory is per user, not shared through the repository.

The content is the same in all three files:

## Known error patterns
- Notification types: update src/types/notification.ts AND the Zod schema in
src/validators/notification.ts together. They must stay in sync.
- Integration tests need Redis: run `docker compose up -d redis` first.

For how long-lived lessons are stored and pruned, see memory patterns.

  • The agent chases the wrong error. It fixes error 14 while error 1 causes it. Recovery: restate “fix only the first error, then re-run” and discard the out-of-order fix.
  • The fix treats the symptom. An optional-chaining ?. stops the crash without explaining why the value was missing. Recovery: ask “where did this value become undefined?” and require a regression test that fails on the old code.
  • The agent suppresses instead of fixing. @ts-ignore, as any, eslint-disable, and empty catch blocks all hide the error. Recovery: the suppression scan in the gate catches them; revert and re-prompt with the ban stated.
  • The test is edited to pass. Recovery: the test-file check in the gate catches it; restore the test from Git and make it read-only for the agent with checks it cannot edit.
  • The loop runs against a flaky test. A non-deterministic failure lets the agent “fix”, see green by chance, and declare victory. Recovery: run the test several times first; if it flakes, quarantine it before starting the loop.
  • The paste hides the culprit. Only the top stack frame was given. Recovery: paste the full trace, or fetch it through the Sentry MCP server, and name the suspect files.
  • Hundreds of errors arrive at once. Recovery: work the first three to five, which are usually the root causes, then re-run and see how many remain.
  • The error carries no information. A segfault or a bare “internal server error” is not a specification. Recovery: switch to hypothesis-driven debugging: reproduce, add logging, narrow by elimination. The systematic-debugging skill from obra/superpowers encodes that process (npx skills add obra/superpowers --skill systematic-debugging); read its SKILL.md before installing, because its directory also ships test fixtures. See testing and debugging skills.
  • The broken file is one the agent wrote from scratch. Recovery: delete it and regenerate from a tighter prompt that includes the error; that usually beats patching.

Where to go next with error-driven development

Section titled “Where to go next with error-driven development”