Error-Driven Development
Error-driven development is a way of working with AI coding agents in which each compiler error, failing test, or runtime exception becomes the agent’s next specification. The agent fixes the first error’s root cause, re-runs the same check, and repeats. The loop is trustworthy only when a final, untouched gate proves it converged.
This page is for developers who run agents day to day and tech leads who want the loop to behave the same way for the whole team. It assumes a check command that runs in seconds; if you do not have one yet, start with giving the session a feedback loop.
CI is red: 14 TypeScript errors, three failing tests, and a stack trace pointing at payment.ts:212. You have spent twenty minutes guessing which of the last six commits broke it, and your last fix turned one failing test into three. Pasting the whole log with “fix everything” makes it worse, because the agent patches symptoms in parallel and each patch creates new errors.
What error-driven development gives you
Section titled “What error-driven development gives you”- A six-step loop you can run the same way in Claude Code, Codex, and Cursor, interactively or headless
- Copy-paste prompts for a production stack trace, a compiler cascade, a failing-test loop, and a stuck loop
- A four-check gate that proves the loop converged honestly, so nobody has to read every changed line
- A way out when fixing one error keeps producing three more
- The Sentry MCP setup that replaces copy-pasted stack traces with an issue ID
- A habit that turns a hard-won fix into a check or a rule, so the same error never returns
How the error-fix-rerun loop works
Section titled “How the error-fix-rerun loop works”- Make a change, or let the agent make one.
- Run the verification command: type-check, lint, the relevant tests, or the build.
- Take the first error only.
- Give it to the agent with the files it names and one sentence on what the code is meant to do.
- Ask for the root cause and the fix, not a suppression.
- Re-run the exact command from step 2. Repeat until it passes, then run the final gate described in how to prove the loop converged.
The discipline that makes the loop work is one error at a time. Errors cascade: one wrong type produces dozens of downstream errors, and fixing the first often removes many of the rest. Fixing error 14 while error 1 is still present wastes a cycle and often introduces a conflicting fix.
The second discipline is a bound. An unbounded loop against a check it cannot satisfy burns context and money. Cap the number of attempts in the prompt, and in headless runs add an outer bound too: a spend flag where the tool has one (--max-budget-usd in Claude Code), or a timeout around the job where it does not.
How to hand an error to Claude Code, Codex, and Cursor
Section titled “How to hand an error to Claude Code, Codex, and Cursor”The cycle is the same in every tool: surface the error, add context, fix, and re-run the command that failed. The tools differ in who runs the command, how you run the loop without a person watching, and how you back out a bad fix.
In an interactive session, Claude Code runs the check through its Bash tool, reads the output, and iterates without you copying anything. Paste this into the prompt:
Run npm run type-check. Fix the FIRST error only: read the file and line itnames, find the root cause, fix it, and re-run npm run type-check. Repeat.Never add @ts-ignore, @ts-expect-error or `as any`. After 6 attempts, or ifthe same error comes back twice, stop and report what you tried.To run the same loop headless from a terminal or a CI job, use print mode with an explicit bound and a narrow tool allowance:
claude -p "Run npm run type-check and fix the errors one at a time, root cause first. Never suppress an error. Stop when it passes." \ --max-budget-usd 2 \ --permission-mode acceptEdits \ --allowedTools "Bash(npm run type-check)"--max-budget-usd stops a loop that cannot converge by capping what it can spend. acceptEdits lets the agent edit files without a prompt, and --allowedTools limits it to the one command it needs. Keep the prompt directly after -p: --allowedTools accepts several values, so a prompt placed after it is read as another tool name and the run stops with no input. Print mode skips the workspace trust dialog, so run it only in a directory you trust. In an interactive session, undo a bad fix with /rewind (or Esc Esc).
Codex runs commands inside its sandbox. For an interactive loop, codex --sandbox workspace-write lets it edit the workspace and run the check; the default approval policy, on-request, asks you before anything outside the sandbox. on-request does not stop a failing loop for you, so put the bound in the prompt:
Run npm run type-check. Fix the first error only, then re-run the check tosee which downstream errors disappeared. One root cause per iteration.Do not suppress errors. After 6 attempts, stop and report.For a headless run, codex exec takes the same prompt:
codex exec --sandbox workspace-write \ "Run npm run type-check and fix the errors one at a time, root cause first. Never suppress an error. Stop after 6 attempts and report."codex exec has no turn or spend limit flag (checked in 0.157.1), so the bound lives in the prompt; in CI, wrap the job in a timeout as well, for example timeout 15m codex exec … or the job’s timeout-minutes. Undo a bad fix with Git, and start a clean thread with /new when the loop goes in circles.
In Cursor’s Agent, the agent runs the check in its terminal, reads the output, and iterates. After a failing command you can ask it to fix the error shown in the terminal. When you paste an error yourself, add the context that makes it actionable:
I ran npm run type-check and got this error:
src/services/notification.ts:42:5 - error TS2345:Argument of type 'string' is not assignable to parameterof type 'NotificationPayload'.
Fix the root cause. The function in @src/services/notification.ts expectsa NotificationPayload object, not a raw string. Check the caller at line 42and the type in @src/types/notification.ts. Do not cast.Commit at every green state so a bad fix is one git restore away. Cursor suits the case where you want to watch the loop and step in the moment it drifts. To run the loop without anyone watching, hand it to a Cursor Cloud Agent, or trigger one from an Automation such as a Sentry alert or a schedule.
Three situations where the loop pays off
Section titled “Three situations where the loop pays off”Fix a production bug from a stack trace
Section titled “Fix a production bug from a stack trace”A user hit a crash and your error tracker captured it. Give the agent the full exception and stack, not the top line: the top frame is where the value was read, rarely where it went wrong.
The right fix addresses where total became undefined (for example, a cart created with no items), not an optional-chaining ?. added at line 212. The regression test is what makes the fix provable: it fails before the change and passes after it.
With the Sentry MCP server connected, the agent fetches the issue, stack trace, and breadcrumbs itself, and the prompt shrinks to “fetch Sentry issue PROJ-1234 and fix the root cause”. Sentry’s official server is remote at https://mcp.sentry.dev/mcp and signs in with OAuth:
claude mcp add --transport http sentry https://mcp.sentry.dev/mcpcodex mcp add sentry --url https://mcp.sentry.dev/mcpcodex mcp login sentryAdd the server to .cursor/mcp.json:
{ "mcpServers": { "sentry": { "url": "https://mcp.sentry.dev/mcp" } } }If you authenticate with a token instead of OAuth, the header is Authorization: Sentry-Bearer <token>, not Bearer. Debugging production from the editor covers Sentry alongside Grafana, Datadog, and PostHog.
Work through a compiler cascade after a refactor
Section titled “Work through a compiler cascade after a refactor”You changed the signature of a core function and the type-checker reports 30 errors across the codebase. This is the loop’s strongest case, because the errors form an exact, machine-generated worklist. Do not fix anything by hand first; let the full list appear, then hand it over.
The instruction to stop on call sites with no currency matters: it turns a guess the agent would otherwise make silently into a short list you decide on.
Drive a failing test to green
Section titled “Drive a failing test to green”The strongest form of the loop is a failing test written first, because the test is an oracle the agent can run but must not change. Test-driven development covers writing that test; the prompt below is the error-driven half.
The last sentence matters. Without it, an agent sometimes makes a failure pass by loosening the assertion. A prompt is a request, not a control, so protect the oracle with checks the agent cannot edit when the stakes are higher.
How do you prove the loop converged?
Section titled “How do you prove the loop converged?”A green re-run of the command that failed proves less than it seems. The agent may have suppressed the error, skipped a test, weakened an assertion, or fixed one check while breaking another. Run this gate before you accept the result. Each check is mechanical, so nobody has to read every changed line.
-
Run every check fresh, not only the one that failed. For example,
npm run type-check && npm run lint && npm test. A type fix that breaks a test elsewhere shows up here. -
Scan the diff for suppressions. Any hit needs a written reason or a revert:
Terminal window git diff main -U0 | grep -nE '^\+.*(@ts-ignore|@ts-expect-error|as any|eslint-disable|\.skip\(|\.only\(|catch( \([^)]*\))? *\{ *\})' -
Confirm the tests were not edited to pass. This command lists only removed or changed lines in test files and snapshots. It must print nothing: added tests are expected, removed or changed lines are not.
Terminal window git diff main -U0 -- '*.test.*' '*.spec.*' '*__snapshots__*' | grep -E '^-[^-]' -
Prove each regression test catches the bug. Put the old source back, keep the new test, and watch it fail. Commit the fix first, then run, for example:
git checkout main -- src/services/payment.ts && npx vitest run src/services/payment.test.ts; git checkout HEAD -- src/services/payment.ts. The test must fail here; the last command then restores the fixed file. A regression test that passes against the old code tests nothing.
The developer who ran the loop signs off on the root-cause explanation and on any hit from steps 2 and 3. CI running the same checks is the merge gate. To move steps 1 to 3 out of the prompt and into the harness, a Claude Code Stop hook that runs them and exits with code 2 on failure keeps the agent working instead of declaring victory. A hook that always fails keeps the agent going indefinitely, so have it check stop_hook_active (or cap its retries) and keep the --max-budget-usd bound on headless runs; see hooks as deterministic guardrails. Codex also has a Stop hook event.
How to break a cascade that will not converge
Section titled “How to break a cascade that will not converge”A compiler worklist shrinks as you work through it. The dangerous cascade does the opposite: fixing error 1 introduces error 2, fixing that introduces error 3, and the context fills with failed attempts. The sign is the same file appearing in every round. At that point, stop fixing errors and look at the structure.
Use /rewind to return to the last good state, or /clear to drop the polluted context, then restart with the analysis prompt below. After two failed correction cycles on the same error, a fresh context usually does better than a longer one, because every failed attempt stays in the window and steers the next one.
Reset the files with Git and start a clean thread with /new. Paste the analysis prompt below so the new thread starts from the structure, not from the last failed patch.
Restore the last green commit, open a new chat, and paste the analysis prompt below with @ references to the file, its types, and its test file. Ask for a proposal before any edit.
How to make sure the same error never returns
Section titled “How to make sure the same error never returns”The most valuable errors are the ones you never see again. When an error was hard to diagnose, capture the lesson in the strongest form available:
- A check that fails deterministically: a lint rule, a type, a test, or a CI step. This works for every agent and every person, whether or not they read instructions.
- A project instruction, when the lesson cannot be expressed as a check. Keep it short and specific.
Add the lesson to the project’s CLAUDE.md so the whole team gets it. Claude Code’s auto memory also stores learnings in ~/.claude/projects/<project>/memory/, but that directory is per user, not shared through the repository.
Add the lesson to AGENTS.md, which Codex reads before doing any work in a trusted project.
Add a project rule under .cursor/rules/ (Cursor’s Rules) that applies to every request.
The content is the same in all three files:
## Known error patterns
- Notification types: update src/types/notification.ts AND the Zod schema in src/validators/notification.ts together. They must stay in sync.- Integration tests need Redis: run `docker compose up -d redis` first.For how long-lived lessons are stored and pruned, see memory patterns.
When error-driven development goes wrong
Section titled “When error-driven development goes wrong”- The agent chases the wrong error. It fixes error 14 while error 1 causes it. Recovery: restate “fix only the first error, then re-run” and discard the out-of-order fix.
- The fix treats the symptom. An optional-chaining
?.stops the crash without explaining why the value was missing. Recovery: ask “where did this value become undefined?” and require a regression test that fails on the old code. - The agent suppresses instead of fixing.
@ts-ignore,as any,eslint-disable, and empty catch blocks all hide the error. Recovery: the suppression scan in the gate catches them; revert and re-prompt with the ban stated. - The test is edited to pass. Recovery: the test-file check in the gate catches it; restore the test from Git and make it read-only for the agent with checks it cannot edit.
- The loop runs against a flaky test. A non-deterministic failure lets the agent “fix”, see green by chance, and declare victory. Recovery: run the test several times first; if it flakes, quarantine it before starting the loop.
- The paste hides the culprit. Only the top stack frame was given. Recovery: paste the full trace, or fetch it through the Sentry MCP server, and name the suspect files.
- Hundreds of errors arrive at once. Recovery: work the first three to five, which are usually the root causes, then re-run and see how many remain.
- The error carries no information. A segfault or a bare “internal server error” is not a specification. Recovery: switch to hypothesis-driven debugging: reproduce, add logging, narrow by elimination. The
systematic-debuggingskill from obra/superpowers encodes that process (npx skills add obra/superpowers --skill systematic-debugging); read itsSKILL.mdbefore installing, because its directory also ships test fixtures. See testing and debugging skills. - The broken file is one the agent wrote from scratch. Recovery: delete it and regenerate from a tighter prompt that includes the error; that usually beats patching.