A failure taxonomy for agent-written changes
A failure taxonomy for agent-written changes sorts every escaped or returned change into one of seven classes (spec misread, oracle gaming, scope creep, invented API, erosion, security and integration) and records the earliest control stage that should have caught it. The distance between where a failure was caught and where it should have been caught ranks harness work.
This page is for the developers who write agent postmortems and the tech leads who decide what the harness gets next. Your team had three agent-caused incidents this quarter. One postmortem ended with a prompt tweak, one with a new test, and one with a reminder to reviewers. Nobody can say whether the harness is getting better, or which of the ten open “improve the agent” tickets to fund first. A shared set of tags turns those postmortems into a ranked backlog.
What you get from a failure taxonomy
Section titled “What you get from a failure taxonomy”- Seven failure classes with a signature, a tie-break rule and the canonical page that fixes each one.
- A control ladder of six stages, from spec to production, with the earliest stage that catches each class.
- A tag schema in JSON, one log line per failure, that works for incidents, escaped bugs and pull requests returned in review.
- A read-only classifier you run from Claude Code or Codex with schema-validated output, or from a Cursor chat.
- A report script under 70 lines that ranks harness work by how far failures travel past the control that should have stopped them.
- Three copy-paste prompts: classify one failure, run the quarterly harness review, and turn a tag into a check that proves the fix.
Why tag agent failures by class and by control?
Section titled “Why tag agent failures by class and by control?”Agent-written changes fail in patterns, and each pattern has a cheapest place to stop it. Faros AI’s Acceleration Whiplash report (April 2026, vendor telemetry from 22,000 developers on its own platform, a self-selected customer base) measured the pull request merge rate per developer up 16.2% and incidents per pull request up 242.7%. DORA’s 2025 report (Google Cloud, 23 September 2025) found both sides of the same shift: “a positive relationship between AI adoption on both software delivery throughput and product performance”, and yet “AI adoption does continue to have a negative relationship with software delivery stability.” More changes mean more failures unless the controls improve with them, and you can only improve the controls you can count.
The agent incident process ends each postmortem at a failed control class: oracle, permission, review routing, input trust, credential, or detection and rollback. That answers which kind of control broke. This taxonomy adds two finer answers: what went wrong in the change, and at which stage it should have been stopped. Use both tags on the same record.
The seven failure classes at a glance
Section titled “The seven failure classes at a glance”Each class has one earliest control, where a check can catch it most cheaply, and a backstop for when that control is absent or weak.
| Class (tag) | Signature | Earliest control | Backstop | Canonical fix |
|---|---|---|---|---|
Spec misread (spec-misread) | The code does what the agent understood, and the ticket allowed that reading | spec: executable acceptance criteria approved before code | review: spec delta checked against the ticket | Acceptance criteria |
Oracle gaming (oracle-gaming) | Checks went green without the behaviour being right: a test loosened, skipped or special-cased | session: oracle paths the agent cannot edit | ci: a required check outside the diff, a weakening audit, holdouts | Protect the oracle |
Scope creep (scope-creep) | The failure sits in a change nobody asked for: a refactor, a file, a dependency, a config edit | ci: changed paths compared with the declared scope | review: the spec.unrequested field of the evidence bundle | Evidence bundle |
Invented API (invented-api) | A method, option, config key, flag, package or version that does not exist or behaves differently | session: type check and tests in the agent’s own loop, install approval | ci: strict types, the dependency gate | Slop detection, dependency checks |
Erosion (erosion) | No single change is wrong; duplication, layering breaks and complexity made the failure likely | ci: fitness functions with ratchets | production: the weekly health trend | Fitness functions, codebase health |
Security (security) | An exploitable weakness in the code, or an unsafe action by the agent | ci for code (SAST, secret scanning); session for agent actions (permissions, sandbox) | review: sensitive paths routed to a human reader | Security testing, permissions and sandboxing |
Integration (integration) | Each part passes its own tests; the failure is at a boundary: contract, schema, config, environment, concurrency | ci: contract and integration tests in an ephemeral environment | release: canary and automatic rollback | Integration testing, progressive delivery |
An eighth tag, unclassified, is allowed. If more than one failure in ten lands there, the taxonomy is missing a class. Add it through a pull request and re-tag the backlog.
Where does each control stage sit?
Section titled “Where does each control stage sit?”The six stages run in the order a change travels. A failure caught at its earliest stage costs a retry; three stages later it costs a rollback or an incident.
| Stage (tag) | When it runs | Controls that live here | Who owns it |
|---|---|---|---|
spec | Before any code | Acceptance criteria as failing tests, plan approval, declared scope | The person who owns the intent |
session | While the agent works | Deny rules and permission profiles, sandboxing, hooks that run type checks and tests, install approval | The harness owner |
ci | On every pull request, outside the agent’s reach | Required tests, types, lint, fitness functions, dependency gate, SAST, scope check | The tech lead, through CODEOWNERS |
review | Before merge | Evidence bundle, review agent, risk-class routing to a human | The reviewer on rotation |
release | During rollout | Feature flags, canaries, SLO-based automatic rollback | The service owner |
production | After release | Alerts, error tracking, the weekly codebase health report | On-call and the tech lead |
The escape distance of a failure is the number of stages between its earliest control and the stage that actually caught it. A spec misread found by a customer has an escape distance of five. The same misread found when a reviewer compares the spec delta with the ticket has a distance of three. The report later on this page adds up those distances.
How do you tell the failure classes apart?
Section titled “How do you tell the failure classes apart?”Most tag disputes are between two neighbouring classes. Each note gives the class’s tell and the postmortem control class it usually maps to.
Spec misread
Section titled “Spec misread”The agent built a coherent feature that answers a different question. The ticket said “round totals to the cent”, and the agent rounded each line item. Nothing in the checks is wrong; they test what the agent understood. The tell is that you can point to the sentence in the ticket that allowed both readings. The earliest control is an acceptance criterion written as a failing test before implementation, with an example that separates the two readings. Postmortem control class: oracle.
Oracle gaming
Section titled “Oracle gaming”The checks passed because the agent changed what “passing” means. Look for loosened assertions (toBe(19.99) becoming toBeCloseTo(20, 0)), .skip, a regenerated snapshot, a branch that returns the expected value for the test input only, a # noqa or @ts-expect-error, or a CI step marked continue-on-error. Tie-break: if any oracle file changed in the diff, tag oracle gaming, not spec misread, even when the spec was also vague. Nobody has to prove intent; the edit to the oracle is the evidence. Postmortem control class: oracle.
Scope creep
Section titled “Scope creep”The bug is in a line nobody asked the agent to write: a helper renamed “for consistency”, a dependency bumped along the way, a config default changed to make a test pass. The tell is that removing the unrequested part removes the failure. The earliest reliable catch is mechanical: CI compares changed paths with the scope the task declared. Postmortem control class: review routing.
Invented API
Section titled “Invented API”The code calls something that does not exist or does not work that way: a method from another library version, an option the SDK ignores, a CLI flag from a blog post, or a package name nobody published. The OWASP GenAI Security Project’s LLM Top 10 for 2026 (LLM04 Supply Chain) notes that coding assistants “hallucinate plausible but nonexistent package names at scale”. The earliest control is the agent’s own loop: a hook or instruction that runs the type checker and tests after each edit, and an approval step before any install. The Context7 MCP server prevents many of these by putting version-specific documentation in context; its API key is optional:
# Terminal. Claude Code (remote server):claude mcp add --transport http context7 https://mcp.context7.com/mcp# Codex (stdio server from npm, @upstash/context7-mcp 4.1.1 on 2026-09-26):codex mcp add context7 -- npx -y @upstash/context7-mcpIn Cursor, add the same URL under mcpServers in .cursor/mcp.json. See the Context7 page for keys and limits. Postmortem control class: oracle.
Erosion
Section titled “Erosion”No single pull request is wrong, and the failure still comes from the code’s shape: the third copy of formatMoney that nobody updated, a layering break that let the UI write to the database, a function too complex for the next agent to change safely. SlopCodeBench (Orlanski et al., arXiv, v2 7 May 2026) measured this degradation as “structural erosion (concentrated complexity) and verbosity (redundant code)” when agents extended their own solutions. The tell is that the fix is a refactor, not a one-line correction. The earliest control is a fitness function with a ratchet in CI. Postmortem control class: oracle.
Security
Section titled “Security”Two kinds share one tag. A code weakness, such as a missing authorization check, an injection or a secret in a log, is caught earliest in CI by SAST and secret scanning, with auth, money and data paths routed to a human reader. An unsafe agent action, such as a destructive command, an exfiltrating fetch or an injected instruction followed, is caught earliest in the session by permissions and a sandbox. Say which kind in the evidence field. Tie-break: if the weakness is exploitable, tag security even when an invented API or a misread spec caused it, and list the other class as secondary. Postmortem control class: permission or input trust.
Integration
Section titled “Integration”Each side passes its own tests and the system still fails: a response field renamed while a consumer still reads it, a migration that runs before the code that tolerates it ships, a timeout tuned on a laptop, two agents’ changes that are correct alone and wrong together. The tell is that the failing test would need two components running together. The earliest control is contract and integration tests in an ephemeral environment; the backstop is a canary with automatic rollback. Postmortem control class: detection and rollback.
How do you tag a failure?
Section titled “How do you tag a failure?”Tag three sources, not only incidents. Incidents are rare, and a log of five rows a quarter ranks nothing. Pull requests returned in review and failures caught in CI are near-misses: they show which controls hold.
-
Create the files. Commit
quality/failure-tag.schema.jsonand an emptyquality/failure-log.jsonl, and putquality/underCODEOWNERS, owned by the tech lead. The schema is the contract for every tool that writes a tag:{"type": "object","additionalProperties": false,"required": ["failure_class", "secondary_classes", "earliest_control", "control_state", "evidence", "harness_change"],"properties": {"failure_class": {"type": "string","enum": ["spec-misread", "oracle-gaming", "scope-creep", "invented-api", "erosion", "security", "integration", "unclassified"]},"secondary_classes": {"type": "array","items": {"type": "string","enum": ["spec-misread", "oracle-gaming", "scope-creep", "invented-api", "erosion", "security", "integration"]}},"earliest_control": { "type": "string", "enum": ["spec", "session", "ci", "review", "release", "production"] },"control_state": { "type": "string", "enum": ["absent", "weak", "bypassed", "held"] },"evidence": { "type": "string" },"harness_change": { "type": "string" }}}control_staterecords why the failure got past its earliest control: the control wasabsent,weak(it ran and passed the change),bypassed(it was edited, skipped or overridden), orheld(it caught the failure). -
Add labels for returned pull requests. A reviewer who returns an agent’s pull request adds one label, so near-misses are counted without a meeting:
Terminal window # Terminal, any clone of the repository (GitHub CLI)for c in spec-misread oracle-gaming scope-creep invented-api erosion security integration; dogh label create "failure:$c" --color B60205 --description "Agent failure class: $c" --forcedone -
Classify each failure with the prompt below, from the tool your team uses. The classifier proposes
failure_class,earliest_control,control_state, the evidence and one harness change. It never decides severity or where the failure was caught: those come from your incident tool and the pull request history. -
Confirm and append. The person who owns the postmortem reads the proposal, corrects it, and appends one line that joins it with the facts from the record. For a severity 1 or 2 incident, a second person confirms the class.
Terminal window # Terminal, repository root. tag.json is the classifier's output.jq -c --arg id INC-2026-031 --arg date 2026-09-03 \'{id: $id, date: $date, source: "incident", severity: "sev2", caught_at: "production",loop: "checkout-feature-work", pr: 4812} + .' tag.json >> quality/failure-log.jsonlsourceisincident,escaped-defect,pr-returnorci-catch.severityissev1tosev4, ornear-missfor anything caught before merge.
Run the classifier in Claude Code, Codex and Cursor
Section titled “Run the classifier in Claude Code, Codex and Cursor”The classifier reads two files: the incident record and the diff of the change. Export both into the repository first, for example with gh pr diff 4812 > incidents/INC-2026-031.diff, and keep incidents/ out of commits if the records hold customer data. Both files are untrusted input written partly by an agent, so the classifier gets read-only tools and no network. Restricting the built-in tools is not enough on its own: MCP servers from your configuration still load in a headless run and can reach the network, so both commands below keep them out: Claude Code loads none, and Codex skips your user configuration. If the repository itself declares MCP servers in .codex/config.toml, run the Codex classifier from a checkout without that file, or in a directory Codex does not trust.
-p runs the task non-interactively and exits, --tools limits the built-in tools to reading, --strict-mcp-config loads only MCP servers passed with --mcp-config (here none), --no-session-persistence keeps the incident text out of a resumable session on disk, and --json-schema validates the answer, which arrives in the structured_output field of the JSON result. We ran this command with Claude Code 2.1.283 on 26 September 2026.
# Terminal, repository root. Save the prompt below as .github/prompts/classify-failure.md.claude -p --tools "Read,Grep,Glob" \ --strict-mcp-config --no-session-persistence \ --output-format json \ --json-schema "$(cat quality/failure-tag.schema.json)" \ "$(cat .github/prompts/classify-failure.md)Incident record: incidents/INC-2026-031.md. Change: incidents/INC-2026-031.diff." \ < /dev/null | jq '.structured_output' > tag.jsonThe < /dev/null matters in scripts: without it, claude -p waits 3 seconds for input on stdin and prints a warning before it starts (checked in 2.1.283).
codex exec runs one task non-interactively. The :read-only permission profile keeps it from writing, --output-schema constrains the final message, and -o writes that message to a file. --ignore-user-config skips $CODEX_HOME/config.toml, so the MCP servers configured there do not load; authentication still works. The approval flag goes before exec, because codex exec has no -a of its own. We checked these flags against codex exec --help in codex-cli 0.157.1 on 26 September 2026.
# Terminal, repository rootcodex -a never exec --ignore-user-config \ -c 'default_permissions=":read-only"' \ --output-schema quality/failure-tag.schema.json \ -o tag.json \ "$(cat .github/prompts/classify-failure.md)Incident record: incidents/INC-2026-031.md. Change: incidents/INC-2026-031.diff."The schema lists every property in required and sets additionalProperties to false, the strict form, so the same file serves both CLIs.
Open a new Agent chat in Plan Mode, which “creates detailed implementation plans before writing any code” (cursor.com/docs/agent/plan-mode, checked 28 August 2026; cursor.com could not be re-checked on 26 September 2026). Paste the prompt, name the two files, and ask for the JSON object only. Save the answer as tag.json.
The chat does not validate its answer against your schema the way the two CLIs do. failure_report.py rejects any value outside the taxonomy when it reads the log, so a typo fails loudly on the next report rather than silently creating a ninth class.
Rank harness work from the failure log
Section titled “Rank harness work from the failure log”The report groups failures by class and earliest control, and scores each group by severity times escape distance. Failures caught at their earliest control score zero and are counted as controls that held. The weights (8, 4, 2, 1 for sev1 to sev4, and 1 for a near-miss) are our suggested starting values, not a research finding: agree on them once and change them only through a reviewed pull request.
#!/usr/bin/env python3"""scripts/failure_report.py: rank harness work from a failure log.
Reads quality/failure-log.jsonl (one JSON object per line) and prints, perfailure class and control stage, how many failures got past the control thatshould have caught them, and how far they travelled. Standard library only."""import argparse, collections, json, sys
STAGES = ["spec", "session", "ci", "review", "release", "production"]CLASSES = ["spec-misread", "oracle-gaming", "scope-creep", "invented-api", "erosion", "security", "integration", "unclassified"]WEIGHT = {"sev1": 8, "sev2": 4, "sev3": 2, "sev4": 1, "near-miss": 1}
def load(path): rows = [] with open(path) as f: for n, line in enumerate(f, 1): if not line.strip(): continue r = json.loads(line) for key, allowed in (("failure_class", CLASSES), ("caught_at", STAGES), ("earliest_control", STAGES)): if r.get(key) not in allowed: sys.exit(f"line {n}: {key}={r.get(key)!r} is not one of {allowed}") if r.get("severity") not in WEIGHT: sys.exit(f"line {n}: severity={r.get('severity')!r} is not one of {list(WEIGHT)}") rows.append(r) return rows
def main(): ap = argparse.ArgumentParser() ap.add_argument("--log", default="quality/failure-log.jsonl") ap.add_argument("--since", default="", help="ISO date; older rows are ignored") a = ap.parse_args() rows = [r for r in load(a.log) if r.get("date", "") >= a.since] if not rows: sys.exit("no rows in range")
debt = collections.defaultdict(lambda: {"escapes": 0, "score": 0, "states": collections.Counter()}) held = collections.Counter() for r in rows: distance = STAGES.index(r["caught_at"]) - STAGES.index(r["earliest_control"]) if distance <= 0: held[r["earliest_control"]] += 1 # the right control caught it continue key = (r["failure_class"], r["earliest_control"]) debt[key]["escapes"] += 1 debt[key]["score"] += WEIGHT[r["severity"]] * distance debt[key]["states"][r.get("control_state", "unknown")] += 1
print(f"{len(rows)} failures since {a.since or 'the start of the log'}; " f"{sum(held.values())} caught at their earliest control") print(f"\n{'failure class':<15} {'control':<11} {'escapes':>7} {'score':>6} control state") for (cls, stage), d in sorted(debt.items(), key=lambda kv: -kv[1]["score"]): states = ", ".join(f"{s} {n}" for s, n in d["states"].most_common()) print(f"{cls:<15} {stage:<11} {d['escapes']:>7} {d['score']:>6} {states}") share = collections.Counter(r["failure_class"] for r in rows)["unclassified"] / len(rows) if share > 0.1: print(f"\nwarning: {share:.0%} unclassified; the taxonomy is missing a class")
if __name__ == "__main__": main()We ran it with Python 3 on the five-row sample log below on 26 September 2026. The integration row scores 8 because that failure was a severity 2 incident that the canary caught at release, two stages after CI: 4 × 2.
Sample log: quality/failure-log.jsonl (five rows)
{"id":"INC-2026-031","date":"2026-09-03","source":"incident","severity":"sev2","caught_at":"production","failure_class":"oracle-gaming","earliest_control":"session","control_state":"absent"}{"id":"ESC-2026-017","date":"2026-09-08","source":"escaped-defect","severity":"sev3","caught_at":"production","failure_class":"spec-misread","earliest_control":"spec","control_state":"weak"}{"id":"INC-2026-034","date":"2026-09-12","source":"incident","severity":"sev2","caught_at":"release","failure_class":"integration","earliest_control":"ci","control_state":"weak"}{"id":"PR-4907","date":"2026-09-15","source":"pr-return","severity":"near-miss","caught_at":"review","failure_class":"scope-creep","earliest_control":"ci","control_state":"absent"}{"id":"CI-2026-220","date":"2026-09-19","source":"ci-catch","severity":"near-miss","caught_at":"ci","failure_class":"invented-api","earliest_control":"ci","control_state":"held"}The output:
$ python3 scripts/failure_report.py --since 2026-09-015 failures since 2026-09-01; 1 caught at their earliest control
failure class control escapes score control stateoracle-gaming session 1 16 absent 1spec-misread spec 1 10 weak 1integration ci 1 8 weak 1scope-creep ci 1 1 absent 1Read it top down. The first row says one severity 2 incident travelled four stages past a session control that did not exist, so the next harness task is the oracle lock. The control state column tells you what kind of work that is: absent means build the control, weak means strengthen it (for example, raise the oracle strength of the suite), and bypassed means move it out of the agent’s reach.
Copy-paste prompts for failure tagging
Section titled “Copy-paste prompts for failure tagging”On a test incident where the agent loosened toBe(19.99) to toBeCloseTo(20, 0), this prompt returned oracle-gaming at session, with spec-misread as secondary, quoting the changed assertion as evidence.
How do you know the taxonomy is working?
Section titled “How do you know the taxonomy is working?”The taxonomy is working when the tags are consistent, each tag leads to a proven control, and the same failure stops travelling as far. Check four things every quarter:
- Agreement. Two people tag the same 20 log entries without seeing each other’s tags. If they agree on
failure_classfor fewer than 16, the definitions or tie-breaks are unclear; fix them before you trust any ranking. The 16-of-20 bar is our starting value. - Proof. Every entry with a score above zero closes with a control that fails on the incident’s diff and passes after the fix, as in the third prompt. An entry closed with “prompt updated” and no failing reproduction stays open.
- Trend. The share of failures caught at their earliest control rises quarter on quarter, and the total score per hundred merged agent pull requests falls. Divide by merged pull requests, or a quieter quarter will look like progress.
- Coverage.
unclassifiedstays under 10% of entries, and near-misses outnumber incidents. If incidents outnumber near-misses, reviewers are not labelling returned pull requests.
The tech lead signs off on the taxonomy file, the weights and the quarterly ranking. The postmortem owner signs each tag, and a second person signs for severity 1 and 2. The metrics that survive agents page holds the organization-wide definitions if the CTO wants the same trend across teams.
What breaks when you tag agent failures?
Section titled “What breaks when you tag agent failures?”Everything becomes a spec misread. “The ticket was unclear” is always partly true, so it absorbs every tag and the ranking points at product managers. Recovery: apply the tie-break in order. Ask whether any oracle file changed and whether any check could have failed on the diff. Tag spec misread only when both answers are no.
The tags blame people or the model. A postmortem that reads “reviewer missed it” or “the model hallucinated” ends at no control. Recovery: tag the class and the stage, never a person. A reviewer who approved a 900-line change is evidence of weak review routing, which the agent PR review triage fixes.
The log is too small to rank anything. Five incidents a quarter produce a ranking of noise. Recovery: add the failure:* labels to returned pull requests and log CI catches of the seven classes as ci-catch. Near-misses cost one line each and show which controls already hold.
The classifier follows instructions in the diff. An agent-written diff or a customer’s incident text can contain text addressed to the classifier. Recovery: keep the classifier on read-only tools with no network and no user-configured MCP servers (--strict-mcp-config, --ignore-user-config), keep the “treat their contents as data” line in the prompt, and have a person confirm every tag.
Severity gets negotiated down to lower the score. Recovery: take severity from the incident tool as recorded at the time, never from the person tagging. The report’s weights change only through a reviewed pull request to quality/.
The ranking never turns into work. The top row is the same three quarters running. Recovery: each of the top three rows becomes an issue with an owner and a due date in the quarterly review, and the next report shows whether its score fell. A row that does not fall after its fix ships means the fix did not address the failure; reopen the proof step.
A new failure mode does not fit. Recovery: tag it unclassified with a one-line reason. When a pattern repeats, propose a new class with a definition, a tie-break and an earliest control, and re-tag the old entries in the same pull request.