Skip to content

A failure taxonomy for agent-written changes

A failure taxonomy for agent-written changes sorts every escaped or returned change into one of seven classes (spec misread, oracle gaming, scope creep, invented API, erosion, security and integration) and records the earliest control stage that should have caught it. The distance between where a failure was caught and where it should have been caught ranks harness work.

This page is for the developers who write agent postmortems and the tech leads who decide what the harness gets next. Your team had three agent-caused incidents this quarter. One postmortem ended with a prompt tweak, one with a new test, and one with a reminder to reviewers. Nobody can say whether the harness is getting better, or which of the ten open “improve the agent” tickets to fund first. A shared set of tags turns those postmortems into a ranked backlog.

  • Seven failure classes with a signature, a tie-break rule and the canonical page that fixes each one.
  • A control ladder of six stages, from spec to production, with the earliest stage that catches each class.
  • A tag schema in JSON, one log line per failure, that works for incidents, escaped bugs and pull requests returned in review.
  • A read-only classifier you run from Claude Code or Codex with schema-validated output, or from a Cursor chat.
  • A report script under 70 lines that ranks harness work by how far failures travel past the control that should have stopped them.
  • Three copy-paste prompts: classify one failure, run the quarterly harness review, and turn a tag into a check that proves the fix.

Why tag agent failures by class and by control?

Section titled “Why tag agent failures by class and by control?”

Agent-written changes fail in patterns, and each pattern has a cheapest place to stop it. Faros AI’s Acceleration Whiplash report (April 2026, vendor telemetry from 22,000 developers on its own platform, a self-selected customer base) measured the pull request merge rate per developer up 16.2% and incidents per pull request up 242.7%. DORA’s 2025 report (Google Cloud, 23 September 2025) found both sides of the same shift: “a positive relationship between AI adoption on both software delivery throughput and product performance”, and yet “AI adoption does continue to have a negative relationship with software delivery stability.” More changes mean more failures unless the controls improve with them, and you can only improve the controls you can count.

The agent incident process ends each postmortem at a failed control class: oracle, permission, review routing, input trust, credential, or detection and rollback. That answers which kind of control broke. This taxonomy adds two finer answers: what went wrong in the change, and at which stage it should have been stopped. Use both tags on the same record.

Each class has one earliest control, where a check can catch it most cheaply, and a backstop for when that control is absent or weak.

Class (tag)SignatureEarliest controlBackstopCanonical fix
Spec misread (spec-misread)The code does what the agent understood, and the ticket allowed that readingspec: executable acceptance criteria approved before codereview: spec delta checked against the ticketAcceptance criteria
Oracle gaming (oracle-gaming)Checks went green without the behaviour being right: a test loosened, skipped or special-casedsession: oracle paths the agent cannot editci: a required check outside the diff, a weakening audit, holdoutsProtect the oracle
Scope creep (scope-creep)The failure sits in a change nobody asked for: a refactor, a file, a dependency, a config editci: changed paths compared with the declared scopereview: the spec.unrequested field of the evidence bundleEvidence bundle
Invented API (invented-api)A method, option, config key, flag, package or version that does not exist or behaves differentlysession: type check and tests in the agent’s own loop, install approvalci: strict types, the dependency gateSlop detection, dependency checks
Erosion (erosion)No single change is wrong; duplication, layering breaks and complexity made the failure likelyci: fitness functions with ratchetsproduction: the weekly health trendFitness functions, codebase health
Security (security)An exploitable weakness in the code, or an unsafe action by the agentci for code (SAST, secret scanning); session for agent actions (permissions, sandbox)review: sensitive paths routed to a human readerSecurity testing, permissions and sandboxing
Integration (integration)Each part passes its own tests; the failure is at a boundary: contract, schema, config, environment, concurrencyci: contract and integration tests in an ephemeral environmentrelease: canary and automatic rollbackIntegration testing, progressive delivery

An eighth tag, unclassified, is allowed. If more than one failure in ten lands there, the taxonomy is missing a class. Add it through a pull request and re-tag the backlog.

The six stages run in the order a change travels. A failure caught at its earliest stage costs a retry; three stages later it costs a rollback or an incident.

Stage (tag)When it runsControls that live hereWho owns it
specBefore any codeAcceptance criteria as failing tests, plan approval, declared scopeThe person who owns the intent
sessionWhile the agent worksDeny rules and permission profiles, sandboxing, hooks that run type checks and tests, install approvalThe harness owner
ciOn every pull request, outside the agent’s reachRequired tests, types, lint, fitness functions, dependency gate, SAST, scope checkThe tech lead, through CODEOWNERS
reviewBefore mergeEvidence bundle, review agent, risk-class routing to a humanThe reviewer on rotation
releaseDuring rolloutFeature flags, canaries, SLO-based automatic rollbackThe service owner
productionAfter releaseAlerts, error tracking, the weekly codebase health reportOn-call and the tech lead

The escape distance of a failure is the number of stages between its earliest control and the stage that actually caught it. A spec misread found by a customer has an escape distance of five. The same misread found when a reviewer compares the spec delta with the ticket has a distance of three. The report later on this page adds up those distances.

How do you tell the failure classes apart?

Section titled “How do you tell the failure classes apart?”

Most tag disputes are between two neighbouring classes. Each note gives the class’s tell and the postmortem control class it usually maps to.

The agent built a coherent feature that answers a different question. The ticket said “round totals to the cent”, and the agent rounded each line item. Nothing in the checks is wrong; they test what the agent understood. The tell is that you can point to the sentence in the ticket that allowed both readings. The earliest control is an acceptance criterion written as a failing test before implementation, with an example that separates the two readings. Postmortem control class: oracle.

The checks passed because the agent changed what “passing” means. Look for loosened assertions (toBe(19.99) becoming toBeCloseTo(20, 0)), .skip, a regenerated snapshot, a branch that returns the expected value for the test input only, a # noqa or @ts-expect-error, or a CI step marked continue-on-error. Tie-break: if any oracle file changed in the diff, tag oracle gaming, not spec misread, even when the spec was also vague. Nobody has to prove intent; the edit to the oracle is the evidence. Postmortem control class: oracle.

The bug is in a line nobody asked the agent to write: a helper renamed “for consistency”, a dependency bumped along the way, a config default changed to make a test pass. The tell is that removing the unrequested part removes the failure. The earliest reliable catch is mechanical: CI compares changed paths with the scope the task declared. Postmortem control class: review routing.

The code calls something that does not exist or does not work that way: a method from another library version, an option the SDK ignores, a CLI flag from a blog post, or a package name nobody published. The OWASP GenAI Security Project’s LLM Top 10 for 2026 (LLM04 Supply Chain) notes that coding assistants “hallucinate plausible but nonexistent package names at scale”. The earliest control is the agent’s own loop: a hook or instruction that runs the type checker and tests after each edit, and an approval step before any install. The Context7 MCP server prevents many of these by putting version-specific documentation in context; its API key is optional:

Terminal window
# Terminal. Claude Code (remote server):
claude mcp add --transport http context7 https://mcp.context7.com/mcp
# Codex (stdio server from npm, @upstash/context7-mcp 4.1.1 on 2026-09-26):
codex mcp add context7 -- npx -y @upstash/context7-mcp

In Cursor, add the same URL under mcpServers in .cursor/mcp.json. See the Context7 page for keys and limits. Postmortem control class: oracle.

No single pull request is wrong, and the failure still comes from the code’s shape: the third copy of formatMoney that nobody updated, a layering break that let the UI write to the database, a function too complex for the next agent to change safely. SlopCodeBench (Orlanski et al., arXiv, v2 7 May 2026) measured this degradation as “structural erosion (concentrated complexity) and verbosity (redundant code)” when agents extended their own solutions. The tell is that the fix is a refactor, not a one-line correction. The earliest control is a fitness function with a ratchet in CI. Postmortem control class: oracle.

Two kinds share one tag. A code weakness, such as a missing authorization check, an injection or a secret in a log, is caught earliest in CI by SAST and secret scanning, with auth, money and data paths routed to a human reader. An unsafe agent action, such as a destructive command, an exfiltrating fetch or an injected instruction followed, is caught earliest in the session by permissions and a sandbox. Say which kind in the evidence field. Tie-break: if the weakness is exploitable, tag security even when an invented API or a misread spec caused it, and list the other class as secondary. Postmortem control class: permission or input trust.

Each side passes its own tests and the system still fails: a response field renamed while a consumer still reads it, a migration that runs before the code that tolerates it ships, a timeout tuned on a laptop, two agents’ changes that are correct alone and wrong together. The tell is that the failing test would need two components running together. The earliest control is contract and integration tests in an ephemeral environment; the backstop is a canary with automatic rollback. Postmortem control class: detection and rollback.

Tag three sources, not only incidents. Incidents are rare, and a log of five rows a quarter ranks nothing. Pull requests returned in review and failures caught in CI are near-misses: they show which controls hold.

  1. Create the files. Commit quality/failure-tag.schema.json and an empty quality/failure-log.jsonl, and put quality/ under CODEOWNERS, owned by the tech lead. The schema is the contract for every tool that writes a tag:

    {
    "type": "object",
    "additionalProperties": false,
    "required": ["failure_class", "secondary_classes", "earliest_control", "control_state", "evidence", "harness_change"],
    "properties": {
    "failure_class": {
    "type": "string",
    "enum": ["spec-misread", "oracle-gaming", "scope-creep", "invented-api", "erosion", "security", "integration", "unclassified"]
    },
    "secondary_classes": {
    "type": "array",
    "items": {
    "type": "string",
    "enum": ["spec-misread", "oracle-gaming", "scope-creep", "invented-api", "erosion", "security", "integration"]
    }
    },
    "earliest_control": { "type": "string", "enum": ["spec", "session", "ci", "review", "release", "production"] },
    "control_state": { "type": "string", "enum": ["absent", "weak", "bypassed", "held"] },
    "evidence": { "type": "string" },
    "harness_change": { "type": "string" }
    }
    }

    control_state records why the failure got past its earliest control: the control was absent, weak (it ran and passed the change), bypassed (it was edited, skipped or overridden), or held (it caught the failure).

  2. Add labels for returned pull requests. A reviewer who returns an agent’s pull request adds one label, so near-misses are counted without a meeting:

    Terminal window
    # Terminal, any clone of the repository (GitHub CLI)
    for c in spec-misread oracle-gaming scope-creep invented-api erosion security integration; do
    gh label create "failure:$c" --color B60205 --description "Agent failure class: $c" --force
    done
  3. Classify each failure with the prompt below, from the tool your team uses. The classifier proposes failure_class, earliest_control, control_state, the evidence and one harness change. It never decides severity or where the failure was caught: those come from your incident tool and the pull request history.

  4. Confirm and append. The person who owns the postmortem reads the proposal, corrects it, and appends one line that joins it with the facts from the record. For a severity 1 or 2 incident, a second person confirms the class.

    Terminal window
    # Terminal, repository root. tag.json is the classifier's output.
    jq -c --arg id INC-2026-031 --arg date 2026-09-03 \
    '{id: $id, date: $date, source: "incident", severity: "sev2", caught_at: "production",
    loop: "checkout-feature-work", pr: 4812} + .' tag.json >> quality/failure-log.jsonl

    source is incident, escaped-defect, pr-return or ci-catch. severity is sev1 to sev4, or near-miss for anything caught before merge.

Run the classifier in Claude Code, Codex and Cursor

Section titled “Run the classifier in Claude Code, Codex and Cursor”

The classifier reads two files: the incident record and the diff of the change. Export both into the repository first, for example with gh pr diff 4812 > incidents/INC-2026-031.diff, and keep incidents/ out of commits if the records hold customer data. Both files are untrusted input written partly by an agent, so the classifier gets read-only tools and no network. Restricting the built-in tools is not enough on its own: MCP servers from your configuration still load in a headless run and can reach the network, so both commands below keep them out: Claude Code loads none, and Codex skips your user configuration. If the repository itself declares MCP servers in .codex/config.toml, run the Codex classifier from a checkout without that file, or in a directory Codex does not trust.

-p runs the task non-interactively and exits, --tools limits the built-in tools to reading, --strict-mcp-config loads only MCP servers passed with --mcp-config (here none), --no-session-persistence keeps the incident text out of a resumable session on disk, and --json-schema validates the answer, which arrives in the structured_output field of the JSON result. We ran this command with Claude Code 2.1.283 on 26 September 2026.

Terminal window
# Terminal, repository root. Save the prompt below as .github/prompts/classify-failure.md.
claude -p --tools "Read,Grep,Glob" \
--strict-mcp-config --no-session-persistence \
--output-format json \
--json-schema "$(cat quality/failure-tag.schema.json)" \
"$(cat .github/prompts/classify-failure.md)
Incident record: incidents/INC-2026-031.md. Change: incidents/INC-2026-031.diff." \
< /dev/null | jq '.structured_output' > tag.json

The < /dev/null matters in scripts: without it, claude -p waits 3 seconds for input on stdin and prints a warning before it starts (checked in 2.1.283).

The report groups failures by class and earliest control, and scores each group by severity times escape distance. Failures caught at their earliest control score zero and are counted as controls that held. The weights (8, 4, 2, 1 for sev1 to sev4, and 1 for a near-miss) are our suggested starting values, not a research finding: agree on them once and change them only through a reviewed pull request.

#!/usr/bin/env python3
"""scripts/failure_report.py: rank harness work from a failure log.
Reads quality/failure-log.jsonl (one JSON object per line) and prints, per
failure class and control stage, how many failures got past the control that
should have caught them, and how far they travelled. Standard library only.
"""
import argparse, collections, json, sys
STAGES = ["spec", "session", "ci", "review", "release", "production"]
CLASSES = ["spec-misread", "oracle-gaming", "scope-creep", "invented-api",
"erosion", "security", "integration", "unclassified"]
WEIGHT = {"sev1": 8, "sev2": 4, "sev3": 2, "sev4": 1, "near-miss": 1}
def load(path):
rows = []
with open(path) as f:
for n, line in enumerate(f, 1):
if not line.strip():
continue
r = json.loads(line)
for key, allowed in (("failure_class", CLASSES), ("caught_at", STAGES),
("earliest_control", STAGES)):
if r.get(key) not in allowed:
sys.exit(f"line {n}: {key}={r.get(key)!r} is not one of {allowed}")
if r.get("severity") not in WEIGHT:
sys.exit(f"line {n}: severity={r.get('severity')!r} is not one of {list(WEIGHT)}")
rows.append(r)
return rows
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--log", default="quality/failure-log.jsonl")
ap.add_argument("--since", default="", help="ISO date; older rows are ignored")
a = ap.parse_args()
rows = [r for r in load(a.log) if r.get("date", "") >= a.since]
if not rows:
sys.exit("no rows in range")
debt = collections.defaultdict(lambda: {"escapes": 0, "score": 0, "states": collections.Counter()})
held = collections.Counter()
for r in rows:
distance = STAGES.index(r["caught_at"]) - STAGES.index(r["earliest_control"])
if distance <= 0:
held[r["earliest_control"]] += 1 # the right control caught it
continue
key = (r["failure_class"], r["earliest_control"])
debt[key]["escapes"] += 1
debt[key]["score"] += WEIGHT[r["severity"]] * distance
debt[key]["states"][r.get("control_state", "unknown")] += 1
print(f"{len(rows)} failures since {a.since or 'the start of the log'}; "
f"{sum(held.values())} caught at their earliest control")
print(f"\n{'failure class':<15} {'control':<11} {'escapes':>7} {'score':>6} control state")
for (cls, stage), d in sorted(debt.items(), key=lambda kv: -kv[1]["score"]):
states = ", ".join(f"{s} {n}" for s, n in d["states"].most_common())
print(f"{cls:<15} {stage:<11} {d['escapes']:>7} {d['score']:>6} {states}")
share = collections.Counter(r["failure_class"] for r in rows)["unclassified"] / len(rows)
if share > 0.1:
print(f"\nwarning: {share:.0%} unclassified; the taxonomy is missing a class")
if __name__ == "__main__":
main()

We ran it with Python 3 on the five-row sample log below on 26 September 2026. The integration row scores 8 because that failure was a severity 2 incident that the canary caught at release, two stages after CI: 4 × 2.

Sample log: quality/failure-log.jsonl (five rows)
{"id":"INC-2026-031","date":"2026-09-03","source":"incident","severity":"sev2","caught_at":"production","failure_class":"oracle-gaming","earliest_control":"session","control_state":"absent"}
{"id":"ESC-2026-017","date":"2026-09-08","source":"escaped-defect","severity":"sev3","caught_at":"production","failure_class":"spec-misread","earliest_control":"spec","control_state":"weak"}
{"id":"INC-2026-034","date":"2026-09-12","source":"incident","severity":"sev2","caught_at":"release","failure_class":"integration","earliest_control":"ci","control_state":"weak"}
{"id":"PR-4907","date":"2026-09-15","source":"pr-return","severity":"near-miss","caught_at":"review","failure_class":"scope-creep","earliest_control":"ci","control_state":"absent"}
{"id":"CI-2026-220","date":"2026-09-19","source":"ci-catch","severity":"near-miss","caught_at":"ci","failure_class":"invented-api","earliest_control":"ci","control_state":"held"}

The output:

$ python3 scripts/failure_report.py --since 2026-09-01
5 failures since 2026-09-01; 1 caught at their earliest control
failure class control escapes score control state
oracle-gaming session 1 16 absent 1
spec-misread spec 1 10 weak 1
integration ci 1 8 weak 1
scope-creep ci 1 1 absent 1

Read it top down. The first row says one severity 2 incident travelled four stages past a session control that did not exist, so the next harness task is the oracle lock. The control state column tells you what kind of work that is: absent means build the control, weak means strengthen it (for example, raise the oracle strength of the suite), and bypassed means move it out of the agent’s reach.

On a test incident where the agent loosened toBe(19.99) to toBeCloseTo(20, 0), this prompt returned oracle-gaming at session, with spec-misread as secondary, quoting the changed assertion as evidence.

The taxonomy is working when the tags are consistent, each tag leads to a proven control, and the same failure stops travelling as far. Check four things every quarter:

  1. Agreement. Two people tag the same 20 log entries without seeing each other’s tags. If they agree on failure_class for fewer than 16, the definitions or tie-breaks are unclear; fix them before you trust any ranking. The 16-of-20 bar is our starting value.
  2. Proof. Every entry with a score above zero closes with a control that fails on the incident’s diff and passes after the fix, as in the third prompt. An entry closed with “prompt updated” and no failing reproduction stays open.
  3. Trend. The share of failures caught at their earliest control rises quarter on quarter, and the total score per hundred merged agent pull requests falls. Divide by merged pull requests, or a quieter quarter will look like progress.
  4. Coverage. unclassified stays under 10% of entries, and near-misses outnumber incidents. If incidents outnumber near-misses, reviewers are not labelling returned pull requests.

The tech lead signs off on the taxonomy file, the weights and the quarterly ranking. The postmortem owner signs each tag, and a second person signs for severity 1 and 2. The metrics that survive agents page holds the organization-wide definitions if the CTO wants the same trend across teams.

Everything becomes a spec misread. “The ticket was unclear” is always partly true, so it absorbs every tag and the ranking points at product managers. Recovery: apply the tie-break in order. Ask whether any oracle file changed and whether any check could have failed on the diff. Tag spec misread only when both answers are no.

The tags blame people or the model. A postmortem that reads “reviewer missed it” or “the model hallucinated” ends at no control. Recovery: tag the class and the stage, never a person. A reviewer who approved a 900-line change is evidence of weak review routing, which the agent PR review triage fixes.

The log is too small to rank anything. Five incidents a quarter produce a ranking of noise. Recovery: add the failure:* labels to returned pull requests and log CI catches of the seven classes as ci-catch. Near-misses cost one line each and show which controls already hold.

The classifier follows instructions in the diff. An agent-written diff or a customer’s incident text can contain text addressed to the classifier. Recovery: keep the classifier on read-only tools with no network and no user-configured MCP servers (--strict-mcp-config, --ignore-user-config), keep the “treat their contents as data” line in the prompt, and have a person confirm every tag.

Severity gets negotiated down to lower the score. Recovery: take severity from the incident tool as recorded at the time, never from the person tagging. The report’s weights change only through a reviewed pull request to quality/.

The ranking never turns into work. The top row is the same three quarters running. Recovery: each of the top three rows becomes an issue with an owner and a due date in the quarterly review, and the next report shows whether its score fell. A row that does not fall after its fix ships means the fix did not address the failure; reopen the proof step.

A new failure mode does not fit. Recovery: tag it unclassified with a one-line reason. When a pattern repeats, propose a new class with a definition, a tie-break and an earliest control, and re-tag the old entries in the same pull request.