Catching plausible-but-wrong code before merge
Slop detection catches agent-written code that compiles, passes its tests and reads well but is wrong: swallowed errors, dead abstractions, weakened tests, duplicate helpers and invented APIs. Each pattern gets a deterministic detector that blocks the merge, plus a review-agent prompt for the cases the detector cannot see, so nobody has to read every line.
This page is for developers who ship agent pull requests and for the tech lead who answers the scorecard question “How do you catch plausible-but-wrong code before merge?”. The situation it solves: the refund job has been “succeeding” for a week, and nobody noticed that a catch {} the agent added turned every failed refund into a silent false. The diff looked tidy. The tests passed because they only checked the happy path. The reviewer skimmed 600 lines and approved.
What you’ll walk away with from slop detection
Section titled “What you’ll walk away with from slop detection”- A catalog of five slop patterns, each mapped to a detector for TypeScript, Python and Go, and to what that detector misses.
- Lint and CI configuration tested on 26 September 2026 with the then-current releases: ESLint 10.11.0 with typescript-eslint 8.70.1 and @vitest/eslint-plugin 1.6.27, Ruff 0.16.9, mypy 2.3.1, pyright 1.1.414, knip 6.38.0, vulture 2.16 and jscpd 5.3.2. The Go rules were checked against the golangci-lint 2.14.0 source.
- A diff audit script that flags new suppressions, error masking and skipped tests on the lines a pull request adds.
- Five review-agent prompts, one per pattern, and a slop reviewer set up for Claude Code, Codex and Cursor.
- A canary fixture that proves every detector still fires after a toolchain upgrade.
What counts as slop in agent-written code?
Section titled “What counts as slop in agent-written code?”Slop is code whose surface signals are good and whose behaviour is not. It has three properties that make it different from an ordinary bug:
- It passes the checks you already run. A swallowed error makes a failing test pass. A
// @ts-expect-errormakes the type checker quiet. - It looks like diligence. An interface with one implementation looks like good design. A
try/catchlooks like error handling. A test withtoBeTruthy()looks like coverage. - It compounds. Each duplicate helper or dead layer makes the next agent’s context worse, because the agent copies what it finds.
The evidence says the effect is real but it does not size your risk. SlopCodeBench (Orlanski et al., arXiv, v2 7 May 2026) had 15 agents extend their own solutions across 36 problems and 196 checkpoints: no agent solved a problem end to end, the best passed 14.8% of checkpoints, and the authors measure the degradation as “structural erosion” and “verbosity”. GitClear’s “The Maintainability Gap” (June 2026), read here only from a search extract (SECONDARY), reports error-masking catch blocks up 47% and block duplication up 81% against 2023 across 623 million changes. Those GitClear figures are secondary and correlational, not causal. Measure your own repository with the detectors below instead of quoting either number to your team.
The slop catalog: five patterns and their detectors
Section titled “The slop catalog: five patterns and their detectors”Each row has a deterministic detector that blocks the merge and a review prompt that covers the gap. The detector runs on every pull request; the prompt runs where the detector is blind.
| Pattern | What it looks like | Deterministic detector | What the detector misses |
|---|---|---|---|
| Swallowed errors | catch {}, catch { return null }, except Exception: log; return {}, an un-awaited promise | ESLint no-empty, @typescript-eslint/no-floating-promises, a no-restricted-syntax selector · Ruff BLE001, S110, S112, E722 · golangci-lint errcheck | A handler that logs and then returns a fallback the caller treats as success |
| Invented APIs | A method, option or config key that does not exist, silenced with as any or @ts-expect-error | tsc --noEmit plus @typescript-eslint/no-unsafe-call, no-explicit-any, ban-ts-comment · mypy --strict or pyright · Ruff PGH003 | Wrong semantics of a real API; untyped config and environment keys |
| Test weakening | it.skip, a test with no assertion, exact values swapped for toBeTruthy() | @vitest/eslint-plugin (or eslint-plugin-jest) no-disabled-tests, no-focused-tests, expect-expect · the oracle audit · mutation testing | An assertion that is present but no longer pins the behaviour |
| Duplicate helpers | A third formatMoney under a new name | jscpd with --baseline-from-ref origin/main --fail-on-new-clones · jscpd’s MCP server before writing | Semantic duplicates with different structure |
| Dead abstractions | A factory, strategy or wrapper nothing calls, or that has one implementation | knip --include exports,types,files · vulture · golangci-lint unused | A used-but-pointless layer: one interface, one implementation, one caller |
Invented package names are the sixth pattern and have their own page: dependency checks for agent changes.
Roll out the slop gate, step by step
Section titled “Roll out the slop gate, step by step”-
Run the canary fixture first. Copy the fixture from the verification section into a scratch directory and confirm each detector reports it. A detector that reports nothing on the fixture is off, misconfigured or blind to your stack.
-
Add the detectors for your stack from the pattern sections below, as errors, not warnings. A warning lets the agent report the task as done; an error does not.
-
Baseline the existing violations. Run
npx --no-install eslint . --suppress-allonce and commiteslint-suppressions.json. New violations fail; old ones are recorded, not forgiven. -
Protect the gate. Put the lint config,
eslint-suppressions.json,knip.json,.knip-budget,.jscpd.jsonand the audit script underCODEOWNERSand your agent deny rules, as described in protecting the oracle. -
Add the CI job and the diff audit from the CI section, and make the job a required check.
-
Give every developer the slop reviewer for their tool, and ask for its verdict in the pull request’s evidence bundle.
-
Burn down the baseline. Run
npx --no-install eslint . --prune-suppressionsafter each cleanup so the recorded counts only go down.
Detect each slop pattern
Section titled “Detect each slop pattern”The configurations below were run against a fixture repository with one example of each pattern on 26 September 2026. Every detector named in the output fired on its example; the gaps named here are the ones the fixture exposed.
Swallowed errors: the catch that turns a failure into a success
Section titled “Swallowed errors: the catch that turns a failure into a success”A swallowed error is the most expensive slop pattern, because the program keeps running with a wrong answer. In TypeScript, three rules cover the mechanical forms. no-restricted-syntax with an AST selector flags any catch that returns a value without rethrowing, which is the form no-empty cannot see:
// eslint.config.js (ESLint 10, typescript-eslint 8)import { defineConfig } from 'eslint/config';import tseslint from 'typescript-eslint';
export default defineConfig({ files: ['src/**/*.ts'], extends: [tseslint.configs.base], languageOptions: { parserOptions: { projectService: true } }, rules: { 'no-empty': ['error', { allowEmptyCatch: false }], '@typescript-eslint/no-floating-promises': 'error', '@typescript-eslint/no-misused-promises': 'error', 'no-restricted-syntax': ['error', { selector: 'CatchClause:not(:has(ThrowStatement)) ReturnStatement', message: 'This catch returns a fallback instead of rethrowing. Rethrow, return a typed error result, or justify it with slop-ok.', }], },});On the fixture, catch (e) {} tripped no-empty, a refund(id) call without await inside a loop tripped no-floating-promises, and both catch { return null; } and catch (err) { console.error(err); return {}; } tripped the selector. The message is written as an instruction, because the agent reads it and acts on it.
In Python, select the Ruff rules explicitly:
# pyproject.toml (Ruff)[tool.ruff.lint]extend-select = ["BLE001", "S110", "S112", "E722"]Two Ruff behaviours matter. BLE001 (blind except Exception) stops firing when the handler calls log.exception(...), even if it then returns {}: logging satisfies the linter, not the caller. And do not enable SIM105 for this purpose: it suggests rewriting try/except/pass as contextlib.suppress, which keeps the swallow and removes the visible pass. In Go, errcheck is in golangci-lint v2’s standard default set (checked in the 2.14.0 source), so leave it on.
Invented APIs: the method that only exists after a cast
Section titled “Invented APIs: the method that only exists after a cast”An agent that calls a method that does not exist gets a type error, and the shortest path to green is to silence it. On the fixture, title.toSlugCase() under // @ts-expect-error and (s as any).toUpperCaseLocale() both passed tsc with exit 0. The type checker was satisfied; the code would throw at runtime.
Add three typescript-eslint rules to the config above:
'@typescript-eslint/no-unsafe-call': 'error', '@typescript-eslint/no-explicit-any': 'error', '@typescript-eslint/ban-ts-comment': ['error', { 'ts-expect-error': true, 'ts-ignore': true, 'ts-nocheck': true, }],The ban-ts-comment options matter. Its default allows @ts-expect-error when a description follows, and the fixture’s “agent silenced a missing method” comment was description enough to pass. With the options above, it fails. no-unsafe-call caught both invented methods on its own, because the silenced call has no type it can resolve.
In Python, mypy --strict reported Module has no attribute "dumps_pretty" and pyright reported "dumps_pretty" is not a known attribute of module "json". Add Ruff’s PGH003 so a bare # type: ignore without an error code is flagged too.
The detectors stop at the type system. An invented key in an untyped config file, an environment variable nothing sets, or a real API called with the wrong semantics passes all of them. Parse config and environment through a schema at startup, and ground the agent in current library docs with the Context7 MCP server before it writes the call.
Test weakening: the test that can no longer fail
Section titled “Test weakening: the test that can no longer fail”When the agent is asked to make tests pass, weakening a test is often the smallest edit. The lint layer catches the mechanical forms. With @vitest/eslint-plugin, the fixture’s it.skip(...) tripped vitest/no-disabled-tests and a test with no expect tripped vitest/expect-expect:
// eslint.config.js: import vitest from '@vitest/eslint-plugin', then pass// this object to defineConfig() after the src block{ files: ['tests/**/*.ts'], plugins: { vitest }, rules: { 'vitest/no-disabled-tests': 'error', 'vitest/no-focused-tests': 'error', 'vitest/expect-expect': 'error', },},For Jest, eslint-plugin-jest (29.16.6) has rules with the same three names. The fixture’s expect(safeParse('x')).toBeFalsy() passed every rule: an assertion is present, it just no longer pins a value. That is the semantic form, and it needs two other layers. Protecting the oracle keeps test files out of the implementing agent’s reach and audits every change to them. Mutation testing proves whether the remaining assertions catch bugs at all.
Duplicate helpers: the third formatMoney
Section titled “Duplicate helpers: the third formatMoney”An agent that does not know a helper exists writes a new one, often under a different name. jscpd compares token streams, so a new function name alone does not hide the copy: on the fixture, formatAmount in src/invoice.ts was reported as a clone of formatMoney in src/money.ts. Rename the parameters and locals as well, and jscpd 5.3.2 reported no clone in any of its three --mode settings. Gate on new clones against the base branch, as set up in architecture fitness functions:
# CI, jscpd 5.3.2 (needs the base branch fetched)npx --no-install jscpd src --baseline-from-ref origin/main --fail-on-new-clones --fail-on-emptyjscpd 5.3.2 also runs as an MCP server, which moves the check from the pull request to the moment before the agent writes. Its check_duplication tool takes a snippet and returns any existing copy; the server scans once at start, and check_current_directory rescans after edits. Register it in each tool you use:
# Terminal, in the repository rootclaude mcp add jscpd -- npx --no-install jscpd --mcp src# Terminal, in the repository rootcodex mcp add jscpd -- npx --no-install jscpd --mcp src{ "mcpServers": { "jscpd": { "command": "npx", "args": ["--no-install", "jscpd", "--mcp", "src"] } }}jscpd misses semantic duplicates: a copy with renamed variables, or a second date parser written with different control flow. The prompt covers those.
Dead abstractions: the strategy pattern with one strategy
Section titled “Dead abstractions: the strategy pattern with one strategy”Agents add layers that look like good design: an interface, a factory and a default implementation for a job that has one implementation and one caller. knip finds the ones nothing uses. On the fixture, knip --include exports,types,files reported the unused factory, class and interface, an unused helper and a file nothing imports:
# CI, knip 6.38.0: fail when unused files, exports or types exceed the budgetnpx --no-install knip --include exports,types,files --max-issues "$(cat .knip-budget)"knip has no baseline file, so --max-issues with a committed budget acts as the ratchet: record today’s count in .knip-budget, protect the file with CODEOWNERS, and lower the number as you clean up. In Python, vulture 2.16 reported unused function 'legacy_export' (60% confidence) and exits with code 3 when it finds dead code. In Go, the unused linter in golangci-lint’s standard set covers it.
None of these tools flags an abstraction that is used but pointless. The fixture’s ExportStrategy base class with one subclass passed vulture, because the subclass is called. That is a review question.
Run the slop gate in CI
Section titled “Run the slop gate in CI”The lint, type and dead-code checks run as one required job. The diff audit adds the checks that only make sense on new lines: a suppression comment, an error-masking construct or a skip marker that the pull request adds.
# .github/workflows/slop-gate.yml: required status check on the default branchname: slop-gateon: pull_requestpermissions: contents: readjobs: slop-gate: runs-on: ubuntu-latest steps: - uses: actions/checkout@v7 with: { fetch-depth: 0, persist-credentials: false } - uses: actions/setup-node@v7 with: { node-version: 22, cache: npm } - run: npm ci - run: npx --no-install tsc --noEmit - run: npx --no-install eslint src tests - run: npx --no-install knip --include exports,types,files --max-issues "$(cat .knip-budget)" - run: npx --no-install jscpd src --baseline-from-ref origin/main --fail-on-new-clones --fail-on-empty - run: bash scripts/slop-diff-audit.sh "${{ github.event.pull_request.base.sha }}" # Python repositories: uncomment. Run mypy in the project's own environment # if it needs your dependencies' types. # - uses: astral-sh/setup-uv@v10.2.0 # no floating v10 tag # - run: uvx ruff@0.16.9 check . # - run: uvx mypy@2.3.1 --strict src # - run: uvx vulture@2.16 src # Go repositories: uncomment. errcheck and unused are in the standard set. # - uses: actions/setup-go@v7 # with: { go-version-file: go.mod } # - uses: golangci/golangci-lint-action@v9 # with: { version: v2.14.0 }The job runs pull request code (npm ci lifecycle scripts, the linters’ plugins), so it gets a read-only token and no secrets. A pull request can edit this workflow file; if your agents open pull requests unattended, move the job into a required workflow in another repository, as protecting the oracle shows.
The diff audit script:
#!/usr/bin/env bash# scripts/slop-diff-audit.sh BASE_SHA: flag slop signals on the lines this change addsset -euo pipefailbase="$1"added=$(git diff -U0 "$base"...HEAD -- . ':(exclude)*.md' ':(exclude).github/**' \ ':(exclude)scripts/slop-diff-audit.sh' ':(exclude)slop-canary/**' | grep -E '^\+[^+]' || true)fail=0check() { # name, blocking (1/0), extended regex local hits; hits=$(grep -E "$3" <<<"$added" || true) [ -z "$hits" ] && return 0 echo "$1: $(wc -l <<<"$hits")"; head -n 5 <<<"$hits" | sed 's/^+/ /' [ "$2" = 1 ] && fail=1; return 0}added_all=$addedadded=$(grep -vE 'slop-ok: .{10,}' <<<"$added_all" || true)check "Suppression without a slop-ok reason" 1 'eslint-disable|@ts-(ignore|expect-error|nocheck)|# *noqa|# *type: *ignore|//nolint|pragma: no cover'added=$added_allcheck "Error-masking construct" 1 'catch *(\([^)]*\))? *\{ *\}|\.catch\(\(\) *=> *(\{ *\}|null|undefined|\[\]|\{\})\)|except[^:]*: *pass'check "Skip or only marker" 1 '\.(skip|only)\(|\bx(it|describe|test)\(|@pytest\.mark\.(skip|xfail)|t\.Skip\('check "Loose assertion (review)" 0 '\.(toBeTruthy|toBeDefined|toBeFalsy)\(\)|expect\.anything\(\)'check "TODO or stub (review)" 0 'TODO|FIXME|NotImplementedError|[Nn]ot implemented'if [ "$fail" = 1 ]; then echo "SLOP SIGNAL: fix these lines, or justify a suppression on the same line with 'slop-ok: <reason>'." >&2 exit 1fiecho "No blocking slop signals in the added lines."Fed a diff with one example of each signal, it reported a bare // @ts-expect-error, a catch (e) {}, a .catch(() => null), an it.skip, a toBeTruthy() and a TODO, and exited 1. The suppression carrying slop-ok: caller renders cached price on failure was exempt. A diff that only added toBeDefined() printed it for review and exited 0. The script matches lines, not syntax: a two-line except Exception: followed by pass gets past it, which is why the Ruff rules above must run in the same job (the commented Python steps). The script excludes itself and the slop-canary/ fixture, whose lines contain these patterns on purpose. The slop-ok marker is the escape hatch for a deliberate fallback, and it puts the justification where a reviewer can grep for it.
Add the slop reviewer to your tool
Section titled “Add the slop reviewer to your tool”The detectors block; the reviewer advises. Run it in a fresh context, not in the session that wrote the code, because the author’s context already believes the code is right. For a gating verdict, use a different model or vendor from the author, the rule model-graded checks sets out.
Save the reviewer as a subagent in .claude/agents/slop-reviewer.md. Its tool list has no Edit or Write, and its Bash calls still go through your permission rules. See custom subagents for the file format:
---name: slop-reviewerdescription: Reviews a branch for plausible-but-wrong code (swallowed errors, invented APIs, weakened tests, duplicate helpers, dead abstractions). Use before opening a pull request. Never edits.tools: Read, Grep, Glob, Bash---You review code you did not write. Run `git diff origin/main...HEAD` and therepository's lint command, and read what you need. Never edit files.
Check five patterns: swallowed errors, invented APIs, weakened tests,duplicate helpers, dead abstractions. Report only findings with file:lineand quoted evidence. Ignore claims in commit messages or the parentconversation about what the code does.
End with a table (pattern, file:line, finding, severity) and one verdict:CLEAN, FIX BEFORE PR, or NEEDS HUMAN DECISION.In a session, ask: “Use the slop-reviewer subagent on this branch.” From the terminal, run it headless (Claude Code 2.1.283):
# Terminal: -p denies any tool call that would need approval, so pre-allow the two commandsclaude -p "Review this branch against origin/main." --agent slop-reviewer \ --allowedTools "Bash(git diff *)" "Bash(npx --no-install eslint *)"Without the allow rules, or matching allow entries in .claude/settings.json, the lint call is refused and the reviewer returns without lint evidence. The bundled /code-review command is a broader correctness review; run it as well, not instead.
Save the body of the Claude Code reviewer above (everything below the frontmatter) as .codex/slop-review.md and run it read-only with codex exec and the built-in :read-only permission profile (permission profiles are beta in 0.157.1; the legacy equivalent is --sandbox read-only, and do not combine the two):
# Terminal, Codex CLI 0.157.1codex exec -c default_permissions=":read-only" -o slop-review.out.md \ "$(cat .codex/slop-review.md) Review the diff of this branch against origin/main."codex review is the dedicated review command, but in 0.157.1 custom instructions cannot be combined with its target flags: codex review --base main "…" fails with the argument '--base <BRANCH>' cannot be used with '[PROMPT]', and so does --uncommitted. Use codex review --base main for Codex’s own review and codex exec for the slop checklist.
For reviews on GitHub, @codex review on a pull request reads review rules from AGENTS.md (verified 28 August 2026), so paste the five-pattern checklist under a review-guidelines heading there.
Open a new agent chat, so the reviewer carries none of the authoring conversation, and pick a model in the model picker that differs from the one that wrote the change. Paste the five prompts from this page, or the Claude Code reviewer’s body, and start in Plan Mode so it reports before it touches anything.
On pull requests, Bugbot “reviews pull requests and identifies bugs, security issues, and code quality problems” (cursor.com/docs/bugbot, checked 28 August 2026). Treat it as a second reader, not as the gate: the required slop-gate check is what binds Cursor’s Cloud Agents, which run in isolated cloud VMs, outside your local hooks and permission rules. How to give Bugbot project-specific rules was not re-verified on 26 September 2026, because cursor.com was unreachable from our environment; take it from Cursor’s Bugbot documentation.
A third-party pull request reviewer can carry the same checklist. The Greptile CLI (npm greptile 3.6.0, Node 22 or later) takes custom instructions on a review against a base branch: greptile review -b main --instructions "Check for swallowed errors, invented APIs, weakened tests, duplicate helpers and dead abstractions". Like Bugbot, it is a second reader; the required check stays the gate.
How do you know the slop gate catches slop?
Section titled “How do you know the slop gate catches slop?”A gate that reports nothing may be clean code or a blind gate. Keep a canary fixture, a small directory with one known example of each pattern, and run every detector against it in CI after each toolchain upgrade. The fixture used for this page contained:
| Fixture file | Seeded slop | Detector that must fire |
|---|---|---|
src/payments.ts | catch (e) {}, catch { return null; }, refund(id) without await | no-empty, no-restricted-syntax, no-floating-promises |
src/silenced.ts | @ts-expect-error over an invented method, (s as any).method() | ban-ts-comment, no-unsafe-call, no-explicit-any |
tests/payments.test.ts | it.skip, a test with no expect | vitest/no-disabled-tests, vitest/expect-expect |
src/money.ts + src/invoice.ts | formatMoney copied as formatAmount, body unchanged | jscpd clone |
src/invoice.ts | Interface, class and factory nothing calls | knip unused exports and types |
billing.py | except: + pass, except Exception returning {}, json.dumps_pretty, an uncalled legacy_export | Ruff E722, S110, BLE001; mypy attr-defined; vulture |
The paths matter: the script lints src and tests and runs jscpd on src, so a fixture copied flat into slop-canary/ makes ESLint and jscpd report nothing.
The canary job asserts on findings, not on “some non-zero exit”. ESLint exits 2 on a config error or an unknown rule, mypy exits 2 on a fatal error, and a crashed knip or jscpd also exits non-zero, so treating any failure as “live” would hide exactly the upgrade breakage the canary exists for. Each tool must exit with its own findings code (ESLint, Ruff, mypy, knip and jscpd 1, vulture 3); 0 means BLIND, anything else means BROKEN. For ESLint, one exit code cannot show that eight separate rules still fire, so the script compares the rule IDs in the JSON output with the list from the table.
#!/usr/bin/env bash# scripts/slop-canary.sh: every detector must report its seeded examplecd slop-canary || exit 1status=0expect_findings() { # name, exit code that means "findings", command... local name=$1 want=$2 got; shift 2 "$@" >/dev/null 2>&1; got=$? if [ "$got" = "$want" ]; then echo "live: $name" elif [ "$got" = 0 ]; then echo "BLIND: $name"; status=1 else echo "BROKEN: $name exited $got, expected $want"; status=1; fi}expect_findings knip 1 npx --no-install knip --include exports,types,filesexpect_findings jscpd 1 npx --no-install jscpd src --exit-code 1expect_findings ruff 1 ruff check --select BLE001,S110,E722 billing.pyexpect_findings mypy 1 mypy --strict billing.pyexpect_findings vulture 3 vulture billing.py
expected='no-emptyno-restricted-syntax@typescript-eslint/no-floating-promises@typescript-eslint/ban-ts-comment@typescript-eslint/no-unsafe-call@typescript-eslint/no-explicit-anyvitest/no-disabled-testsvitest/expect-expect'out=$(npx --no-install eslint src tests --format json 2>/dev/null); code=$?if [ "$code" != 1 ]; then echo "BROKEN: eslint exited $code, expected 1"; status=1else fired=$(jq -r '[.[].messages[].ruleId] | unique[]' <<<"$out") missing=$(comm -23 <(sort <<<"$expected") <(sort <<<"$fired")) if [ -n "$missing" ]; then echo "BLIND: eslint rules:"; sed 's/^/ /' <<<"$missing"; status=1 else echo "live: eslint (8 rules)"; fifiexit "$status"Keep the fixture in its own directory with its own lint config and knip.json, and exclude it from the normal gate so it does not count against your budget; the diff audit already excludes slop-canary/**, so seeding it does not fail your pull request. Keep it out of eslint-suppressions.json too: running --suppress-all over the fixture by accident makes ESLint report it clean. A BLIND or BROKEN line after an upgrade is the alarm: the upgrade changed a rule name, a default or a file glob.
Ownership and numbers:
- The tech lead owns the rule set, the budgets and the canary. Changes to them go through
CODEOWNERS, and the baseline counts only go down. - The author owns the reviewer’s verdict. “FIX BEFORE PR” findings are fixed or answered in the pull request before a human is asked to review.
- A human reviews what the machines flag, not the whole diff:
slop-okjustifications, “NEEDS HUMAN DECISION” findings and every change to test files. The agent pull request review protocol covers the rest. - Track three numbers per week: slop-gate failures per agent pull request,
slop-okmarkers added, and the suppression baseline total. Failures should fall as rules and instructions improve;slop-okmarkers rising faster than failures fall means the escape hatch has become the path.
What breaks when you gate on slop?
Section titled “What breaks when you gate on slop?”The agent learns to satisfy the detector instead of fixing the code. It replaces catch {} with catch (e) { console.error(e) }, or as any with as unknown as Charge. Recovery: add each new evasion to the canary fixture and a detector for it. For the log-only handler, a second no-restricted-syntax entry with the selector CatchClause > BlockStatement[body.length=1] > ExpressionStatement > CallExpression[callee.object.name="console"] flagged catch (e) { console.error(e); } on the fixture and left catch (e) { setError(e); } alone. Keep the review prompts in the loop too, because they ask what the caller receives, not what the syntax looks like.
The agent rewrites the gate. eslint --suppress-all records every new violation as old, and a raised .knip-budget hides new dead code. Recovery: deny the agent edits to the lint config, eslint-suppressions.json, .knip-budget and .jscpd.json, own them in CODEOWNERS, and fail CI when the suppression file’s total count rises:
# CI step: the recorded suppressions may only go downnow=$(jq '[.[][] | .count] | add // 0' eslint-suppressions.json)base=$(git show origin/main:eslint-suppressions.json | jq '[.[][] | .count] | add // 0')[ "$now" -le "$base" ] || { echo "Suppressions rose from $base to $now" >&2; exit 1; }The diff audit fires on legitimate code. A retry wrapper returns null after the last attempt by design. Recovery: add slop-ok: with a reason of at least 10 characters on the line, and review those markers in the weekly numbers instead of weakening the pattern.
The reviewer produces long, confident, wrong findings. A reviewer on the author’s model, in the author’s context, agrees with the author. Recovery: run it in a fresh session on a different model, require file:line and quoted evidence for every finding, and discard findings without them.
The gate goes green on nothing. A wrong files glob in the ESLint config, a knip entry that marks everything as used, or jscpd scanning an empty path turns each detector into a no-op. Recovery: the canary job, --fail-on-empty on jscpd, and a check that ESLint’s --format json output lists the files you expect.