Skip to content

Keeping an agent-written codebase healthy

Keeping an agent-written codebase healthy means watching its trend, not only gating each pull request: measuring line rework, revert rate, complexity over budget, duplication and change hotspots every week, lowering a loop’s autonomy when a signal crosses its threshold, and running scheduled refactor loops that prove each change with before-and-after numbers.

This page is for tech leads and the developers who run agent loops, with a monthly view for the CTO. Your fitness gates are green, every pull request passes its tests, and the team merges more than ever. Yet the checkout module now takes three agent sessions to change instead of one, the same files appear in every sprint, and a third of last month’s code has already been rewritten. No single pull request caused it, so no per-PR gate catches it. You need a signal over time and a rule for what happens when it turns bad.

  • Five health signals with exact definitions, all computed from git and two open-source tools, with no vendor dashboard.
  • A 120-line script and a weekly GitHub Actions job that produce the report, compare it with a baseline and open an issue when the status turns red.
  • A threshold table that says which signal lowers a loop’s autonomy, by how much, and who signs the change.
  • A scheduled refactor loop for Claude Code, Codex and Cursor, with guardrails that keep it from editing the oracle.
  • Three copy-paste prompts: the unattended refactor run, the weekly triage and a duplicate consolidation.

Why does an agent-written codebase decay when every pull request passes?

Section titled “Why does an agent-written codebase decay when every pull request passes?”

Per-PR gates check that a change does not break a rule. They cannot see that the rules are passing while the code is getting harder to change. Architecture fitness functions stop new layering breaks and budget violations; this page covers the drift those gates allow: rework, concentration of complexity in a few files, and copies that each stay under the duplication budget.

Three independent sources describe the same pattern:

  • Faros AI (“AI Engineering Report 2026: The Acceleration Whiplash”, April 2026, telemetry from 22,000 developers on its own platform) reports throughput up (epics completed per developer +66.2%, pull request merge rate per developer +16.2%) and quality down (bugs per developer +54%, incidents per pull request +242.7%, code churn +861%). The customer base is self-selected.
  • DORA (Google Cloud, 2025 report, 23 September 2025): “we observe a positive relationship between AI adoption on both software delivery throughput and product performance.” And: “However, AI adoption does continue to have a negative relationship with software delivery stability.”
  • SlopCodeBench (Orlanski et al., arXiv, v2 7 May 2026) had agents repeatedly extend their own solutions; the best passed 14.8% of 196 checkpoints, and the authors measure the degradation as “structural erosion (concentrated complexity) and verbosity (redundant code)”.

GitClear’s June 2026 study of 623 million changes, seen only through a secondary search extract, points the same way: moved (refactored) code fell from 13% to 3.8% of changed lines against 2023. GitClear presents it as correlation, not cause. The practical reading is the same in all four: generation scales on its own, and refactoring does not happen unless someone schedules it.

Five signals cover the failure modes above. Each is a trend, read against the repository’s own baseline, never against another team.

SignalDefinitionWhat it catchesSource
14-day line reworkOf the lines added 14 to 28 days ago, the share no longer present today (whitespace and in-file moves ignored)Code written, merged and rewritten within weeks: the spec was wrong or the design did not holdgit log --numstat and git blame -w -M
Revert rateCommits whose subject starts with Revert " ÷ all non-merge commits in 30 days, against the 30 days beforeChanges that were wrong enough to undogit log
Complexity over budgetShare of code lines inside functions whose cyclomatic complexity exceeds the budget (default 10)Complexity concentrating in fewer, larger functionslizard 1.24.0
DuplicationDuplicated lines as a percentage of all linesHelpers copied instead of reused, below the per-PR clone gatejscpd 5.3.2
HotspotsCommits in 90 days × summed complexity, per file, top 10The files where decay costs the most because agents touch them every weekgit log plus lizard

The first two measure the process, the next two the code, and hotspots tell the refactor loop where to spend its effort. Complexity weighted by change frequency follows Adam Tornhill’s hotspot analysis (Your Code as a Crime Scene): complex code nobody touches is cheap, and complex code everyone touches is where the time goes.

Revert rate here is a git proxy. The organization-wide definitions, including the 14-day follow-up fix rate computed from pull request links, live in metrics that survive agents; use those on the CTO’s dashboard and this page’s numbers inside one repository.

  1. Add the script. Save it as scripts/codebase_health.py. It needs Python 3, git with full history and lizard; jscpd is optional and reported as null when missing. We ran its complexity and duplication functions against a TypeScript codebase on 26 September 2026 with lizard 1.24.0 and jscpd 5.3.2.

    #!/usr/bin/env python3
    """scripts/codebase_health.py: weekly health signals for an agent-written codebase.
    Run from the repository root with full history (fetch-depth: 0 in CI).
    Needs git and lizard (pip install lizard==1.24.0); jscpd is optional.
    Prints one JSON report to stdout. Compares against health/baseline.json if present.
    """
    import argparse, csv, io, json, os, re, subprocess, tempfile
    from collections import Counter, defaultdict
    SHA = re.compile(r"^[0-9a-f]{40,64} ")
    WORSE_IF_HIGHER = ["revert_rate_30d", "rework_14d", "over_budget_nloc_share", "duplication_pct"]
    # A zero baseline would never alert, so compare against at least this absolute floor.
    FLOOR = {"revert_rate_30d": 0.01, "rework_14d": 0.02, "over_budget_nloc_share": 0.01, "duplication_pct": 0.5}
    def git(*args):
    return subprocess.run(["git", *args], check=True, capture_output=True, text=True).stdout
    def revert_rate(src, since, until):
    subjects = git("log", "--no-merges", f"--since={since} days ago", f"--until={until} days ago",
    "--format=%s", "--", src).splitlines()
    reverts = sum(1 for s in subjects if s.startswith('Revert "'))
    return round(reverts / len(subjects), 4) if subjects else None
    def rework(src, as_of):
    """Share of lines added 14-28 days before `as_of` that are gone at `as_of`."""
    rev = "HEAD" if as_of == 0 else git("rev-list", "-1", f"--before={as_of} days ago", "HEAD").strip()
    if not rev:
    return None
    log = git("log", rev, "--no-merges", "--no-renames", f"--since={as_of + 28} days ago",
    f"--until={as_of + 14} days ago", "--numstat", "--format=@%H", "--", src)
    cohort, added = set(), defaultdict(int)
    for line in log.splitlines():
    if line.startswith("@"):
    cohort.add(line[1:])
    elif line.strip():
    plus, _minus, path = line.split("\t", 2)
    if plus != "-": # "-" marks a binary file
    added[path] += int(plus)
    total = sum(added.values())
    if not total:
    return None
    survived = 0
    for path, count in added.items():
    if subprocess.run(["git", "cat-file", "-e", f"{rev}:{path}"], capture_output=True).returncode:
    continue # deleted (or renamed) since: nothing survived at this path
    blame = git("blame", "-w", "-M", "--line-porcelain", rev, "--", path)
    alive = sum(1 for l in blame.splitlines() if SHA.match(l) and l.split(" ", 1)[0] in cohort)
    survived += min(alive, count)
    return round(1 - survived / total, 4)
    def complexity(src, budget):
    out = subprocess.run(["lizard", "--csv", src], capture_output=True, text=True).stdout
    per_file, funcs, over, nloc_all, nloc_over, worst = Counter(), 0, 0, 0, 0, 0
    for row in csv.reader(io.StringIO(out)):
    nloc, ccn, path = int(row[0]), int(row[1]), os.path.normpath(row[6])
    funcs, nloc_all, worst = funcs + 1, nloc_all + nloc, max(worst, ccn)
    per_file[path] += ccn
    if ccn > budget:
    over, nloc_over = over + 1, nloc_over + nloc
    share = round(nloc_over / nloc_all, 4) if nloc_all else None
    return {"functions": funcs, "over_budget": over, "over_budget_nloc_share": share, "max_ccn": worst}, per_file
    def duplication(src):
    with tempfile.TemporaryDirectory() as out:
    run = subprocess.run(["npx", "--no-install", "jscpd", src, "--reporters", "json",
    "--output", out, "--silent"], capture_output=True, text=True)
    report = os.path.join(out, "jscpd-report.json")
    if run.returncode or not os.path.exists(report):
    return None
    with open(report) as f:
    return round(json.load(f)["statistics"]["total"]["percentage"], 2)
    def main():
    ap = argparse.ArgumentParser()
    ap.add_argument("--src", default="src")
    ap.add_argument("--ccn-budget", type=int, default=10)
    ap.add_argument("--baseline", default="health/baseline.json")
    a = ap.parse_args()
    stats, per_file = complexity(a.src, a.ccn_budget)
    touches = Counter(os.path.normpath(p) for p in git(
    "log", "--no-merges", "--since=90 days ago", "--format=", "--name-only", "--", a.src).splitlines() if p.strip())
    hotspots = sorted(({"file": f, "commits_90d": n, "ccn_sum": per_file[f], "score": n * per_file[f]}
    for f, n in touches.items() if per_file.get(f)), key=lambda h: -h["score"])[:10]
    trend = [rework(a.src, d) for d in (28, 14, 0)] # oldest first
    report = {
    "revert_rate_30d": revert_rate(a.src, 30, 0),
    "revert_rate_prev_30d": revert_rate(a.src, 60, 30),
    "rework_14d": trend[-1],
    "rework_trend": trend,
    **stats,
    "duplication_pct": duplication(a.src),
    "hotspots": hotspots,
    }
    signals, status = {}, "green"
    if os.path.exists(a.baseline):
    with open(a.baseline) as f:
    base = json.load(f)
    for key in WORSE_IF_HIGHER:
    now, then = report.get(key), base.get(key)
    if now is None or then is None:
    continue
    ratio = now / max(then, FLOOR[key])
    signals[key] = "red" if ratio > 1.5 else "amber" if ratio > 1.25 else "green"
    states = list(signals.values())
    if "red" in states or states.count("amber") >= 2:
    status = "red"
    elif "amber" in states:
    status = "amber"
    report.update(signals=signals, status=status if signals else "no-baseline")
    print(json.dumps(report, indent=2))
    if __name__ == "__main__":
    main()
  2. Run it locally once from the repository root, and read the hotspot list before anything else. If the top 10 surprises the team, the script is pointed at the wrong directory or the history is shallow.

    Terminal window
    # Terminal, repository root
    python3 -m pip install lizard==1.24.0
    python3 scripts/codebase_health.py --src src > health.json
  3. Calibrate for four weeks, then commit a baseline. Keep each weekly report, take the median of each signal, and commit it as health/baseline.json with the four reports’ dates in the pull request description. Put health/ and the script under CODEOWNERS, owned by the tech lead. The thresholds compare against this file, so a baseline anyone can edit is a threshold anyone can switch off.

  4. Schedule the job. It runs no agent and needs no model key. It fails loudly when the analysis saw nothing, writes the report to the job summary, uploads it as an artifact and opens an issue on red.

    .github/workflows/codebase-health.yml
    name: codebase-health
    on:
    schedule:
    - cron: '23 5 * * 1' # Mondays 05:23 UTC
    workflow_dispatch:
    permissions:
    contents: read
    issues: write
    jobs:
    health:
    runs-on: ubuntu-latest
    steps:
    - uses: actions/checkout@v7
    with:
    fetch-depth: 0 # rework and hotspots need the history
    persist-credentials: false
    - uses: actions/setup-node@v7
    with:
    node-version: 22
    - run: npm ci # provides jscpd, pinned to 5.3.2 in devDependencies
    - uses: actions/setup-python@v7
    with:
    python-version: '3.13'
    - run: python -m pip install lizard==1.24.0
    - run: python scripts/codebase_health.py --src src > health.json
    - name: Fail if the analysis was blind
    run: jq -e '.functions > 0 and .rework_14d != null' health.json
    - uses: actions/upload-artifact@v7
    with:
    name: codebase-health
    path: health.json
    - name: Summary, and an issue on red
    env:
    GH_TOKEN: ${{ github.token }}
    run: |
    status=$(jq -r .status health.json)
    { echo "## Codebase health: $status"; echo '```json'; cat health.json; echo '```'; } >> "$GITHUB_STEP_SUMMARY"
    if [ "$status" = "red" ]; then
    gh issue create --title "Codebase health is red ($(date +%F))" \
    --body "Signals: $(jq -c .signals health.json). Apply the red actions from loops.yaml. Report: $GITHUB_SERVER_URL/$GITHUB_REPOSITORY/actions/runs/$GITHUB_RUN_ID"
    fi

The job checks out only the default branch, holds no secret beyond a token that can write issues, and runs no code from a pull request, so it is safe to schedule.

Which thresholds should lower a loop’s autonomy?

Section titled “Which thresholds should lower a loop’s autonomy?”

A signal is only useful if crossing it changes what the agents are allowed to do. The script marks a signal amber above 1.25 times its baseline and red above 1.5 times; two ambers count as red. These are our suggested starting values, not research findings: tune them after the four calibration weeks, and change them only through a CODEOWNERS-approved pull request. Each ratio uses the baseline or a small absolute floor, whichever is larger (1 percentage point for revert rate and complexity share, 2 for rework, 0.5 for duplication), because a baseline of zero, normal for reverts in a clean repository, would otherwise never raise an alert.

StatusTriggerWhat changes for the loopsWho signs
GreenNo signal above 1.25× baselineLoops keep their level; the refactor loop runs on the top hotspotNobody; the report is filed
AmberOne signal above 1.25×Promotions freeze. Changes to this week’s top five hotspot files need a human reading the diff, even in L4 loops. The refactor loop targets the signal that turned amberTech lead, in the weekly triage
RedAny signal above 1.5×, or two ambersEvery loop whose scope includes a top-five hotspot drops one level on the autonomy ladder: L5 loses auto-merge, L4 goes back to humans reading every diff, as at L3. Unattended feature loops in that scope pause; only the refactor loop runs thereTech lead demotes; the CTO sees it in the monthly trend
RecoveryTwo consecutive green weeksThe loop may go back up one level, after the usual promotion auditTech lead, with the audit sample

Write the rule into the loop register, so a red week is a lookup, not a debate:

# loops.yaml (excerpt): see the one-map page for the full register
- loop: checkout-feature-work
owner: checkout-team
level: L4
health_scope: [src/checkout/, src/payments/]
on_amber: freeze promotion; humans read diffs touching this week's top-5 hotspots
on_red: demote to L3 until two consecutive green weeks, then re-audit
last_health_status: green (2026-09-21)

Demotion is by scope, not by team. A dependency-bump loop that never touches src/checkout/ keeps running at its level while the checkout feature loop drops a rung.

The fix for decay is refactoring on a schedule, done by an agent under tighter rules than feature work. Five rules make it safe to run unattended, and they are an instance of protecting the oracle: the agent cannot edit the checks that judge it.

  1. One target per run: the top hotspot from health.json, unless it is listed in health/skip.txt.
  2. Behaviour-preserving only. No change to tests, fixtures, fitness rules, baselines, health/, scripts/codebase_health.py, CI files, agent config (.claude/, .codex/, .cursor/), lockfiles or package.json. Tests must pass unchanged. If the target has weak tests, a human-started task writes characterization tests first.
  3. A size cap: 400 changed lines. A larger refactor is a planned task, not a loop run.
  4. A measured win. The pull request body quotes the target’s complexity and clone count before and after, from the same tools the health job uses. No measurable win, no pull request.
  5. Normal gates, human merge. Refactor pull requests go through the same required checks and a code owner’s approval. Start the loop itself at L3 in loops.yaml and promote it on evidence like any other loop.

Store the run prompt (the first tip below) as .github/prompts/refactor-loop.md under CODEOWNERS, so all three tools run the same instructions.

Routines run a saved prompt against a repository on Anthropic-managed cloud infrastructure, or on your organization’s self-hosted environment (Team and Enterprise), on a schedule, an API call or a GitHub event. They are a research preview on Pro, Max, Team and Enterprise plans, with a minimum interval of one hour (checked against the routines documentation and Claude Code 2.1.283 on 26 September 2026). Create one from a session:

/schedule weekly on Tuesday at 6:41, in acme/shop: follow .github/prompts/refactor-loop.md exactly and open a pull request only if it reports a measured win

Claude asks for the repository, environment and prompt before saving. Know four things before you rely on it:

  • A routine runs as a full cloud session with no permission-mode picker, and commits and pull requests carry your GitHub identity. Remove every connector it does not need, because it can call their tools without asking.
  • It pushes to claude/-prefixed branches. Your branch ruleset, required checks and CODEOWNERS are what stop it; deny rules in your local settings are not a boundary for a cloud run.
  • Install lizard in the environment’s setup script (on a self-hosted environment, on the runner itself), and keep network access at the default Trusted level unless the run needs more.
  • A green status in the routine’s run list means the session started and exited without an infrastructure error, not that the refactor worked. The pull request with its before-and-after numbers is the only success signal, so a week with no pull request and no skip note in the transcript needs a look.

For the human side, the bundled /simplify skill reviews code for reuse of existing helpers, simplification, efficiency and abstraction level, and applies the fixes. Point it at the top hotspot in the weekly triage. For a deeper monthly session on the worst module, Matt Pocock’s improve-codebase-architecture skill surveys recently changed code and then interviews you about one refactor before touching anything:

Terminal window
claude plugin install mattpocock-skills

The prompts call npm run fitness, the single script defined on architecture fitness functions. Substitute your own gate command if you named it differently.

How do you know the health loop itself works?

Section titled “How do you know the health loop itself works?”

A health report that reads zero is worse than none, because a green status buys trust it has not earned. The workflow’s jq -e step fails the job when lizard analysed no functions or git returned no rework cohort, which is what a shallow clone, a wrong --src or a missing tool produces. A fortnight with no commits under --src also yields a null rework cohort. On a low-traffic repository, check only .functions > 0 and treat a null rework as “no data” in the triage, so a red job keeps meaning a blind analysis. Check the hotspot list against the team’s intuition after every toolchain change.

A refactor pull request is verified without reading every line:

  • Behaviour: the existing tests pass unchanged, and the protected-path check (or your branch ruleset and CODEOWNERS) proves no test moved. That proof is only as strong as the suite. Before the loop touches a hotspot, check how strong the oracle is for that file, and add it to health/skip.txt until it is strong enough.
  • Shape: the fitness gates from architecture fitness functions pass, with no rule or baseline edited.
  • Value: the before and after numbers in the body reproduce when the reviewer re-runs lizard or jscpd on the branch. A refactor that moves complexity instead of removing it shows up as a new hotspot next week.
  • Reviewer focus: the reviewer reads the risk notes and samples the changed call sites, as in reviewing an agent’s pull request. The before and after numbers belong in the pull request’s evidence bundle.
WhoOwnsSigns
Tech leadThe baseline, the thresholds, health/skip.txt and the refactor promptDemotions, promotions and every baseline change
Developer on rotationThe weekly triage and the refactor pull requestsMerge of refactor pull requests, as code owner
CTOThe monthly trend across repositoriesThe threshold policy, and which repositories must run the loop

For the CTO, three numbers per repository and month are enough: weeks spent in each status, loops demoted and re-promoted, and the share of refactor pull requests merged. A repository that is always green with no refactor merges is either healthy or measuring the wrong directory; the hotspot list tells you which.

What breaks when you run codebase health loops?

Section titled “What breaks when you run codebase health loops?”

The agent games the complexity number. It splits one complex function into five helpers with seven parameters each, and the CCN falls while the design gets worse. Recovery: keep max-params and function-length budgets in the fitness gates, and read the rework number next to complexity: a split that did not help gets rewritten within weeks.

Refactoring raises the rework signal. A refactor deletes recent lines on purpose, so a busy refactor week can push rework to amber. Recovery: read rework with the week’s refactor pull requests beside it in the triage, and do not demote a loop on a rise the refactor loop caused.

The baseline drifts up. Someone re-baselines during a bad month, and the decay becomes the new normal. Recovery: the baseline changes only through a CODEOWNERS-approved pull request with its four source reports, and never while the status is red.

Red becomes background noise. A weekly issue nobody closes trains the team to ignore it. Recovery: red changes what loops may do, through loops.yaml, so ignoring it has a cost; cap the open health issues at one and close it at the next green week.

The revert rate miscounts. It counts commits, so a merge-commit workflow with many commits per pull request dilutes it, and teams that fix forward never revert. Recovery: use the pull-request-based follow-up fix rate from the metrics page on the dashboard, and keep this proxy for trends inside one repository.

The refactor loop keeps choosing the same file. Its fix did not reduce the score enough, so the file stays on top. Recovery: after two runs on one target without a merged pull request, add it to health/skip.txt and hand it to a human-led session with improve-codebase-architecture or an architecture decision.

An agent incident starts in a hotspot. A red signal is often the early warning for the incident that follows. Recovery: tag the incident with the hotspot and the loop in the agent incident process, and classify it with the failure taxonomy.

The next page in this section turns the failures these signals point at into a taxonomy you can tag incidents with.