Skip to content

Making a codebase agent-ready

An agent-ready codebase is a repository in which automated checks, not a human reading the diff, can decide whether an agent’s change is correct. Six properties make that possible: one command that runs every gate, fast feedback, deterministic tests, strict types and lint, enforced module boundaries, and seeded, reproducible data. Each property is scored 0–3 with an audit prompt.

This page is for developers who want to stop reading every line an agent writes, and for tech leads deciding which repositories are ready for background runs. You gave the agent a well-specified ticket. It reported “all tests pass”, and you still spent 40 minutes reading the diff, because the test command it ran skipped the integration suite, two tests fail one run in five anyway, and nothing would have stopped it importing the billing module from the UI layer. The ticket was fine; the repository could not prove anything.

What you’ll walk away with from an agent-readiness audit

Section titled “What you’ll walk away with from an agent-readiness audit”
  • A 0–3 scoring rubric for six properties, with the evidence that earns each score.
  • A copy-paste audit prompt that makes the agent measure, not guess.
  • The order to fix things in, and three prompts that do the fixing: the one-command check, flaky tests, and enforced boundaries.
  • A canary test that proves your checks catch defects before unattended runs.
  • The failure modes, above all the agent that makes the check green by weakening it.

Why does the codebase decide whether you can stop reading diffs?

Section titled “Why does the codebase decide whether you can stop reading diffs?”

An agent’s claim that its work is done is worth exactly as much as the check behind it. If the only way to know a change is correct is to read it, you read it, no matter how good the model is. Every property on this page converts one kind of reading into a check that runs without you.

The strongest third-party evidence for this comes from DORA. The 2025 DORA report (Google Cloud blog, Nathen Harvey and Derek DeBellis, 2025-09-23) found a positive relationship between AI adoption and delivery throughput, and a negative one with delivery stability. Its explanation: “Without robust control systems, like strong automated testing, mature version control practices, and fast feedback loops, an increase in change volume leads to instability. Teams working in loosely coupled architectures with fast feedback loops see gains, while those constrained by tightly coupled systems and slow processes see little or no benefit.” The same report’s summary line is that AI “amplifies what’s already there”.

Stripe shows the other end: its Minions merge more than 1,300 pull requests a week with no human-written code (vendor-internal figure, Alistair Gray, stripe.dev, 2026-02-19), a loop that stops after “at most two rounds of CI” (Minions Part 1, stripe.dev, 2026-02-09) and runs against “over three million” tests.

In harness terms, the codebase sits next to the seven harness layers rather than inside them. The harness bounds what the agent can do; the codebase decides whether anything can check what it did.

Which six properties make a codebase agent-ready?

Section titled “Which six properties make a codebase agent-ready?”
#PropertyWhat the agent does without itThe check it makes possible
1One-command checkGuesses the test command, runs the unit suite only, and reports “tests pass”“Done” means one named command exited 0, and CI runs the same command
2Fast feedbackSkips the slow suite, or batches 30 edits before the first run, so failures arrive late and tangledThe agent runs the check after every meaningful edit and fixes failures while the cause is still local
3Deterministic testsLearns that red is sometimes noise, retries until green, or “fixes” a flake by editing the testA red result always means a defect, so a green result means something
4Strict types and lint as gatesPasses any, None or an unchecked index through, and invents a method the compiler would have rejectedWhole classes of error (wrong shape, missing field, invented API) fail in seconds without a test
5Enforced module boundariesImports whatever makes the change work, so the diff is correct today and the architecture erodesA layer violation fails the check, so you review the interface, not every import
6Seeded, reproducible data and setupNeeds a shared database, your credentials or a manual step, so it either cannot verify or verifies against the wrong stateA clean clone reaches a green check with one bootstrap command, locally, in a worktree or in a cloud environment

Two things are deliberately not on the list. Test coverage is not: a high number with weak assertions passes wrong code, and the oracle-strength page covers how to measure what tests actually catch. A context file (AGENTS.md, CLAUDE.md, Cursor Rules) is not either: it names the command, but a repository without the properties above cannot be fixed by describing them. See concise AGENTS.md and CLAUDE.md files for what belongs there.

How do you score each property from 0 to 3?

Section titled “How do you score each property from 0 to 3?”

Score each property on evidence you can point at: a command, its output, its run time, a config file. A score without evidence is a 0. The thresholds for run times are this page’s recommendation for an agent’s inner loop, not a published benchmark; adjust them once, write your version into the scorecard, and keep them fixed so scores stay comparable over time.

Property0123
1. One-command checkNo single command; the gates live only in CI YAML or in people’s headsA command exists but skips a gate CI runs, or needs manual setup firstOne command runs every CI gate from a clean clone and is named in AGENTS.md or CLAUDE.mdAlso has a scoped variant (changed package or files), prints failures in a form an agent can act on, and a hook or goal runs it before the agent may stop
2. Fast feedbackThe check the agent can run takes over 15 minutes, or runs only in CI5–15 minutesScoped check under 2 minutes; full check under 15Scoped check under 30 seconds; full check parallelized under 10 minutes
3. Deterministic testsKnown flaky tests, retries on by default, tests that reach the real network or wall clockFlakes are quarantined by hand; some tests still depend on time, order or networkThree consecutive runs and one shuffled-order run give identical results; clock, randomness and network are controlled in testsAlso: flaky results on the main branch are tracked, quarantined tests have an owner and an expiry date, unexpected network calls fail the test
4. Strict types and lintUntyped, or strict mode off with lint as warnings onlyTypes exist but strict mode is off, or type and lint errors do not fail CIStrict mode on ("strict": true in tsconfig.json, mypy --strict or your language’s equivalent) and both gates fail CIAlso: zero-warning policy, and every suppression (@ts-expect-error, # type: ignore, lint disables) needs a reason and is counted, with a ceiling that only goes down
5. Module boundariesNo stated layering; circular imports; small changes touch many unrelated modulesThe architecture is described in a document but nothing enforces itA tool in the check command fails on forbidden imports (for example dependency-cruiser for JavaScript and TypeScript, import-linter for Python)Also: each module exposes a public interface with its own tests that run in isolation, and ownership is mapped (CODEOWNERS)
6. Seeded data and setupTests or the dev server need a shared database, a teammate’s credentials or production dataA seed script exists but is run by hand and drifts from the schemaOne bootstrap command installs, migrates and loads deterministic fixtures from a clean clone, with no production secretsAlso: per-run isolation (a port block and a database per worktree or container), fixtures versioned with the migrations, fakes for third-party services

The total is out of 18, but the minimum matters more than the sum, because one weak property undermines the others. A fast suite that is flaky teaches the agent to ignore red. This page recommends three thresholds:

ProfileWhat it meansRuns this supports
Property 1 or 3 at 0Nothing can reliably say “no”Interactive work only; you read the diff
Every property at 1 or more, at least one still at 1Checks catch some defects, and you know whichInteractive and supervised background runs; you review the risky areas by reading
Every property at 2 or more (total 12 or more)Checks decide most outcomesBackground and overnight runs on low-risk tickets; you review evidence, not diffs

Autonomy is decided per change, not per repository, so pair the score with the ticket’s blast radius and reversibility on shaping a backlog for agents, and with the run posture on permissions, sandboxes and approval modes.

The audit is itself agent output, so the agent must show evidence and you rerun what it cites. Run it read-only.

  1. Run the audit prompt below in your usual tool (see the tabs for each). Allow it to run the check, test and type commands, because timing and rerunning them is the point.

  2. Verify the scorecard, not the prose. For each score of 2 or 3, rerun the cited command and compare result and timing. This catches the most common audit failure: a score awarded because a script with the right name exists.

  3. Commit the scorecard as docs/agent-readiness.md with the date, the tool and model that produced it, and the thresholds you used. The tech lead owns it; the next audit is diffed against it.

  4. Fix in this order: one-command check, determinism, speed, types, boundaries, data. The check makes every later fix verifiable; a flaky check makes every other signal untrustworthy.

  5. Run the canary test (below) before you raise a repository to background runs. The score says the gates exist; the canary proves they catch defects.

  6. Re-run the audit after any change to the toolchain, the CI pipeline or the default model, and at least once a quarter. Diff it against the committed scorecard.

Interactive: paste the prompt into a session started in Manual mode (claude --permission-mode manual) so you approve each command the audit runs, or into plan mode (/plan) if you only want a static read and will time the commands yourself.

Headless, for a scripted audit you can diff each quarter, pre-approve only the read tools and the commands the audit needs. claude -p starts in Manual, so anything you did not list is refused rather than prompted (checked against Claude Code 2.1.283):

Terminal window
# terminal, repository root; readiness-audit.txt holds the prompt above
claude -p "$(cat readiness-audit.txt)" \
--allowedTools "Read,Grep,Glob,Bash(git log *),Bash(pnpm check),Bash(pnpm check:changed),Bash(pnpm lint),Bash(pnpm test),Bash(pnpm typecheck),Bash(time pnpm check),Bash(time pnpm check:changed)" \
> docs/agent-readiness.new.md

Bash(time pnpm check) and Bash(time pnpm check:changed) let the audit time the two checks and nothing else; Bash(pnpm test) stays exact so the audit cannot pass extra arguments such as a snapshot update. Reject a property 2 score with no timing. Replace the pnpm commands with your repository’s.

The first fix is almost always a single entry point that runs exactly what CI runs, plus a scoped variant for the agent’s inner loop. For a TypeScript monorepo with pnpm and Vitest:

{
"scripts": {
"typecheck": "tsc --noEmit",
"lint": "eslint . --max-warnings 0",
"boundaries": "depcruise src --ignore-known",
"test": "vitest run",
"check": "pnpm typecheck && pnpm lint && pnpm boundaries && pnpm test",
"check:changed": "pnpm typecheck && vitest run --changed origin/main"
}
}

Then make CI call pnpm check and nothing else, so the local command and the gate cannot drift. (In Python: a Makefile target running ruff check, mypy --strict, lint-imports and pytest.)

Finally, name the command in the context file:

## Checks
Run `pnpm check:changed` after each change and `pnpm check` before you say you are done.
Both must exit 0. Do not edit test files, tsconfig.json or .dependency-cruiser.js to make them pass.

A line in the context file is advice. To make the check non-negotiable, run it at the point where the agent tries to finish.

A Stop hook (hooks as deterministic guardrails covers the event model) runs every time Claude tries to end its turn; Stop takes no matcher. Exit code 2 prevents Claude from stopping and continues the conversation, so print the failing output to stderr for Claude to act on (per the Claude Code hooks reference, checked 2026-09-26). Commit it in .claude/settings.json:

{
"hooks": {
"Stop": [
{
"hooks": [{ "type": "command", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/require-check.sh" }]
}
]
}
}
.claude/hooks/require-check.sh
#!/usr/bin/env bash
input=$(cat)
# Avoid an endless loop: if a Stop hook already kept Claude going, let it stop.
if echo "$input" | grep -q '"stop_hook_active": *true'; then exit 0; fi
if ! out=$(pnpm check:changed 2>&1); then
echo "pnpm check:changed failed. Fix the cause, not the test:" >&2
echo "$out" | tail -n 60 >&2
exit 2
fi

Make it executable: chmod +x .claude/hooks/require-check.sh.

The loop guard matters: without it, a check the agent cannot fix keeps blocking until Claude Code overrides the hook after more than eight consecutive blocks (default cap, CLAUDE_CODE_STOP_HOOK_BLOCK_CAP) and ends the turn. With it, Claude gets one retry; if the check is still red, the turn ends and CI is the backstop. For long runs, a /goal condition does the same job.

A flaky test does more damage to an agent than to a person. You remember that checkout.spec.ts is flaky; the agent sees red, rewrites working code or edits the test until it passes. So quarantine flakes before you tune speed.

Once the suite is deterministic, speed is usually a matter of scope: run only what the change can affect (vitest run --changed origin/main, which covers committed, staged and unstaged changes on the branch plus new files Git does not ignore, or per-package test targets in a monorepo), keep slow end-to-end suites in the full check and in CI, and parallelize the full check. Unit testing with agents and test data management cover the test-side patterns in depth.

Turn the architecture into enforced boundaries

Section titled “Turn the architecture into enforced boundaries”

Agents follow the path of least resistance through your imports: if the UI layer can import the database client, eventually one change will. A boundary rule turns that review comment into a failing check.

The install commands the prompt expects are pnpm add -D dependency-cruiser (npm i -D dependency-cruiser with npm; version 18.4.0 on npm, checked 2026-09-26) and pip install import-linter or uv add --dev import-linter (version 2.15 on PyPI, checked 2026-09-26).

Strict types follow the same ratchet logic. Turning on "strict": true in a large TypeScript codebase can produce thousands of errors at once, so enable it per directory (a stricter tsconfig.json for new or migrated packages) and count suppressions in the check so the number can only fall. Large codebases covers splitting the work across agent sessions.

Prove the checks catch defects with a canary test

Section titled “Prove the checks catch defects with a canary test”

A high score says the gates exist and run. It does not say they catch anything. Before you move a repository to background runs, plant known defects and confirm the check turns red for each one. This is a small, manual form of mutation testing; oracle strength covers the automated version.

Any defect that survives marks a score that is too high: lower it and fix the gap before raising autonomy. Repeat the canary quarterly; a check that caught five defects in March can lose one when someone adds --passWithNoTests in June.

Who signs off. The tech lead owns the scorecard and the canary result. A repository moves up an autonomy profile only when the scorecard meets the threshold above and the latest canary caught every planted defect. After that, changes are accepted on their evidence, as described in evidence, not diffs.

What breaks when you make a codebase agent-ready?

Section titled “What breaks when you make a codebase agent-ready?”

The agent makes the check green by weakening the check. It adds @ts-expect-error, marks a test .skip or edits the lint or boundary config; once the check is enforced, it is the goal. Recovery: deny agent writes to test, lint, boundary and CI configs; count suppressions; route those files to human review with CODEOWNERS. The full pattern is on protecting the oracle.

The audit reports scores it did not measure. A test:e2e script that has not passed in months earns a 2. Recovery: reject any row without an exit code and timing, and rerun the cited commands (audit step 2).

Strict mode produces thousands of errors, and the effort stalls. Recovery: ratchet one package at a time, baseline existing boundary violations, and let the agent burn down one module per session.

The scoped check passes and CI fails. check:changed missed a consumer in another package. Recovery: keep the full check in the Stop hook or goal for any change touching a shared package, and add the missed dependency path to the scoped selection.

Quarantined tests become a graveyard. Twenty skipped tests, no owner, and the agent adds a twenty-first. Recovery: every skip has an owner and an expiry date; expired skips cap determinism at 1.

Seed data drifts from the schema. Recovery: the bootstrap command migrates a fresh database before seeding, inside the full check, so drift fails instead of hiding.

Two agents share one database or one port. Recovery: one worktree and one port block per agent; see parallel agents in worktrees and ephemeral environments.

Where to go next after the readiness audit

Section titled “Where to go next after the readiness audit”

The audit tells you where the repository stands. The harness overview is the prerequisite; tech leads continue with shared agent rules, developers with fitness functions.