Making a codebase agent-ready
An agent-ready codebase is a repository in which automated checks, not a human reading the diff, can decide whether an agent’s change is correct. Six properties make that possible: one command that runs every gate, fast feedback, deterministic tests, strict types and lint, enforced module boundaries, and seeded, reproducible data. Each property is scored 0–3 with an audit prompt.
This page is for developers who want to stop reading every line an agent writes, and for tech leads deciding which repositories are ready for background runs. You gave the agent a well-specified ticket. It reported “all tests pass”, and you still spent 40 minutes reading the diff, because the test command it ran skipped the integration suite, two tests fail one run in five anyway, and nothing would have stopped it importing the billing module from the UI layer. The ticket was fine; the repository could not prove anything.
What you’ll walk away with from an agent-readiness audit
Section titled “What you’ll walk away with from an agent-readiness audit”- A 0–3 scoring rubric for six properties, with the evidence that earns each score.
- A copy-paste audit prompt that makes the agent measure, not guess.
- The order to fix things in, and three prompts that do the fixing: the one-command check, flaky tests, and enforced boundaries.
- A canary test that proves your checks catch defects before unattended runs.
- The failure modes, above all the agent that makes the check green by weakening it.
Why does the codebase decide whether you can stop reading diffs?
Section titled “Why does the codebase decide whether you can stop reading diffs?”An agent’s claim that its work is done is worth exactly as much as the check behind it. If the only way to know a change is correct is to read it, you read it, no matter how good the model is. Every property on this page converts one kind of reading into a check that runs without you.
The strongest third-party evidence for this comes from DORA. The 2025 DORA report (Google Cloud blog, Nathen Harvey and Derek DeBellis, 2025-09-23) found a positive relationship between AI adoption and delivery throughput, and a negative one with delivery stability. Its explanation: “Without robust control systems, like strong automated testing, mature version control practices, and fast feedback loops, an increase in change volume leads to instability. Teams working in loosely coupled architectures with fast feedback loops see gains, while those constrained by tightly coupled systems and slow processes see little or no benefit.” The same report’s summary line is that AI “amplifies what’s already there”.
Stripe shows the other end: its Minions merge more than 1,300 pull requests a week with no human-written code (vendor-internal figure, Alistair Gray, stripe.dev, 2026-02-19), a loop that stops after “at most two rounds of CI” (Minions Part 1, stripe.dev, 2026-02-09) and runs against “over three million” tests.
In harness terms, the codebase sits next to the seven harness layers rather than inside them. The harness bounds what the agent can do; the codebase decides whether anything can check what it did.
Which six properties make a codebase agent-ready?
Section titled “Which six properties make a codebase agent-ready?”| # | Property | What the agent does without it | The check it makes possible |
|---|---|---|---|
| 1 | One-command check | Guesses the test command, runs the unit suite only, and reports “tests pass” | “Done” means one named command exited 0, and CI runs the same command |
| 2 | Fast feedback | Skips the slow suite, or batches 30 edits before the first run, so failures arrive late and tangled | The agent runs the check after every meaningful edit and fixes failures while the cause is still local |
| 3 | Deterministic tests | Learns that red is sometimes noise, retries until green, or “fixes” a flake by editing the test | A red result always means a defect, so a green result means something |
| 4 | Strict types and lint as gates | Passes any, None or an unchecked index through, and invents a method the compiler would have rejected | Whole classes of error (wrong shape, missing field, invented API) fail in seconds without a test |
| 5 | Enforced module boundaries | Imports whatever makes the change work, so the diff is correct today and the architecture erodes | A layer violation fails the check, so you review the interface, not every import |
| 6 | Seeded, reproducible data and setup | Needs a shared database, your credentials or a manual step, so it either cannot verify or verifies against the wrong state | A clean clone reaches a green check with one bootstrap command, locally, in a worktree or in a cloud environment |
Two things are deliberately not on the list. Test coverage is not: a high number with weak assertions passes wrong code, and the oracle-strength page covers how to measure what tests actually catch. A context file (AGENTS.md, CLAUDE.md, Cursor Rules) is not either: it names the command, but a repository without the properties above cannot be fixed by describing them. See concise AGENTS.md and CLAUDE.md files for what belongs there.
How do you score each property from 0 to 3?
Section titled “How do you score each property from 0 to 3?”Score each property on evidence you can point at: a command, its output, its run time, a config file. A score without evidence is a 0. The thresholds for run times are this page’s recommendation for an agent’s inner loop, not a published benchmark; adjust them once, write your version into the scorecard, and keep them fixed so scores stay comparable over time.
| Property | 0 | 1 | 2 | 3 |
|---|---|---|---|---|
| 1. One-command check | No single command; the gates live only in CI YAML or in people’s heads | A command exists but skips a gate CI runs, or needs manual setup first | One command runs every CI gate from a clean clone and is named in AGENTS.md or CLAUDE.md | Also has a scoped variant (changed package or files), prints failures in a form an agent can act on, and a hook or goal runs it before the agent may stop |
| 2. Fast feedback | The check the agent can run takes over 15 minutes, or runs only in CI | 5–15 minutes | Scoped check under 2 minutes; full check under 15 | Scoped check under 30 seconds; full check parallelized under 10 minutes |
| 3. Deterministic tests | Known flaky tests, retries on by default, tests that reach the real network or wall clock | Flakes are quarantined by hand; some tests still depend on time, order or network | Three consecutive runs and one shuffled-order run give identical results; clock, randomness and network are controlled in tests | Also: flaky results on the main branch are tracked, quarantined tests have an owner and an expiry date, unexpected network calls fail the test |
| 4. Strict types and lint | Untyped, or strict mode off with lint as warnings only | Types exist but strict mode is off, or type and lint errors do not fail CI | Strict mode on ("strict": true in tsconfig.json, mypy --strict or your language’s equivalent) and both gates fail CI | Also: zero-warning policy, and every suppression (@ts-expect-error, # type: ignore, lint disables) needs a reason and is counted, with a ceiling that only goes down |
| 5. Module boundaries | No stated layering; circular imports; small changes touch many unrelated modules | The architecture is described in a document but nothing enforces it | A tool in the check command fails on forbidden imports (for example dependency-cruiser for JavaScript and TypeScript, import-linter for Python) | Also: each module exposes a public interface with its own tests that run in isolation, and ownership is mapped (CODEOWNERS) |
| 6. Seeded data and setup | Tests or the dev server need a shared database, a teammate’s credentials or production data | A seed script exists but is run by hand and drifts from the schema | One bootstrap command installs, migrates and loads deterministic fixtures from a clean clone, with no production secrets | Also: per-run isolation (a port block and a database per worktree or container), fixtures versioned with the migrations, fakes for third-party services |
What the total means for autonomy
Section titled “What the total means for autonomy”The total is out of 18, but the minimum matters more than the sum, because one weak property undermines the others. A fast suite that is flaky teaches the agent to ignore red. This page recommends three thresholds:
| Profile | What it means | Runs this supports |
|---|---|---|
| Property 1 or 3 at 0 | Nothing can reliably say “no” | Interactive work only; you read the diff |
| Every property at 1 or more, at least one still at 1 | Checks catch some defects, and you know which | Interactive and supervised background runs; you review the risky areas by reading |
| Every property at 2 or more (total 12 or more) | Checks decide most outcomes | Background and overnight runs on low-risk tickets; you review evidence, not diffs |
Autonomy is decided per change, not per repository, so pair the score with the ticket’s blast radius and reversibility on shaping a backlog for agents, and with the run posture on permissions, sandboxes and approval modes.
How to run the agent-readiness audit
Section titled “How to run the agent-readiness audit”The audit is itself agent output, so the agent must show evidence and you rerun what it cites. Run it read-only.
-
Run the audit prompt below in your usual tool (see the tabs for each). Allow it to run the check, test and type commands, because timing and rerunning them is the point.
-
Verify the scorecard, not the prose. For each score of 2 or 3, rerun the cited command and compare result and timing. This catches the most common audit failure: a score awarded because a script with the right name exists.
-
Commit the scorecard as
docs/agent-readiness.mdwith the date, the tool and model that produced it, and the thresholds you used. The tech lead owns it; the next audit is diffed against it. -
Fix in this order: one-command check, determinism, speed, types, boundaries, data. The check makes every later fix verifiable; a flaky check makes every other signal untrustworthy.
-
Run the canary test (below) before you raise a repository to background runs. The score says the gates exist; the canary proves they catch defects.
-
Re-run the audit after any change to the toolchain, the CI pipeline or the default model, and at least once a quarter. Diff it against the committed scorecard.
Interactive: paste the prompt into a session started in Manual mode (claude --permission-mode manual) so you approve each command the audit runs, or into plan mode (/plan) if you only want a static read and will time the commands yourself.
Headless, for a scripted audit you can diff each quarter, pre-approve only the read tools and the commands the audit needs. claude -p starts in Manual, so anything you did not list is refused rather than prompted (checked against Claude Code 2.1.283):
# terminal, repository root; readiness-audit.txt holds the prompt aboveclaude -p "$(cat readiness-audit.txt)" \ --allowedTools "Read,Grep,Glob,Bash(git log *),Bash(pnpm check),Bash(pnpm check:changed),Bash(pnpm lint),Bash(pnpm test),Bash(pnpm typecheck),Bash(time pnpm check),Bash(time pnpm check:changed)" \ > docs/agent-readiness.new.mdBash(time pnpm check) and Bash(time pnpm check:changed) let the audit time the two checks and nothing else; Bash(pnpm test) stays exact so the audit cannot pass extra arguments such as a snapshot update. Reject a property 2 score with no timing. Replace the pnpm commands with your repository’s.
Interactive: start the CLI with the same read-only profile, codex -c 'default_permissions=":read-only"', or pick it with /permissions in the TUI, then paste the prompt. The default on-request approval policy only prompts for actions that leave the sandbox; with the default profile Codex can write inside the workspace without asking.
Headless, codex exec reads the prompt from stdin when you pass - and writes the final message to a file with -o (checked against Codex CLI 0.157.1). The :read-only permission profile (beta) lets it run commands but not write files, which matches an audit that changes nothing; in codex exec nothing is prompted, so a command that needs to leave the sandbox fails and the audit should report it as unmeasured:
# terminal, repository rootcodex exec -c 'default_permissions=":read-only"' \ -o docs/agent-readiness.new.md - < readiness-audit.txtTest runners that write caches or coverage files can fail under :read-only. If yours do, switch to :workspace, which allows writes in the workspace; the prompt’s “Do not edit any file” is then the only thing preventing edits, so diff the tree afterwards.
Before pasting, set Cursor to ask before terminal commands; Cursor’s Run modes documentation (Agent security) describes the options, and the mode names are not re-verified here. Then paste the prompt into the Agent and approve each command as it appears. Save the resulting table as docs/agent-readiness.md yourself.
Cursor also has a CLI with a print mode (-p, --print) for headless use (per cursor.com, checked 2026-08-28). The executable name was not re-verified on 2026-09-26, so check Cursor’s CLI documentation before scripting it.
Build the one-command check first
Section titled “Build the one-command check first”The first fix is almost always a single entry point that runs exactly what CI runs, plus a scoped variant for the agent’s inner loop. For a TypeScript monorepo with pnpm and Vitest:
{ "scripts": { "typecheck": "tsc --noEmit", "lint": "eslint . --max-warnings 0", "boundaries": "depcruise src --ignore-known", "test": "vitest run", "check": "pnpm typecheck && pnpm lint && pnpm boundaries && pnpm test", "check:changed": "pnpm typecheck && vitest run --changed origin/main" }}Then make CI call pnpm check and nothing else, so the local command and the gate cannot drift. (In Python: a Makefile target running ruff check, mypy --strict, lint-imports and pytest.)
Finally, name the command in the context file:
## ChecksRun `pnpm check:changed` after each change and `pnpm check` before you say you are done.Both must exit 0. Do not edit test files, tsconfig.json or .dependency-cruiser.js to make them pass.Make “done” mean the check passed
Section titled “Make “done” mean the check passed”A line in the context file is advice. To make the check non-negotiable, run it at the point where the agent tries to finish.
A Stop hook (hooks as deterministic guardrails covers the event model) runs every time Claude tries to end its turn; Stop takes no matcher. Exit code 2 prevents Claude from stopping and continues the conversation, so print the failing output to stderr for Claude to act on (per the Claude Code hooks reference, checked 2026-09-26). Commit it in .claude/settings.json:
{ "hooks": { "Stop": [ { "hooks": [{ "type": "command", "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/require-check.sh" }] } ] }}#!/usr/bin/env bashinput=$(cat)# Avoid an endless loop: if a Stop hook already kept Claude going, let it stop.if echo "$input" | grep -q '"stop_hook_active": *true'; then exit 0; fiif ! out=$(pnpm check:changed 2>&1); then echo "pnpm check:changed failed. Fix the cause, not the test:" >&2 echo "$out" | tail -n 60 >&2 exit 2fiMake it executable: chmod +x .claude/hooks/require-check.sh.
The loop guard matters: without it, a check the agent cannot fix keeps blocking until Claude Code overrides the hook after more than eight consecutive blocks (default cap, CLAUDE_CODE_STOP_HOOK_BLOCK_CAP) and ends the turn. With it, Claude gets one retry; if the check is still red, the turn ends and CI is the backstop. For long runs, a /goal condition does the same job.
Codex has a Stop hook event among its 12 hook events, and hooks need persisted trust before they run; review them with /hooks (Codex CLI 0.157.1). Point the hook at the same check:changed command. For unattended runs, /goal with “pnpm check exits 0” as the stopping condition keeps Codex looping until its evaluator judges the condition met; a model makes that judgement, so CI still runs the command. Check the input fields in Codex’s own hooks documentation before reusing the Claude Code script from the other tab; do not assume the two payloads match.
Put the two check commands in a project Rule so the Agent loads them in every session. Cursor Hooks are “spawned processes that communicate over stdio using JSON in both directions” that “can observe, block, or modify behavior” (cursor.com, checked 2026-08-28); the event names were not re-verified on 2026-09-26, so pick the event that fires when the agent finishes from Cursor’s hooks reference and run check:changed there.
Fix determinism before speed
Section titled “Fix determinism before speed”A flaky test does more damage to an agent than to a person. You remember that checkout.spec.ts is flaky; the agent sees red, rewrites working code or edits the test until it passes. So quarantine flakes before you tune speed.
Once the suite is deterministic, speed is usually a matter of scope: run only what the change can affect (vitest run --changed origin/main, which covers committed, staged and unstaged changes on the branch plus new files Git does not ignore, or per-package test targets in a monorepo), keep slow end-to-end suites in the full check and in CI, and parallelize the full check. Unit testing with agents and test data management cover the test-side patterns in depth.
Turn the architecture into enforced boundaries
Section titled “Turn the architecture into enforced boundaries”Agents follow the path of least resistance through your imports: if the UI layer can import the database client, eventually one change will. A boundary rule turns that review comment into a failing check.
The install commands the prompt expects are pnpm add -D dependency-cruiser (npm i -D dependency-cruiser with npm; version 18.4.0 on npm, checked 2026-09-26) and pip install import-linter or uv add --dev import-linter (version 2.15 on PyPI, checked 2026-09-26).
Strict types follow the same ratchet logic. Turning on "strict": true in a large TypeScript codebase can produce thousands of errors at once, so enable it per directory (a stricter tsconfig.json for new or migrated packages) and count suppressions in the check so the number can only fall. Large codebases covers splitting the work across agent sessions.
Prove the checks catch defects with a canary test
Section titled “Prove the checks catch defects with a canary test”A high score says the gates exist and run. It does not say they catch anything. Before you move a repository to background runs, plant known defects and confirm the check turns red for each one. This is a small, manual form of mutation testing; oracle strength covers the automated version.
Any defect that survives marks a score that is too high: lower it and fix the gap before raising autonomy. Repeat the canary quarterly; a check that caught five defects in March can lose one when someone adds --passWithNoTests in June.
Who signs off. The tech lead owns the scorecard and the canary result. A repository moves up an autonomy profile only when the scorecard meets the threshold above and the latest canary caught every planted defect. After that, changes are accepted on their evidence, as described in evidence, not diffs.
What breaks when you make a codebase agent-ready?
Section titled “What breaks when you make a codebase agent-ready?”The agent makes the check green by weakening the check. It adds @ts-expect-error, marks a test .skip or edits the lint or boundary config; once the check is enforced, it is the goal. Recovery: deny agent writes to test, lint, boundary and CI configs; count suppressions; route those files to human review with CODEOWNERS. The full pattern is on protecting the oracle.
The audit reports scores it did not measure. A test:e2e script that has not passed in months earns a 2. Recovery: reject any row without an exit code and timing, and rerun the cited commands (audit step 2).
Strict mode produces thousands of errors, and the effort stalls. Recovery: ratchet one package at a time, baseline existing boundary violations, and let the agent burn down one module per session.
The scoped check passes and CI fails. check:changed missed a consumer in another package. Recovery: keep the full check in the Stop hook or goal for any change touching a shared package, and add the missed dependency path to the scoped selection.
Quarantined tests become a graveyard. Twenty skipped tests, no owner, and the agent adds a twenty-first. Recovery: every skip has an owner and an expiry date; expired skips cap determinism at 1.
Seed data drifts from the schema. Recovery: the bootstrap command migrates a fresh database before seeding, inside the full check, so drift fails instead of hiding.
Two agents share one database or one port. Recovery: one worktree and one port block per agent; see parallel agents in worktrees and ephemeral environments.
Where to go next after the readiness audit
Section titled “Where to go next after the readiness audit”The audit tells you where the repository stands. The harness overview is the prerequisite; tech leads continue with shared agent rules, developers with fitness functions.