Characterization tests: pin legacy behavior before agents touch it
Characterization tests record what legacy code does today, not what it should do, and fail when that output changes. Snapshot, approval, and golden-master tests at a seam turn untested code into an oracle an agent can refactor against. Record them on the untouched code, prove they can fail with mutation testing, then lock them before the agent starts.
You need an agent to split a 1,400-line statement.js that nobody has touched since 2019. It has no tests, it prints customer statements, and finance reconciles against those statements every month. If you ask the agent to “refactor and add tests”, it writes tests that describe its own refactored code, and they pass whether or not the monthly fee is still charged the same way. This page is for the developer who has to hand that refactor to an agent and still sign off on it without reading every line.
What you get from pinning legacy behavior first
Section titled “What you get from pinning legacy behavior first”- A decision table for choosing between snapshot, approval, golden-master, and combination tests, by the seam you have.
- An eight-step procedure that records goldens on untouched code, removes nondeterminism, proves the goldens can fail, and locks them.
- Two tested harnesses: combination approvals with ApprovalTests in Python and file snapshots with a frozen clock in Vitest, both run on 26 September 2026.
- Four copy-paste prompts: a seam inventory, a recording task, a quirk audit, and a refactor under the locked oracle.
- Tool setups that keep the agent out of the legacy code while it records, and out of the goldens while it refactors, in Claude Code, Codex, and Cursor.
Why can’t the agent write the tests after the refactor?
Section titled “Why can’t the agent write the tests after the refactor?”A test is an oracle only if it was written against something other than the code it judges. Legacy code has no spec, so the only independent source of truth is the running system. Michael Feathers named the technique in Working Effectively with Legacy Code (2004): a characterization test describes the actual behavior of a piece of code, and it stays the reference until someone decides that behavior should change.
Agents make the order of operations decisive. An agent that refactors first and tests second writes assertions that match the refactored output, so a dropped fee or a changed rounding rule becomes “expected”. An agent that records first, on a commit it did not change, produces goldens that the refactor must reproduce. The difference is not the test code; it is which commit the goldens came from.
Characterization tests pin behavior, including bugs. That is the point. A refactor should change structure, not behavior, and every bug you find while recording goes on a quirks list to be fixed later in its own, deliberate change.
Which kind of characterization test fits your seam?
Section titled “Which kind of characterization test fits your seam?”A seam is a place where you can observe the code’s behavior without editing the code: a function’s return value, an HTTP response, a batch job’s output file, the rows a job writes. Choose the highest seam that is deterministic and fast enough to run on every pull request. A high seam survives the refactor; a low seam pins the implementation the agent is supposed to change.
| Form | What it pins | Best seam | Tooling (checked 2026-09-26) |
|---|---|---|---|
| File snapshot | One serialized output per input case, stored as a file you can diff | Pure-ish functions, renderers, API responses | Vitest toMatchFileSnapshot (5.0.2), Jest snapshots (30.5.2), syrupy for pytest (6.1.1) |
| Approval test | A received output compared with an approved file; a human approves by promoting the file | Text reports, documents, anything a human can judge by reading a diff | ApprovalTests: approvaltests 19.1.1 (PyPI), approvals 7.3.0 (npm) |
| Combination approval | Every combination of a few input values, in one approved file | Functions with a handful of parameters and many branches (pricing, tax, eligibility) | verify_all_combinations in approvaltests 19.1.1 |
| Golden master | The whole output of a run over a large, fixed input set | Batch jobs, exports, CLI tools, HTTP endpoints recorded against a deployment | Any runner plus a diff; the HTTP version is on legacy modernization |
When you have more than one seam, pin both: a golden master at the edge proves the system still answers the same, and a combination approval on the hot function tells you where it changed.
Pin legacy behavior step by step
Section titled “Pin legacy behavior step by step”Do the recording as its own pull request, merged before any refactor starts. The agent can do most of the typing; you own the seam choice, the quirks list, and the lock.
-
Map the seams, read-only. Ask the agent for an inventory of entry points, side effects, and sources of nondeterminism in the module (first prompt below). Run it in plan mode so nothing is edited. Pick one seam per slice of work.
-
Neutralize nondeterminism without editing production logic. Freeze the clock, seed or stub randomness, fix the locale and time zone, and sort anything whose order is not guaranteed. Prefer test-side controls (fake timers, a stubbed
Math.random,TZ=UTC) over code changes. If the code reads the clock deep inside, the only production edit allowed in this pull request is a behavior-free seam, such as anowparameter with the old default. -
Collect inputs that exercise the branches. Start from real inputs: anonymized production requests, the support team’s list of odd accounts, and month-end and year-end dates. Add boundaries for every comparison you can see (zero, negative, the threshold, just past it). Run with branch coverage on the seam and add inputs until every reachable branch runs. Vitest needs a coverage provider at the same version as Vitest itself (the provider’s peer dependency pins the exact version):
Terminal window # Terminal (@vitest/coverage-v8 5.0.2 matches Vitest 5.0.2, checked 2026-09-26)npm install --save-dev @vitest/coverage-v8@5.0.2npx vitest run --coverage --coverage.include=src/legacy/statement.js -
Record the goldens on the untouched commit. Generate the snapshot or approved files from the code as it is on the default branch. Commit the harness, the inputs, and the goldens together, with the commit hash they were recorded on in the pull request description.
-
Read the goldens once, for quirks. Skim each golden for output that looks wrong: a fee charged on a zero amount, a negative balance with the wrong sign, a
nullcountry that is priced as foreign. Do not fix any of them now. Write each one in aQUIRKS.mdnext to the goldens, with an owner who decides later. -
Prove the goldens can fail. Run mutation testing on the legacy module only. Every surviving mutant is either an input you have not recorded or an equivalent mutant; add inputs for the first kind and note the second. Details and thresholds are on how strong is your oracle.
Terminal window # Terminal, JavaScript or TypeScript (StrykerJS 10.0.0, checked 2026-09-26)npm install --save-dev @stryker-mutator/core @stryker-mutator/vitest-runnernpx stryker run --testRunner vitest --mutate src/legacy/statement.js# Terminal, Python (mutmut 3.8.0 on PyPI, checked 2026-09-26)pip install mutmutFor mutmut, point
source_pathsat the legacy package and the test selection at the characterization suite inpyproject.toml, then runmutmut run:# pyproject.toml (keys from the mutmut 3.8.0 README)[tool.mutmut]source_paths = ["legacy/"]pytest_add_cli_args_test_selection = ["tests/characterization/"] -
Lock the goldens. Put the goldens directory under
CODEOWNERS, deny the agent edit access to it (tool tabs below), and add the CI guard below as a required check. The recording pull request itself changestests/characterization/, so the maintainer labels itbehavior-change(or you merge the guard in a follow-up pull request right after the recording pull request). Every later pull request runs the guard unlabeled. From here on, a refactor that changes a golden fails by rule.QUIRKS.mdsits inside the guarded directory, so a quirk you add or update after the lock also goes through a pull request labeledbehavior-change. -
Hand the refactor to the agent. The definition of done is “characterization suite green, goldens unchanged”. The refactor pull request carries the suite result and the guard result in its evidence bundle.
Pin a many-branch function with combination approvals
Section titled “Pin a many-branch function with combination approvals”Pricing, tax, and eligibility functions have few parameters and many branches. A combination approval records every combination of a few values per parameter in one file, which a human can skim in a minute. This is the legacy function:
def shipping_cost(weight_kg, country, express): if weight_kg <= 0: raise ValueError("weight must be positive") base = 4.99 if country == "PL" else 9.99 if weight_kg > 20: base += (weight_kg - 20) * 0.5 if express: base *= 1.8 return round(base, 2)The characterization test chooses values on both sides of every comparison, plus junk the legacy code accepts today:
from approvaltests import verify_all_combinations
from legacy.shipping import shipping_cost
def test_shipping_cost_pinned(): verify_all_combinations( shipping_cost, [ [-1, 0, 0.1, 20, 20.01, 35], # boundaries around 0 and 20 ["PL", "DE", "", None], # home, abroad, and junk the legacy code accepts [True, False], ], )The first pytest run fails and writes test_shipping_char.test_shipping_cost_pinned.received.txt with one line per combination, 48 in total. Exceptions are recorded as output, so the ValueError branch is pinned too:
args: (-1, 'PL', True) => ValueError('weight must be positive')...args: (35, None, True) => 31.48args: (35, None, False) => 17.49You approve by renaming the file to .approved.txt and committing it. The next run passes. Two details from our run with approvaltests 19.1.1 and pytest: the first run also creates an empty .approved.txt, so never commit a zero-byte approved file; and None as a country is priced as abroad, which is a quirk for QUIRKS.md, not something to fix during the refactor.
We then mutated the function by hand. Changing weight_kg <= 0 to < 0 failed the test, because 0 is in the inputs. Changing the express multiplier from 1.8 to 1.9 failed it. Changing weight_kg > 20 to >= 20 passed: at exactly 20 kg the surcharge is (20 - 20) * 0.5, which is zero either way. That is an equivalent mutant, and triaging it is the difference between a mutation score and a guess.
Pin a renderer with file snapshots and a frozen clock
Section titled “Pin a renderer with file snapshots and a frozen clock”When the output is a document, a file snapshot per input case keeps each golden small enough to review. This legacy renderer reads the clock and Math.random, so the harness freezes both instead of editing the code:
import { readdirSync, readFileSync } from 'node:fs';import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest';import { renderStatement } from '../../src/legacy/statement.js';
const dir = new URL('./cases/', import.meta.url);const cases = readdirSync(dir).filter((f) => f.endsWith('.json')).sort();
describe('renderStatement: pinned legacy behavior', () => { beforeEach(() => { vi.useFakeTimers(); vi.setSystemTime(new Date('2026-02-01T09:00:00Z')); // freeze the clock vi.spyOn(Math, 'random').mockReturnValue(0.123456789); // freeze the random reference }); afterEach(() => { vi.useRealTimers(); vi.restoreAllMocks(); });
it.each(cases)('%s', async (file) => { const account = JSON.parse(readFileSync(new URL(file, dir), 'utf8')); await expect(renderStatement(account)).toMatchFileSnapshot(`./golden/statement/${file}.txt`); });});Each JSON file in cases/ is one account. On a developer machine, npx vitest run writes each missing golden. With CI=true set, the same command fails every test whose golden is missing instead of creating it (tested with Vitest 5.0.2). That default is what stops an empty oracle from passing in a pipeline. A golden written from a basic-tier account looks like this:
Statement for Bo (2026-02-01)2026-01-02 -0.01 fee 1.502026-01-03 0.00Balance: -1.51Ref: 4fzzzxjyA one-cent debit costs a 1.50 fee: a quirk for the list, not a fix for the refactor. When we changed the fee condition from tx.amount < 0 to tx.amount <= 0, the case with a zero amount failed, which is why that case is in the inputs.
vitest run -u accepts true, "new", "all", or "none" in 5.0.2. Any update rewrites goldens, so it is a behavior change and never part of a refactor branch.
Scrub nondeterminism without hiding behavior
Section titled “Scrub nondeterminism without hiding behavior”When you cannot freeze a value at the source, replace it in the output before comparison. ApprovalTests calls this a scrubber:
from approvaltests import Options, verifyfrom approvaltests.scrubbers import combine_scrubbers, create_regex_scrubber
from legacy.report import legacy_report # your legacy function under test
scrubber = combine_scrubbers( create_regex_scrubber(r"\d{4}-\d{2}-\d{2}T[\d:.]+", "<timestamp>"), create_regex_scrubber(r"Ref: \w+", "Ref: <ref>"),)verify(legacy_report(), options=Options().with_scrubber(scrubber))Check the received file after the first run, every time. In our test with approvaltests 19.1.1, the built-in scrub_all_dates left an ISO 8601 timestamp with microseconds (2026-09-26T12:14:35.785033) untouched, so the golden would have failed on the next run. A scrubber you did not look at is a guess.
Scrub only values that differ between two runs of the same code: timestamps, request IDs, random references. If you scrub a field because the refactored code formats it differently, you have decided that difference does not matter. That decision goes on the quirks list with an owner, not into a regular expression.
Keep the goldens out of the agent’s reach in CI
Section titled “Keep the goldens out of the agent’s reach in CI”This guard runs on every pull request, whichever tool opened it. It fails when a pull request changes a golden, unless a maintainer has labeled the pull request as a deliberate behavior change:
name: characterization-guardon: pull_request: types: [opened, synchronize, reopened, labeled, unlabeled]
permissions: contents: read
jobs: goldens-unchanged: if: ${{ !contains(github.event.pull_request.labels.*.name, 'behavior-change') }} runs-on: ubuntu-latest steps: - uses: actions/checkout@v7 with: fetch-depth: 0 persist-credentials: false - name: Fail if this pull request changes pinned behavior env: BASE_REF: ${{ github.base_ref }} run: | changed=$(git diff --name-only "origin/$BASE_REF...HEAD" -- tests/characterization/) if [ -n "$changed" ]; then printf 'Pinned behavior changed:\n%s\n' "$changed" >&2 echo "Split it into its own pull request labeled behavior-change." >&2 exit 1 fiOnly people with triage or write access can add a label, and a CODEOWNERS entry for tests/characterization/ still requires the oracle owner’s review on a labeled pull request. Make this job and the characterization run required checks in the branch ruleset.
Copy-paste prompts for characterization tests
Section titled “Copy-paste prompts for characterization tests”How do you run this in Claude Code, Codex, and Cursor?
Section titled “How do you run this in Claude Code, Codex, and Cursor?”The harnesses, prompts, and CI guard are identical in all three tools. What differs is how you enforce the two locks: during recording the agent may not edit production code, and during the refactor it may not edit the goldens. Claude Code and Codex commands were checked against Claude Code 2.1.283 and Codex CLI 0.157.1 on 2026-09-26; Cursor features were checked on cursor.com on 2026-08-28.
Recording phase. Run the seam inventory in plan mode (/plan), then record with production code denied in the committed .claude/settings.json. A leading / resolves to the project root:
{ "permissions": { "deny": ["Edit(/src/**)", "Edit(/.github/**)", "Edit(/.claude/**)"] }}Refactor phase. Swap the rule in a separate commit so the goldens are locked instead: "deny": ["Edit(/tests/characterization/**)", "Edit(/.github/**)", "Edit(/.claude/**)"]. Edit deny rules cover the edit tools and the shell writes Claude Code recognizes, such as sed and > redirects, but not every script the agent can write and run; when the agent has shell access, add the sandbox and the hook from protecting the oracle.
Headless recording. A claude -p session starts in Manual permission mode, so grant only what recording needs:
# Terminal, from the repository rootclaude -p "$(cat prompts/record-characterization.md)" \ --allowedTools "Read,Grep,Glob,Write,Edit,Bash(npx vitest run *)" \ --max-budget-usd 3The deny rule in .claude/settings.json still applies, so Write and Edit reach the tests but not src/.
Recording phase. Define a permission profile in ~/.codex/config.toml that makes production code under src/ read-only inside the sandbox. Directories, exact file paths, and a trailing /** work; any other glob with read is rejected in 0.157.1:
[permissions.record-only]extends = ":workspace"
[permissions.record-only.filesystem.":project_roots"]"src" = "read"".github" = "read"".codex" = "read"# Terminal, from the repository root (Codex CLI 0.157.1)codex exec -c default_permissions=record-only "$(cat prompts/record-characterization.md)"Refactor phase. Use a second profile with "tests/characterization" = "read" in place of "src", and select it the same way. Permission profiles are beta in Codex CLI 0.157.1 and do not compose with --sandbox; do not combine them. The profile syntax is the one tested on protecting the oracle.
Recording phase. Start the seam inventory in Plan Mode, which plans before it writes any code. Add a project Rule for the recording task: “never edit anything under src/; record behavior, never fix it”.
Refactor phase. Replace the Rule with “never edit tests/characterization/ or QUIRKS.md; if a golden looks wrong, stop and report it”. A Rule is an instruction, not an enforcement, and Cloud Agents run in their own VMs. Cursor’s Hooks can observe, block, or modify agent behavior, but as of 2026-08-28 we could not verify a hook event that blocks a file edit, so we do not show one: the characterization-guard job and the CODEOWNERS entry are the only binding controls on Cursor’s runs. Cursor’s limits are summarized on protecting the oracle. The editor-side workflow for one legacy module is on refactoring legacy code with Cursor.
How do you know the characterization suite is good enough?
Section titled “How do you know the characterization suite is good enough?”You do not read the goldens line by line, apart from the one quirk pass. You check evidence that a machine produced:
| Check | Evidence | Who signs off |
|---|---|---|
| Recorded on untouched code | The recording commit’s parent is on the default branch, and git diff on src/ in the recording pull request is empty or behavior-free | Reviewer of the recording pull request |
| Deterministic | Two consecutive runs with CI=true pass with no golden changes | CI |
| Branch coverage of the seam | Coverage report on the legacy module; every uncovered branch is listed and explained | Developer who owns the slice |
| Can fail | Mutation score on the legacy module at or above the team’s floor, every survivor triaged as missing input or equivalent | Developer; floor set by the tech lead |
| Quirks owned | QUIRKS.md lists every suspicious output with an owner and a decision date | Product or domain owner |
| Locked | CODEOWNERS entry, agent deny rule or profile, and the guard job as a required check | Tech lead |
Only when all six hold does the suite count as the oracle for an agent refactor. The refactor pull request then needs no line-by-line review of the moved code for behavior: the suite result, the guard result, and the mutation score go in its evidence bundle, and a reviewer reads those.
What breaks when you characterize legacy code?
Section titled “What breaks when you characterize legacy code?”The goldens were recorded after the change. The agent refactored first, then generated snapshots, and everything is green. Recovery: check out the default-branch commit, rerun the suite against it, and treat any failure as a behavior change the refactor introduced. Make “recorded on commit X” a required line in the recording pull request.
Someone ran the update flag. A refactor branch contains vitest -u or --snapshot-update output, and the guard was bypassed with the label. Recovery: revert the golden changes, rerun against the refactored code, and split any genuine behavior change into its own labeled pull request that the oracle owner approves.
The goldens are flaky. They pass on one machine and fail in CI because of the time zone, the locale, map iteration order, or a timestamp the scrubber missed. Recovery: run the suite twice with CI=true and TZ=UTC, diff the two runs, and freeze or scrub each difference at its source.
The scrubber hides real behavior. A mask added to “fix” a flaky golden also swallows a field the refactor changed. Recovery: every scrubber carries a comment saying which run-to-run difference it removes; the quirk-audit prompt lists what each one could hide.
The seam is too low. Tests pin private helpers or call order, so every structural change breaks them and the agent is tempted to regenerate. Recovery: move the seam up to a public function, an HTTP response, or an output file, and delete the low-level goldens once the high ones cover the same branches.
The goldens are too big to review. One 30,000-line golden master hides a changed line in noise. Recovery: split the output per input case, as the Vitest harness does, so a failure names the case and the diff fits on a screen.
A pinned bug becomes the spec forever. Nobody owns QUIRKS.md, so wrong behavior survives the rewrite. Recovery: each quirk gets an owner and a date; fixing one is a separate pull request labeled behavior-change, with its golden updated and approved on purpose.
Where to go next with characterization tests
Section titled “Where to go next with characterization tests”Frequently asked questions
What is a characterization test?
A characterization test records what existing code does today, not what it should do, and fails when that output changes. Snapshot, approval, and golden-master tests are the usual forms. They give untested legacy code an oracle before an agent refactors it.
Why not let the agent write the tests after it refactors?
Tests written against the new code describe the new code, bugs included, and pass by construction. Characterization tests are recorded on the untouched code in a separate change, so they describe the behavior users depend on, and the refactor has to reproduce it.
How do you know a characterization suite is strong enough?
Every golden was recorded on the untouched commit, the inputs cover the branches of the seam, mutation testing on the legacy module kills the mutants that matter, every survivor is triaged, and the goldens are locked against agent edits in the session and in CI.