Skip to content

Legacy modernization with agents: a program, not a prompt

Legacy modernization with coding agents works as a gated program, not a single prompt: pin today’s behavior with characterization tests, cut the system into slices behind seams, migrate them in parallel waves of isolated agent runs, compare old and new paths on live traffic, then cut over on recorded evidence. Agents write the code; pinned behavior decides what merges.

Your team owns a 12-year-old PHP monolith. Billing lives in it, coverage is thin, and the two people who understood the tax rules have left. The board wants it on a TypeScript service by next year. Someone already asked an agent to “migrate billing/ to TypeScript” and got a 900-file pull request nobody can review or dares to merge.

Typing speed is no longer the constraint. Anthropic’s Claude Opus 5.5 announcement (22 September 2026) quotes an unnamed tester who finished “a 680,000-line code migration in less than a day”, a vendor-reported anecdote that does not say whether the migrated system still charges the same tax on the same invoice. This page is about producing that proof.

It is for the tech lead who runs the program, the CTO who funds it and signs the cut-over, and the developers who run the waves.

What you’ll walk away with from an agent-run modernization

Section titled “What you’ll walk away with from an agent-run modernization”
  • A five-phase program with an owned exit gate per phase, and a slices.yaml plan agents and humans share.
  • A characterization harness at the HTTP seam (tested with Vitest 5.0.2, and with pytest plus syrupy 6.1.1) and a CI guard that keeps migration pull requests from changing the pinned behavior.
  • Wave setups for Claude Code, Codex and Cursor, a shadow-comparison helper, a cut-over checklist and four copy-paste prompts.

Why does “migrate this to TypeScript” fail as one prompt?

Section titled “Why does “migrate this to TypeScript” fail as one prompt?”

It fails for four reasons, none of them coding ability.

  1. There is no oracle. Legacy code rarely has tests. Tests the agent writes after the new code describe the new code, bugs included, and pass by construction (see oracle strength).
  2. The batch is too large to verify. DORA’s guidance on its AI capabilities model (Nathen Harvey and Allison Park, Google Cloud, 10 December 2025) puts it directly: “AI can easily generate massive blocks of code, which are hard to review and test. Enforcing the discipline of small batches counteracts this risk.”
  3. Undocumented behavior is load-bearing. Rounding quirks, an empty customer ID meaning “walk-in”, a report sorted in insertion order: someone depends on each, and a rewrite from reading the code drops them.
  4. Intent gets mixed. Porting, upgrading a framework and “fixing” a calculation in one change leaves no way to tell which one broke production.

Each phase ends with evidence a person can check without reading the migrated code. Agents do most of the typing; humans own the gates.

PhaseWhat agents doWhat humans ownExit gate (evidence)
0. MapInventory entry points, callers, data stores and side effectsScope, target architecture, slice boundariesslices.yaml approved by the tech lead
1. PinRecord characterization tests and goldens against the legacy systemWhich inputs are representative; which quirks are bugsGoldens committed, mutation score at or above the floor, oracle owner signs off
2. SeamAdd a routing point (facade, proxy route or flag) with no behavior changeWhere the seam sitsCharacterization suite green through the seam, zero golden changes
3. Migrate in wavesPort each slice in an isolated worktree or VM, one pull request per sliceWave order, merge queue, disputesPer slice: goldens unchanged and green on the new path, fitness functions green, an evidence bundle on the pull request
4. Prove and cut overRun shadow comparison, triage mismatches, draft the cut-over recordMismatch decisions, rollout, the go decisionZero unexplained mismatches over a full business cycle, rollback rehearsed, cut-over record signed
  1. Map the system and write the slice plan. Run the first prompt below in plan mode (/plan in Claude Code and Codex, Plan Mode in Cursor), review the plan, then approve it so the agent writes slices.draft.yaml and map.md. Edit it until each slice has one seam, one owner and a risk class. A slice is right-sized when its suite runs in minutes and its migration fits one reviewable pull request. For very large systems, see million-line codebases.

  2. Pin the behavior of the first slices. Record goldens against the running legacy system with the harness below, on real inputs: anonymized production requests, the support team’s odd accounts, month-end cases. A golden that looks like a bug goes on the slice’s quirks list, to be decided later in its own change. The full technique is on characterization tests.

  3. Prove the goldens can fail. Run mutation testing on the legacy slice and hold the score to the agreed floor (tooling below). A surviving mutant is an input you have not recorded: in our test, an empty-invoice golden stayed green when we changed the tax rate, because it totals zero at any rate.

  4. Lock the oracle. Put tests/characterization/ under the oracle owner in CODEOWNERS, deny agents edit access (tool tabs below), and add the CI guard. From here on, a rule, not a reviewer’s attention, rejects a migration pull request that touches a golden.

  5. Cut the seam. Route the slice’s traffic through one point with a flag that sends each request to the legacy or the new path. Merge it with the legacy path still answering everything, and confirm the suite is green through it (the strangler fig pattern).

  6. Migrate in waves. A wave is a set of independent slices, each an isolated agent job ending in one pull request. Start with one low-risk pilot slice, then widen. Keep a slice to one intent: port, upgrade or fix, never two.

  7. Shadow, then cut over. Deploy the new path dark, compare it with the legacy path on live traffic, and fix or explicitly accept every mismatch. Then move traffic in steps behind the flag, as in progressive delivery, and file the cut-over record.

  8. Delete the legacy path. After the agreed period at 100%, remove the old code, flag and shadow wiring in one pull request. The characterization suite stays as the regression suite.

The mutated code has to run under the characterization tests, or the score means nothing.

  • JavaScript or TypeScript. Stryker’s command runner can boot the legacy app from the mutated sandbox, then run the HTTP suite. Set coverageAnalysis: "off" and thresholds.break to the floor. Install it with npm i -D @stryker-mutator/core (10.0.0 on npm in September 2026).
  • PHP. Infection 0.35.4 (current on Packagist on 26 September 2026; composer require --dev infection/infection) runs only PHPUnit, Pest or Codeception suites in-process, so it cannot score an HTTP suite against a separate server. Give it a PHPUnit suite for the slice that replays the recorded cases through the front controller against the same goldens.

Keep the plan in the repository as data, so agents can read it, CI can check it, and status is a grep away:

# modernization/slices.yaml: the program's single source of truth
program: billing-php-to-ts
oracle_owner: "@acme/billing-leads"
mutation_score_floor: 80 # percent, per slice, measured on the legacy code
slices:
- id: customer-search
seam: "GET /api/customers"
risk: low # the pilot: first wave, one low-risk slice
depends_on: []
goldens: tests/characterization/golden/customer-search/
quirks: []
status: shadow # mapped | pinned | seamed | migrating | shadow | cut-over | deleted
wave: 1
- id: invoice-total
seam: "POST /api/invoices/total"
risk: high # money: code owner review plus staged rollout
depends_on: []
goldens: tests/characterization/golden/invoice-total/
quirks:
- "Totals round half-down on line items (legacy). Keep until INV-311 decides."
status: seamed
wave: 2
- id: tax-lookup
seam: "GET /api/tax/rate" # an HTTP seam, so the wave's done command can verify it
risk: high
depends_on: []
goldens: tests/characterization/golden/tax-lookup/
quirks: []
status: seamed
wave: 2
- id: vat-id-check
seam: "POST /api/vat-ids/validate"
risk: medium
depends_on: []
goldens: tests/characterization/golden/vat-id-check/
quirks: []
status: seamed
wave: 2
- id: invoice-pdf
seam: "GET /api/invoices/:id/pdf"
risk: medium
depends_on: [invoice-total]
goldens: tests/characterization/golden/invoice-pdf/
quirks: []
status: mapped
wave: 3
# ...one entry per slice

Three rules keep the plan truthful: a slice enters a wave only at seamed; a slice waits until everything in its depends_on reaches cut-over; and quirks is the only place a known bug may live. Prefer HTTP seams, because the wave commands below use the HTTP suite as their definition of done. A function-signature seam needs in-process tests in each language and its own done command.

A characterization harness at the HTTP seam

Section titled “A characterization harness at the HTTP seam”

Recording at the HTTP seam lets one suite pin a PHP system and verify its TypeScript replacement, because both answer the same requests. Point TARGET_URL at the legacy deployment to record and at the new one to verify.

tests/characterization/invoice-total.char.test.ts
import { readFileSync } from 'node:fs';
import { describe, expect, it } from 'vitest';
// Point TARGET_URL at the legacy system to record, at the new one to verify.
const BASE_URL = process.env.TARGET_URL ?? 'http://localhost:8080';
const cases: { id: string; body: unknown }[] = JSON.parse(
readFileSync(new URL('./invoice-total.cases.json', import.meta.url), 'utf8'),
);
// Mask what differs on every run, so the golden file pins behavior, not noise.
const scrub = (body: Record<string, unknown>) => ({ ...body, requestId: '<id>', generatedAt: '<ts>' });
describe('POST /api/invoices/total: pinned legacy behavior', () => {
it.each(cases)('$id', async ({ id, body }) => {
const res = await fetch(`${BASE_URL}/api/invoices/total`, {
method: 'POST',
headers: { 'content-type': 'application/json' },
body: JSON.stringify(body),
});
const observed = { status: res.status, body: scrub(await res.json()) };
await expect(JSON.stringify(observed, null, 2) + '\n').toMatchFileSnapshot(
`./golden/invoice-total/${id}.json`,
);
});
});

Record with TARGET_URL=https://legacy.internal npx vitest run tests/characterization. Vitest writes a missing golden locally but fails the test when CI is set (tested with Vitest 5.0.2), so an empty oracle cannot pass in a pipeline. Never run vitest -u on a migration branch: updating goldens is a behavior change.

The scrub function is the part people get wrong. Mask only values that differ on every run of the same system, such as request IDs and timestamps. Masking a field because the new system formats it differently is a decision that belongs in the slice’s quirks list.

Keep the pinned behavior out of the agent’s reach

Section titled “Keep the pinned behavior out of the agent’s reach”

During the waves, the agent makes the goldens pass without changing them. Enforce that in the agent’s session and in CI. This guard works for every tool, because it checks the pull request, not the session:

.github/workflows/oracle-guard.yml
name: oracle-guard
on:
pull_request:
types: [opened, synchronize, reopened, labeled, unlabeled]
permissions:
contents: read
jobs:
goldens-unchanged:
if: ${{ !contains(github.event.pull_request.labels.*.name, 'behavior-change') }}
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
with:
fetch-depth: 0
persist-credentials: false
- name: Fail if a migration PR changes pinned behavior
env:
BASE_REF: ${{ github.base_ref }}
run: |
changed=$(git diff --name-only "origin/$BASE_REF...HEAD" -- tests/characterization/)
if [ -n "$changed" ]; then
printf 'Pinned behavior changed:\n%s\n' "$changed" >&2
echo "Split it into its own PR labelled behavior-change, approved by the oracle owner." >&2
exit 1
fi

Only people with triage or write access can add the behavior-change label, but an agent with a write-scoped GitHub token (a Cloud Agent, or gh in the session) can add it itself: give agent tokens no label permission, or make the guard also require the oracle owner’s approval. Make both the guard and the characterization run required checks. A pull_request workflow runs the pull request’s own copy of the guard, so an agent could edit it until it passes: put .github/workflows/ and modernization/slices.yaml under the oracle owner in CODEOWNERS too and turn on Require review from Code Owners, so a guard edit needs the same approval as a golden edit.

Deny edits to the oracle in the committed project settings, .claude/settings.json, where a leading / is relative to the project that holds the settings file (the primary working directory):

{
"permissions": {
"deny": [
"Edit(/tests/characterization/**)",
"Edit(/modernization/slices.yaml)",
"Edit(/.github/**)",
"Edit(/.claude/settings.json)",
"Edit(/.claude/settings.local.json)",
"Edit(/.claude/hooks/**)"
]
},
"sandbox": {
"enabled": true,
"filesystem": {
"denyWrite": [
"./tests/characterization",
"./modernization/slices.yaml",
"./.github",
"./.claude/settings.json",
"./.claude/settings.local.json",
"./.claude/hooks"
]
},
"network": { "allowLocalBinding": true }
}
}

Deny rules alone stop the edit tools, not a script the agent runs. With the sandbox enabled, Claude Code also applies Edit(...) deny rules to sandboxed commands (v2.1.283); the explicit denyWrite list makes that visible and covers any path you do not deny for Edit. The .claude entries name files rather than all of .claude/, because Claude Code creates its worktrees under .claude/worktrees/ (v2.1.283), and a blanket rule could block every /batch unit’s edits.

On macOS the sandbox refuses local port binding by default, so a wave unit could not start its server; sandbox.network.allowLocalBinding: true allows it. The key is macOS only (checked on v2.1.283) and changes nothing on Linux or WSL2. Add the hook from protecting the oracle, and prove the setup on the pilot slice before you launch a wave: ask a /batch unit to edit a golden in its worktree (under .claude/worktrees/<name>/) and confirm the edit is refused. If it is not, the root-anchored deny rules do not reach that worktree, so rely on the sandbox and the oracle-guard check.

How do you run a migration wave in each tool?

Section titled “How do you run a migration wave in each tool?”

The shape is the same in every tool: one slice per isolated checkout, one pull request per slice, the characterization suite as the definition of done.

For one slice, run claude --worktree invoice-total (-w for short) and paste the migration prompt below.

For a whole wave, the bundled /batch skill “decomposes the work into 5 to 30 independent units, and presents a plan”; after you approve it, it “spawns one background subagent per unit in an isolated worktree”, and each one “implements its unit, runs tests, and publishes its change” (Claude Code commands reference, checked 26 September 2026 against v2.1.283). Name the slices so it does not invent its own:

/batch migrate wave 2 from modernization/slices.yaml (invoice-total, tax-lookup, vat-id-check). One unit per slice. Give unit i its own port, PORT=3001+i (3001, 3002, 3003). Each unit ports only its slice to services/billing-ts, leaves tests/characterization/ untouched, starts services/billing-ts from its own worktree on its PORT, and is done only when TARGET_URL=http://localhost:$PORT npx vitest run tests/characterization/<slice> passes.

For larger waves, or orchestration you can rerun, use a dynamic workflow: Anthropic lists “large migrations” as a primary use, at “Dozens to hundreds of agents per run”. On Pro, enable them in /config first.

Size a wave by the review and merge capacity behind it, not by how many agents you can start: 12 pull requests on one reviewer in one afternoon moves the bottleneck. See orchestration patterns.

Copy-paste prompts for legacy modernization

Section titled “Copy-paste prompts for legacy modernization”

How do you prove the new path in production before cut-over?

Section titled “How do you prove the new path in production before cut-over?”

Characterization tests cover inputs you thought of; a parallel run covers what users send. It runs the old path as control and the new one as candidate, returns the control’s result, and reports any difference:

src/modernization/shadow.ts
type Outcome = 'match' | 'mismatch' | 'error';
export interface ShadowOptions<I, O> {
slice: string;
legacy: (input: I) => Promise<O>;
modern: (input: I) => Promise<O>; // must have no side effects while shadowing
enabled: (input: I) => boolean; // your feature flag: per tenant or a percentage
normalize?: (output: O) => unknown; // sort keys, drop fields that may differ
record: (event: { slice: string; outcome: Outcome; detail?: unknown }) => void;
}
export function shadow<I, O>(opts: ShadowOptions<I, O>): (input: I) => Promise<O> {
const norm = opts.normalize ?? ((o: O) => o);
return async (input) => {
const control = await opts.legacy(input); // users still get the legacy answer
if (opts.enabled(input)) {
// Not awaited: the comparison never adds latency or errors to the request.
// Recording failures are swallowed by the final catch, so they cannot surface either.
void opts
.modern(input)
.then(
(candidate) => {
const same = JSON.stringify(norm(candidate)) === JSON.stringify(norm(control));
opts.record({ slice: opts.slice, outcome: same ? 'match' : 'mismatch', detail: same ? undefined : { control, candidate } });
},
(err: unknown) => opts.record({ slice: opts.slice, outcome: 'error', detail: String(err) }),
)
.catch(() => {});
}
return control;
};
}

Two rules make it safe. The candidate must be read-only while it shadows (point its writes at a shadow store or stub them), or every request happens twice. And record receives real data, so redact personal data first.

Define the numbers before you look at them, so nobody argues the threshold after seeing the result:

MetricDefinitionCut-over gate we recommend
Compared callsRequests where both paths ran and record firedA full business cycle, such as a month-end close
Unexplained mismatch rateMismatches not classified as a fixed port-bug or an accepted legacy-quirk, per compared callZero at the end of the window
Candidate error rateerror outcomes per compared callZero, or each error traced to a ticket
Golden coverageShare of golden case IDs whose input shape appeared in live trafficEvery golden seen at least once, or the gap is explained

The CTO or named service owner signs the cut-over; the agent drafts the record. Keep it in modernization/cutover/<slice>.md and link it from the pull request that flips the flag:

  • Characterization suite green on the new path, goldens unchanged since pinning (link the oracle-guard runs).
  • Mutation score on the legacy slice at or above mutation_score_floor, measured when the goldens were recorded.
  • Fitness functions and type checks green on the new service.
  • Shadow window dates, compared-call count, and zero unexplained mismatches.
  • Every legacy-quirk either ported on purpose or changed in its own approved behavior-change pull request.
  • Rollout steps and the metric that stops each step, written down before the first step (see progressive delivery).
  • Rollback rehearsed: flag flipped back in staging, legacy path answered correctly.
  • Deletion date for the legacy path and its flag.

The approver can say yes without reading the migrated code, as in evidence, not diffs.

What breaks in an agent-run modernization?

Section titled “What breaks in an agent-run modernization?”

The agent “fixes” a golden. Symptom: a green migration pull request touches tests/characterization/. Recovery: oracle-guard rejects it; if a golden already changed on main, revert and re-record from legacy.

The goldens pass but production disagrees. Symptom: shadow mismatches on inputs the suite never recorded. Recovery: add each representative input as a golden case (the fourth prompt drafts them); many mismatches mean the slice was pinned too thinly.

Wave pull requests conflict. Symptom: slices in one wave edit the same router, manifest or shared type. Recovery: land shared scaffolding before the wave and tighten depends_on. Slices that keep colliding were one slice.

The shared database is the hidden coupling. Symptom: the new service reads a table the legacy code writes, so slices cannot cut over independently. Recovery: make the data seam its own slice with goldens, and sequence schema changes with expand, migrate, contract from database development.

The long tail stalls and burns budget. Symptom: 80% of slices are cut over and agent attempts on the rest pile up without merges. Recovery: give each remaining slice an owner and a deletion date, and after two failed runs route it to a human for a better seam or more goldens. Use best-of-N attempts in Codex cloud or a goal-driven loop on the hardest ones, and track merged slices per week, not agent runs. A modernization stalled at 80% maintains two systems forever.

Where to go next with legacy modernization

Section titled “Where to go next with legacy modernization”