Skip to content

Spec-driven development: the spec as the source of truth

Spec-driven development makes a versioned spec.md the authoritative description of what a system does. Every pull request that changes behaviour carries a spec delta, the agent implements against it, and each requirement is tied to a test. People review the spec delta and its evidence instead of the diff, which works only while spec-code drift is caught mechanically.

Three months ago your team started writing spec.md before every feature, and the first specs were excellent. Today the audit-log spec says CSV exports stop at 10,000 rows, the code streams without a limit, and nobody knows which one is right. Worse, the agent reads the spec, so last week it “fixed” the code back to 10,000 rows and broke a customer’s nightly export. A spec that nobody keeps true is more dangerous than no spec, because agents believe it.

This page is for the developer who writes specs for agents and the tech lead who wants the team to review specs instead of 2,000-line diffs.

  • The three rules that make a spec authoritative rather than a planning note you discard after the merge.
  • A spec.md template with requirement IDs, observable outcomes and non-goals, ready to commit.
  • The spec-delta pull request: the order of commits, a PR template section, and who approves what.
  • Two drift checks: a 30-line traceability script for CI and a scheduled read-only agent audit for Claude Code, Codex and Cursor.
  • Four copy-paste prompts: draft a delta, implement against it, audit drift, and retrofit a spec onto existing code.
  • A decision table for plain markdown versus Spec Kit, OpenSpec, Kiro specs and BMAD, with verified install commands.

Most teams already write something before they prompt an agent. The difference between a prompt and a source of truth is what happens after the merge. Two widely cited practitioners put the spec at the centre: Andrej Karpathy put it as “You are in charge of the spec and plan” (Sequoia Ascent 2026 summary, 2026-04-30), and Dan Shapiro’s Level 4 developer is the one who will “write a spec” and check whether the tests pass (The Five Levels, 2026-01-23). Neither says how to keep the spec true. These three rules do.

RuleWhat it means in the repositoryWhat breaks without it
1. The spec describes current behaviourspecs/<capability>/spec.md says what the system does today, not what one feature added. Feature documents are deltas against it.Nobody can answer “what does export do?” without reading ten feature folders and the code.
2. Behaviour changes go through the spec firstA PR that changes observable behaviour contains a spec delta, committed before the code.The code moves and the spec silently becomes fiction.
3. Every requirement is traceable to a checkEach requirement has an ID; at least one test names that ID. CI fails on untested or orphaned IDs.Drift is invisible until an agent “corrects” the code back to a stale spec.

Rule 1 is the easiest to lose. Planning tools produce a spec per feature, and after twenty features the current behaviour is scattered across twenty folders. Whatever tool you use, keep one merged spec.md per capability, and treat feature documents as change requests against it.

The prerequisite for this page is writing the first spec.md in the design stage. This page is about keeping that file true for the life of the product. Where the spec sits in the full chain of handoffs is on the artifact chain.

An authoritative spec states behaviour a test can observe, and nothing a refactor would change. Use one file per capability, stable requirement IDs, and “when … the system shall …” sentences with a concrete example underneath. That sentence shape is the EARS notation, which Kiro’s documentation describes for its requirements files (as reported in secondary sources, September 2026; see Kiro).

# Audit log export
Owner: @team-compliance · Last reviewed: 2026-09-26
## Scope
Team admins export audit-log entries as CSV. Non-goals: scheduled exports, PDF.
## Requirements
### EXPORT-1: Filter by date range and actor
WHEN an admin requests an export with `from`, `to` and optional `actor`
THE SYSTEM SHALL return only entries with `from <= created_at < to`
and, when `actor` is set, only that actor's entries.
Example: from=2026-09-01, to=2026-09-02 excludes an entry at 2026-09-02T00:00:00Z.
### EXPORT-2: Streaming without a row limit
WHEN the result has more than 10,000 rows
THE SYSTEM SHALL stream all rows and SHALL NOT truncate.
(Changed 2026-09-12 in #4812; previously capped at 10,000.)
### EXPORT-3: Authorisation
WHEN a non-admin requests an export
THE SYSTEM SHALL respond 403 and SHALL record the denied attempt in the audit log.
## Open questions
- EXPORT-2: is there a maximum export duration? (owner: @pm-audit)

Four things in that file do the work. The IDs give tests and PRs something to point at. The example under each requirement is the input the test will use. The non-goals stop an agent from adding scheduled exports because they seemed helpful. The change note on EXPORT-2 answers the question from the opening scenario before anyone asks it.

Leave out class names, file paths, library choices and database schemas. Those belong in plan.md or an architecture decision record. If a pure refactor forces a spec change, the spec is describing implementation, and it will drift every week.

How does a pull request carry a spec delta?

Section titled “How does a pull request carry a spec delta?”

A spec delta is the part of the PR that changes specs/. It lands first, in its own commit, so the reviewer can read the behaviour change before any code exists and the agent implements against an approved target.

  1. Draft the delta from the change request. The agent reads the ticket and the current spec.md, then edits only the spec: new or changed requirements, a change note, and open questions. Use the first prompt below.

  2. Get the delta approved before code. The product owner approves user-facing behaviour; the tech lead approves non-functional requirements such as limits, latency and security. Push the spec commit and request review on it alone, or approve it in the session if the change is low-risk.

  3. Implement against the approved spec. The agent writes a failing test per changed requirement, named with the requirement ID, then the code. It must not edit specs/ in this phase. If it finds the spec is wrong, it stops and reports.

  4. Run the gates. The traceability check and the behaviour-path check (next section) run in CI next to types, lint and tests.

  5. Review the delta and the evidence. The reviewer reads the spec diff, checks that each changed ID has a test that failed before and passes now, and reads code only for the escalation classes: auth, money, schema and migrations.

Add this section to the pull request template so every agent-opened PR fills it in:

## Spec delta
- Spec file(s): specs/audit-log/spec.md
- Requirements added/changed/removed: EXPORT-2 (changed)
- Approved by: @pm-audit (behaviour), @lead-platform (non-functional)
## Evidence per requirement
| ID | Test | Failed before | Passes now |
|----|------|---------------|------------|
| EXPORT-2 | tests/export/streaming.test.ts "[EXPORT-2] streams 25,000 rows" | yes | yes |
## No behaviour change?
If this PR changes no observable behaviour, add the label `no-behaviour-change` and say why.

The full contract for what an agent’s PR must prove is on the evidence bundle; the review protocol that reads it is reviewing an agent’s pull request.

The workflow is the same in all three tools: propose the delta without touching code, approve it, then write only the spec file. The mechanism that keeps code untouched differs.

Start in plan mode (/plan, or Shift+Tab to cycle to it). Plan mode blocks edits until you approve, so the agent presents the delta as its plan. Approve it, and let it write only the files under specs/. Then run git diff --stat and confirm that nothing outside specs/ changed.

How do you detect drift between the spec and the code?

Section titled “How do you detect drift between the spec and the code?”

Drift has two shapes. Structural drift is mechanical: a requirement with no test, a test that names a requirement that no longer exists, or a behaviour change with no spec delta. Semantic drift is the opening scenario: a test exists, but the code and the spec now disagree about what it should check. Catch the first on every PR with deterministic checks, and the second on a schedule with a read-only agent audit.

This script fails the build when a requirement has no test or a test cites a requirement that does not exist. It assumes requirement headings like ### EXPORT-1: under specs/ and test names containing [EXPORT-1] under tests/; change the two patterns to match your layout.

// scripts/check-spec-trace.mjs: fails CI on untested or orphaned requirement IDs
import { readdirSync, readFileSync, statSync } from 'node:fs';
import { join } from 'node:path';
const walk = (dir) =>
readdirSync(dir).flatMap((name) => {
const path = join(dir, name);
return statSync(path).isDirectory() ? walk(path) : [path];
});
const idsIn = (dir, fileRe, idRe) => {
const ids = new Set();
for (const file of walk(dir).filter((f) => fileRe.test(f))) {
for (const m of readFileSync(file, 'utf8').matchAll(idRe)) ids.add(m[1]);
}
return ids;
};
const specIds = idsIn('specs', /\.md$/, /^###\s+([A-Z]+-\d+):/gm);
const testIds = idsIn('tests', /\.(test|spec)\.[cm]?[jt]sx?$/, /\[([A-Z]+-\d+)\]/g);
const untested = [...specIds].filter((id) => !testIds.has(id));
const orphaned = [...testIds].filter((id) => !specIds.has(id));
if (untested.length) console.error(`Requirements without a test: ${untested.join(', ')}`);
if (orphaned.length) console.error(`Tests citing unknown requirements: ${orphaned.join(', ')}`);
if (untested.length || orphaned.length) process.exit(1);
console.log(`Spec trace OK: ${specIds.size} requirements, all tested.`);

The test-side pattern matches any bracketed token of that shape, so [UTF-8] in a test title counts as an orphaned requirement. Give each capability a distinct ID prefix that no other bracketed text uses, and widen both regexes (for example to [A-Z]+(?:-[A-Z]+)*-\d+) if your IDs contain several hyphens, such as BILLING-INVOICE-1.

The second check catches rule 2: a PR that touches behaviour paths without touching specs/. The label escape hatch keeps refactors cheap while making the claim visible to the reviewer. Because the workflow also listens for labeled and unlabeled, adding or removing the label re-runs the check without a new push.

.github/workflows/spec.yml
name: spec
on:
pull_request:
types: [opened, synchronize, reopened, labeled, unlabeled]
jobs:
spec-gates:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- name: Requirements are traceable to tests
run: node scripts/check-spec-trace.mjs
- name: Behaviour changes carry a spec delta
if: ${{ !contains(github.event.pull_request.labels.*.name, 'no-behaviour-change') }}
env:
BASE: ${{ github.event.pull_request.base.sha }}
run: |
changed=$(git diff --name-only "$BASE"...HEAD)
if echo "$changed" | grep -qE '^src/(api|domain)/' && ! echo "$changed" | grep -q '^specs/'; then
echo "src/api or src/domain changed without a spec delta. Update specs/ or add the no-behaviour-change label."
exit 1
fi

Set the src/(api|domain)/ pattern to the directories where observable behaviour lives in your repository. Protect specs/ with a CODEOWNERS entry for the capability owners, so a spec change always needs their approval, whoever wrote it.

Once a week, or before a release, run a read-only agent pass that compares each requirement with the code and the tests and reports disagreements with file-and-line evidence. The agent reports; a person decides which side is wrong. If the code is wrong, it is a bug with a failing test. If the spec is wrong, the code encodes an undocumented decision, and that decision gets a spec delta and an approver like any other.

Run it headless with only the read-only tools available. --tools limits the session to the listed built-in tools; --allowedTools would only pre-approve them and leave Edit or Bash available if your settings allow them. Put the prompt before --tools: the flag takes a variadic list and would otherwise swallow the prompt as a tool name (checked against Claude Code 2.1.283).

Terminal window
claude -p "$(cat .github/prompts/spec-drift.md)" \
--tools "Read,Grep,Glob" > drift-report.md

Treat each DRIFTED row as a hypothesis until a person confirms it, ideally by writing the disagreeing input as a test. An audit that cites no file and line is not evidence.

What do humans review instead of the code?

Section titled “What do humans review instead of the code?”

Once the spec is authoritative and traced, the reviewer’s job changes from “is this diff correct?” to “is this the behaviour we want, and is it proven?”. Split the sign-off so each approver judges what they can judge.

What is reviewedWho signs offEvidence they needCode reading
Spec delta: user-facing behaviourProduct ownerThe spec diff, examples, non-goalsNone
Spec delta: limits, latency, securityTech leadThe spec diff and the test or benchmark that enforces itNone
Implementation matches the deltaAuthor, then reviewer or a review agentTrace check green, per-ID table (failed before, passes now), full suite greenOnly where the evidence is thin
Escalation classes: auth, money, schema, migrationsCode ownerAll of the aboveYes, always
Drift audit findingsCapability ownerDRIFTED rows with a reproducing testOnly the cited lines

The reasoning behind this split, and how to build the trust to use it, is on reading evidence instead of code. Turning a story into the failing test the table relies on is covered in executable acceptance criteria.

Should you use plain markdown or a spec framework?

Section titled “Should you use plain markdown or a spec framework?”

Start with plain markdown plus the two CI checks above. Every framework below generates good documents, but none of them makes your spec current or traced by itself, and the checks work with all of them. Adopt a framework when you want its ceremony: gated stages, a project constitution, or a change-proposal workflow. Versions and stars were read on 2026-09-26 from npm, PyPI and GitHub.

OptionArtifactsHow it keeps “current behaviour”ToolsPick it whenAvoid when
Plain markdownspecs/<capability>/spec.md + PR templateYou merge each delta by hand; the CI checks enforce itAny agentYou want the smallest process that worksNobody owns the spec files
Spec Kit (GitHub; specify-cli 1.0.12; 138.9k stars)Constitution, then spec, plan and tasks per featurePer-feature documents; merging them into a current spec is your job40+ integrations, including Claude Code, Codex and CursorGreenfield features that need a written trail and a constitutionSmall changes in a large codebase: each feature produces several documents
OpenSpec (Fission AI; @fission-ai/openspec 1.13.2; 70.4k stars)Per change: proposal.md, spec deltas, design.md, tasks.mdArchiving a change merges its deltas into openspec/specs/, the living specClaude Code, Codex, Cursor and othersExisting codebases with many incremental changesYou need a heavyweight governance trail
Kiro specs (AWS; built into Kiro).kiro/specs/<feature>/: requirements.md (EARS), design.md, tasks.mdPer-feature, like Spec KitKiro only; cc-sdd 3.1.0 gives the same shape to other agentsYour team already works in KiroYour team is on Claude Code, Codex or Cursor
BMAD Method (BMad Code; bmad-method 6.12.0; 53.5k stars)Product brief, PRD, architecture, spec, storiesPer-document; its value is the hand-offs between personasClaude Code, Codex, Cursor and ~40 other tool IDsProduct work with several stakeholders and epicsA solo developer shipping small features

Two traps show up in old tutorials. Spec Kit removed specify init --ai claude in 0.10.0; the flag is now --integration. OpenSpec’s /openspec:proposal, /openspec:apply and /openspec:archive are the retired workflow; the current one is /opsx:*. The npm package openspec is an unrelated placeholder, so install the scoped @fission-ai/openspec. The whole field, including Tessl and Superpowers, is compared on the frameworks overview.

Install the two frameworks closest to this workflow

Section titled “Install the two frameworks closest to this workflow”

OpenSpec fits this page’s model most directly, because its archive step produces the merged, current spec that rule 1 requires. Spec Kit fits greenfield work where you want each stage gated. The command spelling differs per tool.

Terminal window
# OpenSpec (Node >= 20.19.0)
npm install -g @fission-ai/openspec@latest
openspec init --tools claude
# in the agent: /opsx:propose audit-export-streaming → /opsx:apply → /opsx:archive
# Spec Kit (needs uv)
uv tool install specify-cli
specify init --here --integration claude
# in the agent: /speckit-specify, /speckit-plan, /speckit-tasks, /speckit-implement

In an existing repository, specify init --here asks for confirmation before it writes into a non-empty directory; add --force to skip the prompt, and --non-interactive as well in scripts and CI (specify init --here --force --non-interactive --integration claude).

After openspec init, openspec validate --all --json checks the change and spec files in CI. It validates the documents, not the code, so keep the traceability check next to it.

A retrofitted spec records what the code does, bugs included. Have the capability owner walk the “Suspicious behaviour” list before the spec becomes authoritative, and add the missing tests before you switch on the traceability check for that capability.

What breaks when the spec is the source of truth?

Section titled “What breaks when the spec is the source of truth?”

The spec rots after the first feature. Symptom: spec.md has not changed in months while the code has. Recovery: run the drift audit, turn every DRIFTED row into a spec delta or a bug, then switch on the behaviour-path check so it cannot recur silently.

The agent edits the spec to match its code. Symptom: an implementation PR contains a spec change nobody asked for, and the tests pass. This is the spec version of an agent weakening its own tests. Recovery: implementation prompts forbid edits under specs/, CODEOWNERS requires the capability owner on any spec change, and implementation sessions deny writes to specs/ (see permissions and sandboxes).

The spec describes implementation. Symptom: every refactor needs a spec delta, so people start adding no-behaviour-change to everything. Recovery: move class names, schemas and library choices into plan.md or an architecture decision record, and keep only observable behaviour in the spec.

Requirements are too vague to test. Symptom: the agent writes a test that passes for any implementation, such as “export works”. Recovery: require one concrete example per requirement and reject a delta without it. Property-based tests turn invariants from the spec into checks an agent cannot satisfy by example.

Framework ceremony swamps small changes. Symptom: a one-line bug fix produces four documents, and developers bypass the process. Recovery: write a spec delta only when observable behaviour changes. A bug fix that restores specified behaviour needs a test that cites the existing ID, not a new spec.

The drift audit invents drift. Symptom: the report flags disagreements that are not there, and people stop reading it. Recovery: the prompt demands file-and-line evidence and a concrete disagreeing input, and a DRIFTED row counts only once someone writes that input as a failing test.

Old tutorial commands fail. Symptom: specify init --ai claude errors, or /openspec:proposal does nothing. Recovery: use --integration (Spec Kit) and /opsx:* or the per-tool spelling (OpenSpec), as in the install tabs above.

Where to go next with spec-driven development

Section titled “Where to go next with spec-driven development”