GitHub Spec Kit: spec-driven development in practice
GitHub Spec Kit is GitHub’s open-source toolkit for spec-driven development with coding agents. Its specify CLI installs agent skills that walk a feature through constitution, specify, plan, tasks, implement and converge, writing a reviewable Markdown file at each step. It works in Claude Code, Codex and Cursor, and it pays off on features large enough to deserve a written trail.
Your team admins want to export the audit log as CSV. Last quarter a feature like this went to an agent as a one-paragraph ticket and came back as 1,400 lines: the export ignored the date filter for one role, and nobody could say which behaviour was intended because nothing was written down. You want the next feature to arrive with a spec a product owner approved, a plan a tech lead approved, and proof that the code does what the spec says.
This page is for the developer who runs the agent and the tech lead who decides whether the team adopts Spec Kit. It runs one real feature, the audit-log export, through every stage in all three tools.
What you get from running Spec Kit on one feature
Section titled “What you get from running Spec Kit on one feature”- A working install in Claude Code, Codex or Cursor, including a repository that serves all three at once.
- The exact command spelling per tool:
/speckit-specifyin Claude Code and Cursor,$speckit-specifyin Codex. - The file every stage writes, shown for the audit-log export, and the question you ask before you accept it.
- An implement-and-converge loop that ends only when the agent reports “Converged”, and a pull request that carries the spec trail instead of asking a reviewer to read the diff.
- The
bugandassessextensions for work that is not a new feature. - A clear rule for when Spec Kit adds ceremony you should skip.
Where does Spec Kit fit in the artifact chain?
Section titled “Where does Spec Kit fit in the artifact chain?”The artifact chain says every stage commits one file the next stage reads. Spec Kit generates most of that chain for you, with its own file names.
| Artifact chain | Spec Kit stage | File Spec Kit writes | Who accepts it |
|---|---|---|---|
| Project rules | speckit-constitution (once per project) | .specify/memory/constitution.md | Tech lead |
intent.md | speckit-assess-* (optional extension) | .specify/assessments/<slug>/ | Product owner |
spec.md | speckit-specify, then speckit-clarify | specs/001-<feature>/spec.md | Product owner |
plan.md | speckit-plan, then speckit-checklist | plan.md, research.md, data-model.md, contracts/, quickstart.md | Tech lead |
| Task list | speckit-tasks, then speckit-analyze | tasks.md | Engineer who runs the agent |
| Diff and tests | speckit-implement | Code, with tasks ticked [X] | CI |
| Proof against the spec | speckit-converge | A ## Phase N: Convergence section in tasks.md, or no change | Engineer, then the PR reviewer |
Two gaps remain your job. Spec Kit writes one folder per feature, so nothing merges specs/001-audit-log-export/spec.md into a current description of the whole system; the spec-driven development page covers keeping a living spec and detecting drift. And speckit-converge is a model’s judgment, so each requirement still needs a test that fails when the behaviour is wrong.
Install Spec Kit and wire it into your agent
Section titled “Install Spec Kit and wire it into your agent”Spec Kit needs Python 3.11 or later (PyPI requires_python >=3.11), uv and Git. The PyPI package is specify-cli; the command it installs is specify.
-
Install the CLI once per machine (terminal):
Terminal window uv tool install specify-clispecify version # expect CLI Version 1.0.12 or laterTo pin a release for the whole team, install from a tag instead:
uv tool install specify-cli --from git+https://github.com/github/spec-kit.git@v1.0.12.specify self checkreports a newer release andspecify self upgradeinstalls it. -
Initialize the repository for your agent. The flag is
--integration; the old--aiflag was removed in 0.10.0.Terminal window cd your-repospecify init --here --integration claude --script sh# writes .claude/skills/speckit-*/SKILL.md (10 skills)claude# in the session: /speckit-constitution, /speckit-specify, ...Terminal window cd your-repospecify init --here --integration codex --script sh# writes .agents/skills/speckit-*/SKILL.md (10 skills)codex# in the session: $speckit-constitution, $speckit-specify, ...Terminal window cd your-repospecify init --here --integration cursor-agent --script sh# writes .cursor/skills/speckit-*/SKILL.md (10 skills)# open the repo in Cursor; in Agent chat: /speckit-constitution, /speckit-specify, ...In a script or CI job use
specify init --here --force --integration claude --script sh --non-interactive;--herein a non-empty repository needs--forcewhen nothing can answer the prompt. Without--integration, a non-interactive run picks Copilot, not your agent. -
If your team uses more than one agent, add the others to the same repository. They share
.specify/and thespecs/folders, so a spec written in Claude Code can be implemented in Codex:Terminal window specify integration install codexspecify integration install cursor-agentspecify integration status # "Installed integrations: claude, codex, cursor-agent" -
Optionally add the bundled extensions:
specify extension add git(feature branches and auto-commits around each stage),specify extension add bugandspecify extension add assess. Extensions register hooks in.specify/extensions.yml; read that file before you commit it. -
Commit
.specify/,specs/and thespeckit-*skill folders. The init output suggests adding the agent folder to.gitignorebecause agents can store credentials there. Ignore files such as.claude/settings.local.json, not the skills your teammates need.
What does specify init write into the repository?
Section titled “What does specify init write into the repository?”This is the tree after specify init --here --integration claude in 1.0.12, before you run any stage. Codex puts the same 10 skills in .agents/skills/, and Cursor in .cursor/skills/.
Directory.claude/skills/
- speckit-constitution/SKILL.md
- speckit-specify/SKILL.md
- speckit-clarify/SKILL.md
- speckit-plan/SKILL.md
- speckit-checklist/SKILL.md
- speckit-tasks/SKILL.md
- speckit-analyze/SKILL.md
- speckit-implement/SKILL.md
- speckit-converge/SKILL.md
- speckit-taskstoissues/SKILL.md
Directory.specify/
- memory/constitution.md a template until you run the constitution stage
Directorytemplates/ spec, plan, tasks, checklist and constitution templates
- …
Directoryscripts/bash/ create-new-feature.sh, setup-plan.sh and helpers
- …
- workflows/speckit/workflow.yml
Directoryintegrations/ manifests of the files Spec Kit manages
- …
- init-options.json feature numbering, script type, version
- .gitignore ignores feature.json, the per-checkout pointer to the current feature
The skills are ordinary Agent Skills in your repository, not a plugin, so there is no marketplace entry and no claude plugin details figure. specify integration status reports whether anyone has edited the managed files.
Run the full cycle on the audit-log export
Section titled “Run the full cycle on the audit-log export”The commands below use the Claude Code and Cursor spelling. In Codex, replace the leading / with $. Run one stage at a time and review its file before you start the next; the stages are separate skills precisely so that a person can stop the chain.
-
Constitution (once per project). Principles that every later stage checks against.
/speckit-constitution Test-first: every functional requirement gets a failing test before code. No new runtime dependency without an ADR in docs/adr/. Every endpoint enforces authorization server-side. p95 latency under 200 ms for interactive endpoints.The skill fills
.specify/memory/constitution.md, gives it a semantic version and a ratification date, and puts a Sync Impact Report in an HTML comment at the top. Accept it when each principle is testable. “Code should be clean” is not. -
Specify. Describe what and why, never how.
The skill picks the next number, creates
specs/001-audit-log-export/spec.mdfrom the template and writes a quality checklist tospecs/001-audit-log-export/checklists/requirements.md. Numbering is sequential unlessinit-options.jsonsays otherwise. Branch creation happens only if you installed thegitextension. An illustrative excerpt in the template’s format (FR and SC numbering, Given/When/Then scenarios):### User Story 1 - Export filtered audit log (Priority: P1)**Acceptance Scenarios**:1. **Given** an admin and 1,200 events in the last 30 days, **When** they exportwith that range, **Then** they receive a CSV with 1,200 data rows and a header row.2. **Given** a member without the admin role, **When** they request an export,**Then** the request is refused and no file is produced.### Functional Requirements- **FR-001**: System MUST let admins export audit events for a date range of at most 90 days.- **FR-004**: System MUST refuse exports over 100,000 rows with a message to narrow the range.- **FR-006**: System MUST record each export as an audit event [NEEDS CLARIFICATION: shouldthe export event appear in exports that cover its own timestamp?]### Success Criteria- **SC-001**: An admin completes an export of 30 days of events in under 10 seconds.The skill allows at most three
[NEEDS CLARIFICATION]markers and asks you about them at the end of the run. Accept the spec when every acceptance scenario has a concrete input and output, and no class, table or library name appears in it. -
Clarify (optional, before plan).
/speckit-clarifyasks structured questions about gaps it finds and writes your answers back intospec.md. Run it whenever the spec still has a marker or an edge case nobody decided. -
Plan. Now the how.
speckit-planwritesplan.mdfrom the template, thenresearch.md(every open question resolved),data-model.md,contracts/(here, the export endpoint’s request and response) andquickstart.md, a validation guide for trying the feature.plan.mdincludes a Constitution Check gate that is evaluated before research and again after design. Accept the plan when the Constitution Check passes without an unjustified violation and every file it names exists or is new on purpose. -
Checklist (optional, after plan).
/speckit-checklist securitygenerates a checklist that tests the requirements for completeness and clarity, such as “Is the authorization rule stated for every entry point?”. It is a unit test for the spec, not for the code. -
Tasks.
/speckit-taskswritestasks.md: phases for setup and foundations, then one phase per user story, each task with an ID, an optional[P]marker for tasks that can run in parallel, and a story tag.## Phase 3: User Story 1 - Export filtered audit log (Priority: P1)- [ ] T010 [P] [US1] Integration test: admin exports 30-day range in tests/integration/audit-export.test.ts- [ ] T011 [P] [US1] Integration test: non-admin gets 403 in tests/integration/audit-export.test.ts- [ ] T012 [US1] Implement streaming CSV route in app/api/audit-log/export/route.tsThe tasks template labels test tasks “OPTIONAL - only if tests requested”. Your constitution’s test-first principle is what requests them, which is why the constitution comes first.
-
Analyze (optional, before implement).
/speckit-analyzeis strictly read-only: it cross-checksspec.md,plan.mdandtasks.mdand reports requirements with no task, tasks with no requirement, and constitution conflicts. Constitution conflicts are always CRITICAL. Resolve every CRITICAL finding before you implement. -
Implement.
/speckit-implementworks throughtasks.mdphase by phase and marks each finished task[X]. Before it starts, it counts unchecked items in every file underchecklists/and asks whether to continue if any remain. -
Converge.
/speckit-convergecompares the code withspec.md,plan.md,tasks.mdand the constitution. It never edits code, the spec or the plan. When something is missing it appends a new## Phase N: Convergencesection with numbered tasks, CRITICAL and HIGH first. When nothing is missing it leavestasks.mdbyte-for-byte unchanged and reports “Converged — the implementation satisfies the spec, plan, and tasks.”
Loop implement and converge until the feature converges
Section titled “Loop implement and converge until the feature converges”Steps 8 and 9 form the loop that replaces reading the diff: implement, converge, and implement the appended tasks again until converge appends nothing. Then you open the pull request with the trail.
Run the loop in one session so the agent keeps the feature context, and let your test command be the gate between rounds.
/speckit-implement/speckit-convergeRepeat the pair while converge appends a Convergence phase. If an implement round goes wrong, /rewind (or Esc Esc) returns the session to its checkpoint. To hand the whole loop to the agent, paste the loop prompt below.
The same two skills, with Codex’s spelling.
$speckit-implement$speckit-convergeCodex reads AGENTS.md, so put the test command there (for example npm test && npx playwright test) and the loop prompt can refer to “the test command in AGENTS.md”.
In Agent chat, type /speckit-implement, then /speckit-converge, with the same skill names as Claude Code. Commit after each round that passes the tests, so a bad implement round is one git restore away and tasks.md keeps a clean history.
The round limit matters. A finding that reappears after two rounds usually means the spec and the plan disagree, and another implement pass cannot fix a disagreement between documents.
That table is the evidence bundle for this feature. The reviewer starts from the spec delta and the requirement-to-test table, as described in reviewing an agent’s pull request, and reads code only where the table shows UNTESTED or a file outside the plan.
How do you verify Spec Kit’s output without reading every line?
Section titled “How do you verify Spec Kit’s output without reading every line?”Each gate catches a different class of error, and only some of them are deterministic.
| Gate | What it proves | What it cannot prove | Signs off |
|---|---|---|---|
| Spec review | The right behaviour is written down, with examples | That the code does it | Product owner |
speckit-checklist | The requirements are complete and unambiguous | Anything about the code | Product owner or tech lead |
Constitution Check in plan.md | The design respects the project rules | That the build will | Tech lead |
speckit-analyze | Spec, plan and tasks agree with each other | That any of them is right | Engineer |
| Tests written from acceptance scenarios | The code behaves as specified for those inputs | Behaviour nobody specified | CI |
speckit-converge “Converged” | A model found no unbuilt requirement | Correctness; it is a judgment, not a test | Engineer, then PR reviewer |
Treat “Converged” as necessary, not sufficient. The deterministic proof is the requirement-to-test table: each acceptance scenario becomes a failing test before implementation, as in executable acceptance criteria, and CI runs it. Keep specs/ outside the implementing agent’s write scope with permissions and sandboxes and a CODEOWNERS rule, so the agent cannot make a test pass by changing the spec.
Fix bugs and triage ideas with the bug and assess extensions
Section titled “Fix bugs and triage ideas with the bug and assess extensions”Not every change is a feature. Spec Kit 1.0 bundles two more processes as extensions.
Bug fixing (specify extension add bug) adds three skills that write to .specify/bugs/<slug>/:
/speckit-bug-assess "Exporting a range that ends today returns an empty CSV." slug=export-empty-today/speckit-bug-fix slug=export-empty-today/speckit-bug-test slug=export-empty-todaybug-assess writes assessment.md with the suspected root cause and a remediation, bug-fix applies it and records fix.md, and bug-test writes test.md with one verdict: verified (the symptom no longer reproduces and the critical checks pass), partial or failed. Only verified closes the bug. A fix that restores behaviour the spec already describes needs no new spec, only a test that cites the existing FR ID.
Idea assessment (specify extension add assess) adds speckit-assess-intake, -research, -define, -shape and -decide, writing to .specify/assessments/<slug>/. assess-decide records a verdict of go, needs-clarification or kill in decision.md, and only a go goes on to speckit-specify. Use it when a product owner brings a vague idea: killing an idea before a spec exists is the cheapest outcome the process has.
What does Spec Kit cost in context and ceremony?
Section titled “What does Spec Kit cost in context and ceremony?”Context. The 10 core skills add about 1,300 characters of descriptions to every session (measured on 1.0.12; roughly 300 tokens at four characters per token). Each stage loads its full instructions when it runs: speckit-specify is 18.7 KB and speckit-converge 13.5 KB. That is small next to plugin bundles, which the frameworks overview compares. The bigger cost is the documents themselves: a single feature produces a spec, a requirements checklist, a plan, research, a data model, contracts, a quickstart and a task list, and every later stage reads them.
Invocation. In Claude Code all 10 skills ship with disable-model-invocation: false, so the agent can start a stage on its own when a request matches its description. If you want every stage started by a person, say so in CLAUDE.md and keep the stages as explicit commands. To enforce it, set disable-model-invocation: true in each speckit-* SKILL.md; specify integration status then reports those files as modified managed files, and specify integration upgrade --force may overwrite them. Choose deliberately.
Ceremony. Use the stage count as the decision:
| Change | Use Spec Kit? | Instead |
|---|---|---|
| New feature with several user stories, or one that a product owner must approve | Yes, full cycle | — |
| Greenfield service or product | Yes, starting with the constitution | — |
| Bug that restores specified behaviour | The bug extension | A failing test that cites the FR ID |
| Small change in a large existing codebase | Usually no | OpenSpec, which records a spec delta per change |
| Refactor with no behaviour change | No | Tests before and after; no spec |
| Spike or prototype you will throw away | No | Plan mode in your agent |
What breaks when you run Spec Kit?
Section titled “What breaks when you run Spec Kit?”Old flags and spellings fail. Symptom: specify init --ai claude errors, or /speckit.specify does nothing. Recovery: use --integration, and the per-tool spelling /speckit-specify or $speckit-specify.
The init picks the wrong agent. Symptom: a CI or scripted init produces Copilot files. Recovery: always pass --integration; switch an existing repo with specify integration switch claude or add one with specify integration install.
The agent works on the wrong feature. Symptom: speckit-plan writes into specs/001-… while you meant 002. The current feature is stored in .specify/feature.json, which is per checkout and ignored by Git. Recovery: name the feature directory in your prompt, and give each parallel agent its own worktree so each has its own pointer.
No tests were generated. Symptom: tasks.md has no test tasks and implement goes straight to code. Recovery: add a test-first principle to the constitution, or say “include test tasks for every acceptance scenario” in the speckit-tasks prompt, then run speckit-analyze.
Converge never converges. Symptom: every round appends the same finding. Recovery: stop the loop and run speckit-analyze. A repeating finding is almost always a conflict between spec.md and plan.md; fix the document, with the owner’s approval, then implement again.
The agent edits the spec to match its code. Symptom: an implementation PR changes spec.md and the tests pass. Recovery: deny writes to specs/ in implementation sessions, require the spec owner in CODEOWNERS, and reject the PR.
The team stops reviewing the documents. Symptom: specs are approved minutes after they are generated, and bugs trace back to requirements nobody read. An unreviewed spec moves the hallucination upstream. Recovery: make the product owner’s spec approval and the tech lead’s plan approval explicit PR checks, and use the adversarial review prompt above before approval.
Ceremony swamps small work. Symptom: developers skip Spec Kit for anything under a day and the process decays. Recovery: publish the decision table above and let small changes go through the bug extension or OpenSpec.