Spec-Driven Frameworks Compared: Spec Kit vs OpenSpec vs BMAD vs Superpowers
Spec-driven frameworks make a coding agent write reviewable Markdown (requirements, design and tasks) before it writes code. GitHub Spec Kit suits greenfield features that need a gated trail; OpenSpec suits changes to existing code, because archiving a change merges it into a living spec. BMAD adds product roles, and Superpowers adds a design document only when the work is architectural.
Your team has agreed that agents should work from specs, and now four people have four favourites. One developer ran Spec Kit on a side project, another swears by OpenSpec, the product manager saw a BMAD demo, and half the team already has Superpowers installed. You need one choice, a pilot that proves it, and a rule for which documents a human signs off.
This page is for the developer who will run the pilot and the tech lead who picks the framework. It runs one feature, “export the audit log as CSV”, through Spec Kit and OpenSpec side by side, then shows what BMAD, Superpowers and cc-sdd do with the same request.
What you’ll get from comparing spec-driven frameworks on one feature
Section titled “What you’ll get from comparing spec-driven frameworks on one feature”- A decision table keyed on whether the code exists and who approves the documents.
- Verified installs for five frameworks in three tools.
- The audit-log export run through Spec Kit and OpenSpec, plus a follow-up change that shows where they really differ.
- A review-gate table: what a person approves and who signs off.
- Each framework’s always-on context cost and the traps old tutorials teach.
Which spec-driven framework should you pick?
Section titled “Which spec-driven framework should you pick?”Start from the unit of work each framework is built around. Spec Kit thinks in features, OpenSpec in changes to an existing system, BMAD in product documents handed between roles, and Superpowers in a single task whose size decides whether a design document is written at all.
| Framework | Unit of work | Artifacts it writes | Keeps a current spec of the system? | Human gates | Claude Code / Codex / Cursor | Pick it when | Avoid when |
|---|---|---|---|---|---|---|---|
| Spec Kit (GitHub) | Feature | .specify/memory/constitution.md, then specs/NNN-<feature>/spec.md, plan.md, tasks.md and supporting files | No, one folder per feature | After every stage, you drive each one | Skills in all three | Greenfield products and large features that need a constitution and a full trail | Small changes in a large codebase |
| OpenSpec (Fission AI) | Change | openspec/changes/<change>/proposal.md, specs/ deltas, design.md, tasks.md | Yes: archive merges the deltas into openspec/specs/ | Proposal review before apply | Commands and skills in Claude Code and Cursor, skills in Codex | Existing codebases with a steady stream of changes | You need a governance trail per stage |
| BMAD Method (BMad Code) | Product document | Product brief, PRD, architecture, SPEC.md, epics and stories | Per document | Per document, with persona hand-offs | Skills or plugins in all three | Product work with several stakeholders and epics | A solo developer shipping small features |
| Superpowers (design-doc path) | Task, classified as spike, bounded or architectural | For architectural work only: docs/superpowers/specs/YYYY-MM-DD-<topic>-design.md, then a plan in docs/superpowers/plans/ | No | Design approval, then plan approval | Plugin in all three | Agents skip tests; you want a spec only for big work | You need a spec for every change |
| cc-sdd (gotalab) | Spec | requirements.md (EARS), design.md, tasks.md per spec | Per spec | Between phases | Skills; Cursor support is beta | You want Kiro’s spec shape outside Kiro | You want the largest community and docs |
Kiro specs, in AWS’s Kiro IDE and CLI, write requirements.md (EARS), design.md and tasks.md under .kiro/specs/<feature>/, per secondary sources (kiro.dev was unreadable on 2026-09-26). cc-sdd gives other agents that shape; see the Kiro comparison.
Greenfield or brownfield: which question decides?
Section titled “Greenfield or brownfield: which question decides?”Answer these in order and stop at the first yes.
- Is most of the code still to be written, or does a product owner need to approve a feature before anyone plans it? Use Spec Kit. Its constitution and per-stage skills give you a gate at every step.
- Are you changing a system that already works, several times a week? Use OpenSpec. A proposal records only the delta, and archive keeps
openspec/specs/describing the system as it is now. - Does the work start as a product idea that needs a brief, a PRD and an architecture before a spec? Use BMAD’s planning track, then build story by story.
- Is your problem that the agent skips tests and claims success without evidence, not that nobody writes specs? Use Superpowers. It writes a design document only for architectural work.
- None of the above? Use plain Markdown and the CI checks from spec-driven development. A framework adds ceremony; it does not make a spec true.
Mixed teams often use OpenSpec for daily changes and Spec Kit for the occasional new service. Never run two on the same change: the agent gets two sources of truth.
Install the spec-driven frameworks in each agent
Section titled “Install the spec-driven frameworks in each agent”Spec Kit and OpenSpec install project files through their own CLIs. BMAD, Superpowers and cc-sdd install as skills or plugins. Run the terminal lines in your repository root.
# Spec Kit (Python 3.11+ and uv): 10 speckit-* skills in .claude/skills/uv tool install specify-clispecify init --here --integration claude
# OpenSpec (Node >= 20.19.0): 6 skills plus /opsx:* commandsnpm install -g @fission-ai/openspec@latestopenspec init --tools claude
# BMAD, stable track (npm 6.12.0)npx bmad-method install --directory . --modules bmm --tools claude-code --yes
# cc-sdd: Claude Code skills are the default targetnpx cc-sdd@latest# In the Claude Code session# BMAD, plugin track (6.13.0-next)/plugin marketplace add bmad-code-org/bmad-plugins/plugin install bmad-method@bmad
# Superpowers (Anthropic's marketplace installed 6.4.1 on 2026-09-26; upstream was 6.4.2)/plugin install superpowers@claude-plugins-official# Spec Kit: skills in .agents/skills/, invoked as $speckit-<stage>uv tool install specify-clispecify init --here --integration codex
# OpenSpec: skills only (no commands for Codex), invoked as $openspec-proposenpm install -g @fission-ai/openspec@latestopenspec init --tools codex
# BMAD, stable track: skills in .agents/skills/npx bmad-method install --directory . --modules bmm --tools codex --yes# BMAD, plugin trackcodex plugin marketplace add bmad-code-org/bmad-pluginscodex plugin add bmad-method@bmad
# cc-sddnpx cc-sdd@latest --codex-skillsFor Superpowers, run /plugins in a Codex session and install Superpowers from the curated list.
# Spec Kit: the integration id is cursor-agent, not cursoruv tool install specify-clispecify init --here --integration cursor-agent
# OpenSpec: .cursor/commands/ and .cursor/skills/, invoked as /opsx-proposenpm install -g @fission-ai/openspec@latestopenspec init --tools cursor
# BMAD, stable track: skills in .agents/skills/npx bmad-method install --directory . --modules bmm --tools cursor --yes
# cc-sdd: Cursor support is betanpx cc-sdd@latest --cursor-skillsFor Superpowers, type /add-plugin superpowers in Agent chat. That command comes from the Superpowers README; Cursor’s own marketplace could not be checked on 2026-09-26.
BMAD has two install tracks on 2026-09-26, and they ship different skills. BMAD’s tutorial uses the stable npm installer (6.12.0). The plugin (bmad-method@bmad, 6.13.0-next, 21 skills) follows the main branch and ships bmad-create-epics-and-stories where the main docs describe bmad-preview-ticketing. Pick one track for the whole team and pin it, or two developers will follow two different processes under one name.
Run the audit-log export through Spec Kit and OpenSpec
Section titled “Run the audit-log export through Spec Kit and OpenSpec”The feature: team admins export their organization’s audit log as CSV, for a date range of at most 90 days; members who are not admins cannot export. Both runs use Claude Code spelling. In Codex, replace /speckit- with $speckit- and /opsx:propose with $openspec-propose; in Cursor, OpenSpec’s colon becomes a hyphen. Spec Kit’s /speckit-* names are the same in Cursor (skills in .cursor/skills/).
| Stage | Spec Kit | OpenSpec |
|---|---|---|
| Project rules | /speckit-constitution once, writes .specify/memory/constitution.md | Optional context: and rules: in openspec/config.yaml |
| Explore / clarify | /speckit-clarify resolves open questions in the written spec | /opsx:explore investigates before the proposal |
| Write the requirement | /speckit-specify, writes specs/001-audit-log-export/spec.md | /opsx:propose, writes proposal.md, specs/audit-log-export/spec.md (a delta), design.md and tasks.md in one step |
| Plan | /speckit-plan, writes plan.md, research.md, data-model.md, contracts/, quickstart.md | Already in design.md |
| Tasks | /speckit-tasks, then /speckit-analyze | Already in tasks.md |
| Build | /speckit-implement, then /speckit-converge until “Converged” | /opsx:apply |
| Close | Open the PR; the feature folder stays as written | /opsx:archive merges the delta into openspec/specs/audit-log-export/spec.md |
Spec Kit splits planning into six skills, each a point to stop and review. OpenSpec has one review point: propose writes all four artifacts, then stops. Its skill states that the request “authorizes planning only” and waits for a new request before apply.
-
Write the requirement. Paste the same sentence into both frameworks so you compare the frameworks, not the prompts.
-
Review what each one wrote. Spec Kit’s
spec.mduses user stories with Given/When/Then scenarios and numbered requirements (FR-001), and it may leave up to three[NEEDS CLARIFICATION]markers for you. OpenSpec writes a delta with### Requirement:and#### Scenario:blocks in WHEN/THEN form. This is the OpenSpec delta from the scratch run, whichopenspec validate --strictaccepted:## PurposeLets organization admins take audit events out of the product as a CSV file for compliance reviews.## ADDED Requirements### Requirement: Admin exports audit log as CSVThe system SHALL let an organization admin export audit events for a date range of at most 90 days as CSV.#### Scenario: Admin exports a 30-day range- **WHEN** an admin exports a 30-day range containing 1,200 events- **THEN** the system returns a CSV with a header row and 1,200 data rows#### Scenario: Non-admin is refused- **WHEN** a member without the admin role requests an export- **THEN** the system refuses the request and produces no fileAccept either document when every scenario has a concrete input and output, and no table, class or library name appears in a requirement.
-
Build. Spec Kit loops
/speckit-implementand/speckit-convergeuntil converge reports “Converged”; the Spec Kit tutorial has the loop prompt. OpenSpec runs/opsx:apply, which works throughtasks.md. In both, the tests written from the scenarios are what prove the code, not the framework’s own verdict. -
Close the change. With OpenSpec,
/opsx:archivevalidates the change, moves it toopenspec/changes/archive/2026-09-26-add-audit-log-csv-export/and createsopenspec/specs/audit-log-export/spec.md. The scratch run printedaudit-log-export: createand+ 1 added. Spec Kit leavesspecs/001-audit-log-export/as the record of that feature.
What happens on the second change?
Section titled “What happens on the second change?”A month later, compliance asks for an actor filter. This is where the two frameworks part ways.
In Spec Kit, the filter becomes feature 002 with its own spec, plan and tasks, or you edit 001’s spec after it shipped. Either way, nothing produces one current description of the export: to learn what the export does today, a reader has to read both folders and work out which one wins.
In OpenSpec, the proposal carries a MODIFIED delta. The skill’s instructions require copying the whole existing requirement block and editing it, not writing a fragment:
## MODIFIED Requirements
### Requirement: Admin exports audit log as CSVThe system SHALL let an organization admin export audit events for a date range of at most 90 days as CSV, optionally filtered to a single actor.
#### Scenario: Admin exports a 30-day range- **WHEN** an admin exports a 30-day range containing 1,200 events- **THEN** the system returns a CSV with a header row and 1,200 data rows
#### Scenario: Non-admin is refused- **WHEN** a member without the admin role requests an export- **THEN** the system refuses the request and produces no file
#### Scenario: Admin filters by actor- **WHEN** an admin exports a 30-day range filtered to an actor who has 40 of the 1,200 events- **THEN** the system returns a CSV with a header row and exactly 40 data rows, all for that actorThe reviewer reads one requirement, sees the new scenario, and after archive openspec/specs/audit-log-export/spec.md describes the export as it is now. That is the whole brownfield argument for OpenSpec in one diff.
How do BMAD, Superpowers and cc-sdd handle the same feature?
Section titled “How do BMAD, Superpowers and cc-sdd handle the same feature?”BMAD sizes the process to the request. A clear, small change goes straight to one skill (Claude Code spelling; in Codex, $bmad-build):
/bmad-build Add CSV export of the audit log for team admins, date range at most 90 days, non-admins refused.bmad-build asks for the decisions it needs, shows a plan, waits for your approval, then writes and checks the code. If the export is one story in a larger compliance epic, the planning track runs first: bmad-product-brief, bmad-prd, bmad-architecture and bmad-spec (which writes specs/spec-<slug>/SPEC.md), then tickets and one build per story, bmad-code-review and bmad-retrospective. On the plugin track the ticket step is bmad-create-epics-and-stories; on the skills route it is bmad-preview-ticketing.
Superpowers classifies the request before it asks anything. Its brainstorming skill defines three paths. A spike needs an approved question and probe. A bounded task gets a short design in chat that you approve. Architectural work gets a written spec at docs/superpowers/specs/YYYY-MM-DD-<topic>-design.md, which you approve before writing-plans saves a plan to docs/superpowers/plans/, and you approve that plan and choose how it is executed before any code. A single CSV route is probably bounded, so no spec file is written. If you want a file for every change, pair Superpowers with OpenSpec.
cc-sdd gives Claude Code, Codex and Cursor Kiro’s artifact shape. Start with discovery and let it route the work:
/kiro-discovery CSV export of the audit log for team admins/kiro-spec-init audit-log-export/kiro-spec-requirements audit-log-export/kiro-spec-design audit-log-export/kiro-spec-tasks audit-log-export/kiro-impl audit-log-exportThat is the Claude Code spelling. In Codex, --codex-skills installs the same skills in .agents/skills/, invoked as $kiro-discovery, $kiro-spec-init and so on (cc-sdd’s agent-compatibility guide, 2026-09-26); Cursor support is beta.
requirements.md comes out in EARS form with acceptance criteria. For an existing system the README adds kiro-steering first and an optional kiro-validate-gap before design. kiro-impl gives each task a fresh implementer running red-green TDD and an independent reviewer when the host has subagents.
Where should humans review, and what proves the code?
Section titled “Where should humans review, and what proves the code?”A spec framework moves review from the diff to the documents. That only works if each gate has an owner and something deterministic stands behind the model’s verdicts.
| Gate | Spec Kit | OpenSpec | BMAD | Superpowers | Who signs off |
|---|---|---|---|---|---|
| Behaviour is right | spec.md, after speckit-clarify | proposal.md and the spec deltas | PRD, then SPEC.md | Written spec (architectural) or chat design (bounded) | Product owner |
| Design fits the system | plan.md with its Constitution Check | design.md | Architecture document | Implementation plan | Tech lead |
| Documents agree | speckit-analyze (model, read-only) | openspec validate --strict (deterministic, structure only) | bmad-code-review after build | Review between tasks | Engineer |
| Code does what the spec says | Tests from acceptance scenarios, then speckit-converge | Tests from scenarios; apply ticks tasks.md | Tests per story | TDD per task, then verification-before-completion | CI, then PR reviewer |
Two rules make the table hold. First, turn every scenario into a failing test before implementation, as in executable acceptance criteria; converge verdicts and review skills are judgments, and only tests are repeatable. Second, keep the spec folders out of the implementing agent’s write scope, with permissions and sandboxes and a CODEOWNERS entry, so an agent cannot make a test pass by editing the requirement.
For OpenSpec, add the structural check to CI:
# CI step: fails the job if any change or spec is malformednpx -y @fission-ai/openspec@1.13.2 validate --all --strict --no-interactiveIt checks documents, not code. It also misses one common slip: a ### Scenario: (three hashes) under a requirement that also has a valid #### Scenario: passes --strict with only an INFO line saying it “is ignored by validation”. (If it is the requirement’s only scenario, --strict fails.) Grep openspec/ for ^### Scenario: if scenarios are your test cases.
That review is the gate a product owner or tech lead reads instead of the code. The pull request then carries the evidence bundle: each requirement mapped to the test that proves it.
What does each spec-driven framework cost in context?
Section titled “What does each spec-driven framework cost in context?”The always-on cost is what a framework adds to every session before you use it. Claude Code 2.1.283 measures it for plugins with claude plugin details <plugin>; for project-local skills, count the descriptions.
| Framework | How it installs | Always-on cost (2026-09-26) | Heaviest single invocation |
|---|---|---|---|
| Spec Kit 1.0.12 | 10 project skills | About 1,150 characters of skill descriptions, roughly 300 tokens (estimate at four characters per token) | speckit-checklist, 22.7 KB of instructions |
| OpenSpec 1.13.2 | 6 project skills plus 6 commands in Claude Code | About 2,200 characters of descriptions, roughly 550 tokens (same estimate) | openspec-explore, 22.8 KB |
BMAD bmad-method@bmad | Plugin, 21 skills | ~1,676 tokens (claude plugin details); bmad-toolbox@bmad adds ~975 | Not measured |
| Superpowers | Plugin, 15 skills and a SessionStart hook | ~838 tokens (claude plugin details) | subagent-driven-development, ~11.8k tokens each time it fires; brainstorming ~6.3k |
The bigger cost is the documents: a Spec Kit feature produces eight, and each later stage reads them. Run /context in Claude Code before and after your pilot’s first stage to measure both.
Pilot a spec-driven framework on one feature
Section titled “Pilot a spec-driven framework on one feature”Only a pilot shows how your reviewers cope with eight documents per feature.
- Pick one real feature that touches existing code and has a product owner, not a toy.
- Install one framework (two at most, if you are choosing between Spec Kit and OpenSpec) in a branch, following the tabs above, and commit the generated files so everyone reviews the same setup.
- Run the feature through it. Record the time from request to approved spec, the number of review rounds per document, and the always-on context cost.
- Turn every scenario into a test before implementation, and let CI decide whether the feature is done.
- Run the second change on the same feature a week later. How easily a reviewer answers “what does this feature do now?” is the result that separates the frameworks.
What breaks when you adopt a spec-driven framework?
Section titled “What breaks when you adopt a spec-driven framework?”Nobody reviews the documents. Symptom: specs are approved minutes after generation, and bugs trace back to requirements nobody read. An unreviewed spec moves the hallucination upstream. Recovery: make spec and design approval explicit pull request checks with named owners, and run the gate review prompt before approval.
The agent edits the spec to match its code. Symptom: an implementation PR changes spec.md or an OpenSpec delta, and the tests pass. Recovery: deny writes to the spec folders in implementation sessions, require the spec owner in CODEOWNERS, and reject the PR.
Two frameworks compete. Symptom: a repository has specs/001-…, openspec/changes/… and docs/superpowers/specs/… for the same capability, and agents cite whichever they found first. Recovery: pick one spec framework per repository and write it in CLAUDE.md or AGENTS.md. Superpowers can stay as build discipline if its design documents point at the one spec.
The per-feature spec drifts from the system. Symptom: in Spec Kit or BMAD, specs/001-… describes the export as shipped in September, and nobody can say whether it is still true. Recovery: either move brownfield work to OpenSpec, or add the traceability and drift checks from spec-driven development.
Ceremony swamps small changes. Symptom: developers bypass the framework for anything under a day, and the process decays. Recovery: publish a size rule. Bugs and refactors go through a failing test; OpenSpec proposals, BMAD’s bmad-build or a Superpowers bounded design cover small changes; full Spec Kit or BMAD planning is for features with several user stories.
Malformed scenarios pass validation. Symptom: an OpenSpec scenario written as ### Scenario: under a requirement that already has one valid #### Scenario: is ignored, and openspec validate --strict stays green (1.13.2). Recovery: grep for three-hash scenarios in CI, and generate tests from scenarios with a script that fails when it finds none.