Agentic development frameworks compared
An agentic development framework is a packaged methodology for coding agents: a fixed sequence of steps, the artifacts each step leaves, and the skills, hooks or plugins that enforce it. The frameworks fall into three families — spec-first pipelines, discipline bundles and orchestration harnesses — and differ most in ceremony and in always-on context cost.
Your team has a Slack thread arguing Superpowers against Spec Kit against “we tried BMAD once,” and the GitHub star counts are doing the arguing. This page is for the developer choosing a framework and the tech lead or CTO standardising one: it compares them on what they leave behind and what they cost, then outlines a one-feature pilot that settles the argument with evidence.
What you’ll walk away with from this comparison
Section titled “What you’ll walk away with from this comparison”- The three families and the problem each one solves.
- Two tables covering 16 frameworks: artifacts and approvals, then agent support, context cost, stars and when to avoid each.
- A verified way to measure context cost with
claude plugin details. - The four pilot steps, an audit prompt and the common adoption failures.
Which family of framework solves your problem?
Section titled “Which family of framework solves your problem?”Spec-first pipelines turn intent into reviewed Markdown before any code: Spec Kit, OpenSpec, BMAD, Tessl, Kiro specs and AI-DLC. Reach for one when nobody can say what “done” meant after the fact. See spec-driven frameworks and, for audit trails, enterprise frameworks.
Discipline bundles change how the agent behaves through skills that fire on their own: Superpowers, Matt Pocock’s skills, Compound Engineering, addyosmani/agent-skills, gstack, Ponytail and karpathy-guidelines. Reach for one when the agent skips tests, plans or review. Discipline packs covers pairing them.
Orchestration harnesses add subagents, hooks, memory and loops: GSD Core, Ralph loops, Everything Claude Code (ECC) and the multi-agent harnesses. Reach for one only when you want hours of unattended work and have tests strong enough to stop a bad run. See autonomous loops.
On the autonomy ladder, our reading is: discipline bundles make Level 2 and 3 sessions dependable, spec-first pipelines move your reading from the diff to the spec on the way to Level 4, and orchestration harnesses assume Level 4 — a written spec, a stop condition and tests you trust.
How do the frameworks compare?
Section titled “How do the frameworks compare?”What each framework leaves and who approves it
Section titled “What each framework leaves and who approves it”“Ladder / stage” is our reading of the autonomy ladder level the framework assumes and the lifecycle stages where its artifacts appear.
| Framework | Family | Artifacts it leaves | Human gates | Ladder / stage |
|---|---|---|---|---|
| Superpowers | discipline | design doc, plan | design and plan approval | L2–3 / plan, build |
| Matt Pocock’s skills | discipline | spec, tickets | alignment interview | L2–3 / plan |
| Everything Claude Code | orchestration | none required | per command | L3–4 / build, automate |
| karpathy-guidelines | discipline | none | none | L2–3 / build |
| Ponytail | discipline | none | none (levels) | L2–3 / build |
| Spec Kit | spec-first | .specify/, specs/ per feature | after every step | L4 / plan |
| gstack | discipline | design doc, test plan (in ~/.gstack/projects/), QA reports | per role command | L2–3 / plan, test, ship |
| addyosmani/agent-skills | discipline | spec, plan | per phase command | L2–3 / plan to ship |
| OpenSpec | spec-first | openspec/changes/, living openspec/specs/ | proposal review | L4 / plan |
| BMAD Method | spec-first | PRD, architecture, stories | per document | L4 / plan |
| Compound Engineering | discipline | plans, docs/solutions/ | brainstorm and plan | L3 / plan, ship |
Ralph loop (ralph-loop plugin) | orchestration | plan file, git history | none inside the loop | L4–5 / automate |
| GSD Core | orchestration | .planning/ (roadmap, state, plans) | discuss and acceptance | L4 / plan, automate |
| AI-DLC (AWS Labs) | spec-first | state file, append-only audit trail | after every stage | L4 / plan |
| Tessl spec-driven-development tile | spec-first | .spec.md files linked to tests | requirement gathering | L4 / plan, test |
| Kiro specs | spec-first | .kiro/specs/<feature>/ requirements, design, tasks | per document | L4 / plan |
Where it runs, what it costs and when to avoid it
Section titled “Where it runs, what it costs and when to avoid it”Stars come from the repository pages, read on 2026-09-26; they measure attention, not use. Always-on cost is what the plugin adds to every Claude Code session, as projected by claude plugin details on Claude Code 2.1.283 on the same date. “n/a” means the framework installs as project files, a Git clone or a separate product, not a measured plugin.
| Framework | Claude Code / Codex / Cursor | Always-on cost | Stars | Avoid it when |
|---|---|---|---|---|
| Superpowers | yes / yes / yes | ~838 | 291.7k | trivial edits, no tests |
| Matt Pocock’s skills | yes / yes / yes | ~1,609 | 269.8k | product owners must approve a spec trail |
| Everything Claude Code | yes / yes / yes | ~41,515 | 268k | you will not prune it |
| karpathy-guidelines | plugin / CLAUDE.md copied to AGENTS.md / Cursor rule | n/a (not measured) | 215.2k | you need a workflow, not behavioural rules |
| Ponytail | yes / yes / yes | ~983 | 146.1k | you want it to decide scope, not size |
| Spec Kit | yes / yes / yes | n/a (10 project skills) | 138.9k | small brownfield changes |
| gstack | yes / via ./setup --host / via ./setup --host | n/a (Git clone) | 134.2k | you already have a /review workflow |
| addyosmani/agent-skills | yes / yes / yes | ~3,620 | 99.1k | you run another pack’s /review or /ship |
| OpenSpec | yes / yes / yes | n/a (6 project skills) | 70.4k | you need heavy governance |
| BMAD Method | yes / yes / yes | ~1,676 | 53.5k | solo small features |
| Compound Engineering | yes / yes / yes | ~2,989 | 25.3k | prototypes, or nobody curates learnings |
| Ralph loop | Claude Code plugin only | ~84 | 21.9k (snarktank/ralph) | no test oracle, secrets on the host |
| GSD Core | yes / yes / yes | ~10,700 | 9.9k (archived original 64.5k) | context-sensitive sessions |
| AI-DLC | yes / yes / yes | n/a (native CLI) | 4.8k | you do not need an audit trail |
| Tessl tile | via Tessl CLI | n/a (Tessl tile) | 54 (tile repo) | you want no proprietary binary in the toolchain |
| Kiro specs | no / no / no (Kiro IDE and CLI only) | n/a (built into Kiro) | n/a (product) | your team works in Claude Code, Codex or Cursor |
The Kiro and Tessl rows rest on their vendors’ documentation, which we could not run here; check both before a pilot. Cursor support comes from each project’s README. For one rules file across agents, see rules sync.
When should you use no framework?
Section titled “When should you use no framework?”Skip frameworks for one-line fixes, throwaway spikes and repositories without tests. A framework adds steps, not a test oracle. If the agent cannot run a suite that fails on a wrong answer, start with a short CLAUDE.md or AGENTS.md, acceptance criteria and a test gate.
Measure two frameworks’ context cost before you choose
Section titled “Measure two frameworks’ context cost before you choose”claude plugin details prints a plugin’s components and projected token cost. Install the candidates into a throwaway config directory so your real setup stays untouched.
# Terminal. A scratch config dir keeps ~/.claude clean.export CLAUDE_CONFIG_DIR="$(mktemp -d)"claude plugin marketplace add anthropics/claude-plugins-officialclaude plugin marketplace add EveryInc/compound-engineering-pluginclaude plugin install superpowers@claude-plugins-officialclaude plugin install compound-engineering@compound-engineering-pluginclaude plugin details superpowersclaude plugin details compound-engineeringExpected output, trimmed (Claude Code 2.1.283, 2026-09-26):
superpowers 6.4.1 Skills (15) brainstorming, ... writing-plans, writing-skills Hooks (1) SessionStart (harness-only — no model context cost)Projected token cost Always-on: ~838 tok added to every session subagent-driven-development ... ~11.8k brainstorming ... ~6.3k
compound-engineering 3.29.0 Skills (36) ce-brainstorm, ... lfgProjected token cost Always-on: ~2,989 tok added to every sessionThe numbers drift slightly between releases of Claude Code and of each plugin.
Always-on tokens are paid in every session; on-invoke tokens each time a skill fires. Superpowers is cheap to keep installed, but its subagent-driven-development skill costs about 11.8k tokens a run. The official marketplace pins a commit, so it installed Superpowers 6.4.1 while the author’s repository was on 6.4.2.
Codex and Cursor have no equivalent we could verify: codex plugin in codex-cli 0.157.1 offers add, list, marketplace and remove, and no cost report. There, compare the context meter in a fresh session before and after enabling the framework. The context cost page explains what that overhead does to long sessions.
Pilot a framework on one feature
Section titled “Pilot a framework on one feature”The pilot is the bake-off protocol at framework scale; the scorecard and decision rule live in pilot design.
-
Pick one framework for the failure you are fixing, at most two, never two from the same family.
-
Install it at project scope on a branch. In Claude Code,
--scope projectwrites the marketplace and plugin into.claude/settings.json; commit it.Terminal window claude plugin marketplace add EveryInc/compound-engineering-plugin --scope projectclaude plugin install compound-engineering@compound-engineering-plugin --scope projectclaude plugin details compound-engineeringTerminal window codex plugin marketplace add EveryInc/compound-engineering-plugincodex plugin add compound-engineering@compound-engineering-pluginCodex has no project scope for plugins (codex-cli 0.157.1); the install is user-wide, so remove it after the pilot with
codex plugin remove compound-engineering@compound-engineering-plugin./add-plugin compound-engineeringType it in the Agent chat. The command comes from the project’s README, not from a test.
-
Run one real feature through it, with acceptance criteria written as tests first, and the same feature without the framework on a baseline branch.
-
Score both branches the same way, then adopt or uninstall. The tech lead reads the scores and the evidence bundle, not the diff; the framework’s documents are inputs, the tests are the proof.
What goes wrong when teams adopt a framework?
Section titled “What goes wrong when teams adopt a framework?”| Failure | What you see | Recovery |
|---|---|---|
| Two process frameworks at once | Two /review commands, conflicting plans, doubled always-on cost | Keep one per family; pair a process pack only with Ponytail |
| Treating artifacts as proof | A beautiful spec, and a bug in production | Gate the merge on tests and acceptance criteria |
| Installing the whole bundle | Sessions start tens of thousands of tokens down (ECC ~41,515) | Measure with claude plugin details; disable what you do not use |
| An unbounded loop | A Ralph loop runs all night on a goal it cannot meet | Pass --max-iterations, run in a sandbox, prefer built-in /goal for one session |
Where to go next with frameworks
Section titled “Where to go next with frameworks”Frequently asked questions
What is an agentic development framework?
A packaged methodology for coding agents: a fixed sequence of steps, the artifacts each step leaves behind, and the skills, commands, hooks or plugins that enforce the sequence. Most now install as Claude Code or Codex plugins or as project-local Agent Skills.
Which framework should a team start with?
Pick by the problem you have. A missing spec trail points to Spec Kit (greenfield) or OpenSpec (brownfield); agents that skip steps point to Superpowers or Matt Pocock's skills; lessons that never stick point to Compound Engineering; long unattended runs need GSD Core or a Ralph loop plus a strong test oracle. Pilot one on a single feature before adopting it.
How much context does a framework cost?
Claude Code's claude plugin details reports it. On Claude Code 2.1.283 (2026-09-26) the always-on cost ranged from about 84 tokens for ralph-loop to about 41,515 for Everything Claude Code, with Superpowers near 838 and Compound Engineering near 2,989.
Do frameworks prove the agent's output is correct?
No. Frameworks produce specs, plans and review notes. Proof still comes from tests, type and lint gates, acceptance criteria and the evidence bundle on the pull request.