Skip to content

Agentic development frameworks compared

An agentic development framework is a packaged methodology for coding agents: a fixed sequence of steps, the artifacts each step leaves, and the skills, hooks or plugins that enforce it. The frameworks fall into three families — spec-first pipelines, discipline bundles and orchestration harnesses — and differ most in ceremony and in always-on context cost.

Your team has a Slack thread arguing Superpowers against Spec Kit against “we tried BMAD once,” and the GitHub star counts are doing the arguing. This page is for the developer choosing a framework and the tech lead or CTO standardising one: it compares them on what they leave behind and what they cost, then outlines a one-feature pilot that settles the argument with evidence.

What you’ll walk away with from this comparison

Section titled “What you’ll walk away with from this comparison”
  • The three families and the problem each one solves.
  • Two tables covering 16 frameworks: artifacts and approvals, then agent support, context cost, stars and when to avoid each.
  • A verified way to measure context cost with claude plugin details.
  • The four pilot steps, an audit prompt and the common adoption failures.

Which family of framework solves your problem?

Section titled “Which family of framework solves your problem?”

Spec-first pipelines turn intent into reviewed Markdown before any code: Spec Kit, OpenSpec, BMAD, Tessl, Kiro specs and AI-DLC. Reach for one when nobody can say what “done” meant after the fact. See spec-driven frameworks and, for audit trails, enterprise frameworks.

Discipline bundles change how the agent behaves through skills that fire on their own: Superpowers, Matt Pocock’s skills, Compound Engineering, addyosmani/agent-skills, gstack, Ponytail and karpathy-guidelines. Reach for one when the agent skips tests, plans or review. Discipline packs covers pairing them.

Orchestration harnesses add subagents, hooks, memory and loops: GSD Core, Ralph loops, Everything Claude Code (ECC) and the multi-agent harnesses. Reach for one only when you want hours of unattended work and have tests strong enough to stop a bad run. See autonomous loops.

On the autonomy ladder, our reading is: discipline bundles make Level 2 and 3 sessions dependable, spec-first pipelines move your reading from the diff to the spec on the way to Level 4, and orchestration harnesses assume Level 4 — a written spec, a stop condition and tests you trust.

What each framework leaves and who approves it

Section titled “What each framework leaves and who approves it”

“Ladder / stage” is our reading of the autonomy ladder level the framework assumes and the lifecycle stages where its artifacts appear.

FrameworkFamilyArtifacts it leavesHuman gatesLadder / stage
Superpowersdisciplinedesign doc, plandesign and plan approvalL2–3 / plan, build
Matt Pocock’s skillsdisciplinespec, ticketsalignment interviewL2–3 / plan
Everything Claude Codeorchestrationnone requiredper commandL3–4 / build, automate
karpathy-guidelinesdisciplinenonenoneL2–3 / build
Ponytaildisciplinenonenone (levels)L2–3 / build
Spec Kitspec-first.specify/, specs/ per featureafter every stepL4 / plan
gstackdisciplinedesign doc, test plan (in ~/.gstack/projects/), QA reportsper role commandL2–3 / plan, test, ship
addyosmani/agent-skillsdisciplinespec, planper phase commandL2–3 / plan to ship
OpenSpecspec-firstopenspec/changes/, living openspec/specs/proposal reviewL4 / plan
BMAD Methodspec-firstPRD, architecture, storiesper documentL4 / plan
Compound Engineeringdisciplineplans, docs/solutions/brainstorm and planL3 / plan, ship
Ralph loop (ralph-loop plugin)orchestrationplan file, git historynone inside the loopL4–5 / automate
GSD Coreorchestration.planning/ (roadmap, state, plans)discuss and acceptanceL4 / plan, automate
AI-DLC (AWS Labs)spec-firststate file, append-only audit trailafter every stageL4 / plan
Tessl spec-driven-development tilespec-first.spec.md files linked to testsrequirement gatheringL4 / plan, test
Kiro specsspec-first.kiro/specs/<feature>/ requirements, design, tasksper documentL4 / plan

Where it runs, what it costs and when to avoid it

Section titled “Where it runs, what it costs and when to avoid it”

Stars come from the repository pages, read on 2026-09-26; they measure attention, not use. Always-on cost is what the plugin adds to every Claude Code session, as projected by claude plugin details on Claude Code 2.1.283 on the same date. “n/a” means the framework installs as project files, a Git clone or a separate product, not a measured plugin.

FrameworkClaude Code / Codex / CursorAlways-on costStarsAvoid it when
Superpowersyes / yes / yes~838291.7ktrivial edits, no tests
Matt Pocock’s skillsyes / yes / yes~1,609269.8kproduct owners must approve a spec trail
Everything Claude Codeyes / yes / yes~41,515268kyou will not prune it
karpathy-guidelinesplugin / CLAUDE.md copied to AGENTS.md / Cursor rulen/a (not measured)215.2kyou need a workflow, not behavioural rules
Ponytailyes / yes / yes~983146.1kyou want it to decide scope, not size
Spec Kityes / yes / yesn/a (10 project skills)138.9ksmall brownfield changes
gstackyes / via ./setup --host / via ./setup --hostn/a (Git clone)134.2kyou already have a /review workflow
addyosmani/agent-skillsyes / yes / yes~3,62099.1kyou run another pack’s /review or /ship
OpenSpecyes / yes / yesn/a (6 project skills)70.4kyou need heavy governance
BMAD Methodyes / yes / yes~1,67653.5ksolo small features
Compound Engineeringyes / yes / yes~2,98925.3kprototypes, or nobody curates learnings
Ralph loopClaude Code plugin only~8421.9k (snarktank/ralph)no test oracle, secrets on the host
GSD Coreyes / yes / yes~10,7009.9k (archived original 64.5k)context-sensitive sessions
AI-DLCyes / yes / yesn/a (native CLI)4.8kyou do not need an audit trail
Tessl tilevia Tessl CLIn/a (Tessl tile)54 (tile repo)you want no proprietary binary in the toolchain
Kiro specsno / no / no (Kiro IDE and CLI only)n/a (built into Kiro)n/a (product)your team works in Claude Code, Codex or Cursor

The Kiro and Tessl rows rest on their vendors’ documentation, which we could not run here; check both before a pilot. Cursor support comes from each project’s README. For one rules file across agents, see rules sync.

Skip frameworks for one-line fixes, throwaway spikes and repositories without tests. A framework adds steps, not a test oracle. If the agent cannot run a suite that fails on a wrong answer, start with a short CLAUDE.md or AGENTS.md, acceptance criteria and a test gate.

Measure two frameworks’ context cost before you choose

Section titled “Measure two frameworks’ context cost before you choose”

claude plugin details prints a plugin’s components and projected token cost. Install the candidates into a throwaway config directory so your real setup stays untouched.

Terminal window
# Terminal. A scratch config dir keeps ~/.claude clean.
export CLAUDE_CONFIG_DIR="$(mktemp -d)"
claude plugin marketplace add anthropics/claude-plugins-official
claude plugin marketplace add EveryInc/compound-engineering-plugin
claude plugin install superpowers@claude-plugins-official
claude plugin install compound-engineering@compound-engineering-plugin
claude plugin details superpowers
claude plugin details compound-engineering

Expected output, trimmed (Claude Code 2.1.283, 2026-09-26):

superpowers 6.4.1
Skills (15) brainstorming, ... writing-plans, writing-skills
Hooks (1) SessionStart (harness-only — no model context cost)
Projected token cost
Always-on: ~838 tok added to every session
subagent-driven-development ... ~11.8k
brainstorming ... ~6.3k
compound-engineering 3.29.0
Skills (36) ce-brainstorm, ... lfg
Projected token cost
Always-on: ~2,989 tok added to every session

The numbers drift slightly between releases of Claude Code and of each plugin.

Always-on tokens are paid in every session; on-invoke tokens each time a skill fires. Superpowers is cheap to keep installed, but its subagent-driven-development skill costs about 11.8k tokens a run. The official marketplace pins a commit, so it installed Superpowers 6.4.1 while the author’s repository was on 6.4.2.

Codex and Cursor have no equivalent we could verify: codex plugin in codex-cli 0.157.1 offers add, list, marketplace and remove, and no cost report. There, compare the context meter in a fresh session before and after enabling the framework. The context cost page explains what that overhead does to long sessions.

The pilot is the bake-off protocol at framework scale; the scorecard and decision rule live in pilot design.

  1. Pick one framework for the failure you are fixing, at most two, never two from the same family.

  2. Install it at project scope on a branch. In Claude Code, --scope project writes the marketplace and plugin into .claude/settings.json; commit it.

    Terminal window
    claude plugin marketplace add EveryInc/compound-engineering-plugin --scope project
    claude plugin install compound-engineering@compound-engineering-plugin --scope project
    claude plugin details compound-engineering
  3. Run one real feature through it, with acceptance criteria written as tests first, and the same feature without the framework on a baseline branch.

  4. Score both branches the same way, then adopt or uninstall. The tech lead reads the scores and the evidence bundle, not the diff; the framework’s documents are inputs, the tests are the proof.

What goes wrong when teams adopt a framework?

Section titled “What goes wrong when teams adopt a framework?”
FailureWhat you seeRecovery
Two process frameworks at onceTwo /review commands, conflicting plans, doubled always-on costKeep one per family; pair a process pack only with Ponytail
Treating artifacts as proofA beautiful spec, and a bug in productionGate the merge on tests and acceptance criteria
Installing the whole bundleSessions start tens of thousands of tokens down (ECC ~41,515)Measure with claude plugin details; disable what you do not use
An unbounded loopA Ralph loop runs all night on a goal it cannot meetPass --max-iterations, run in a sandbox, prefer built-in /goal for one session

Frequently asked questions

What is an agentic development framework?

A packaged methodology for coding agents: a fixed sequence of steps, the artifacts each step leaves behind, and the skills, commands, hooks or plugins that enforce the sequence. Most now install as Claude Code or Codex plugins or as project-local Agent Skills.

Which framework should a team start with?

Pick by the problem you have. A missing spec trail points to Spec Kit (greenfield) or OpenSpec (brownfield); agents that skip steps point to Superpowers or Matt Pocock's skills; lessons that never stick point to Compound Engineering; long unattended runs need GSD Core or a Ralph loop plus a strong test oracle. Pilot one on a single feature before adopting it.

How much context does a framework cost?

Claude Code's claude plugin details reports it. On Claude Code 2.1.283 (2026-09-26) the always-on cost ranged from about 84 tokens for ralph-loop to about 41,515 for Everything Claude Code, with Superpowers near 838 and Compound Engineering near 2,989.

Do frameworks prove the agent's output is correct?

No. Frameworks produce specs, plans and review notes. Proof still comes from tests, type and lint gates, acceptance criteria and the evidence bundle on the pull request.

Edit page

Last updated:

Cite this page — https://developertoolkit.ai/en/ecosystem/frameworks/, developertoolkit.ai