Skip to content

Superpowers: A Disciplined Development Workflow for Coding Agents

Superpowers (obra/superpowers, MIT, version 6.4.2) is a plugin of 15 skills plus a session-start bootstrap that makes a coding agent agree a design, isolate work in a worktree, write a reviewed plan, implement with strict TDD, review each task, and show test output before claiming done. It steers the model; CI and human review still enforce.

You ask the agent for “organization API keys with rotation”, and 20 minutes later it has touched 14 files, invented a permission model nobody asked for, and says “all tests pass” without showing a single run. This page is for developers and tech leads who want the agent to ask first, plan in the open, and prove its claims — and who need to know what that discipline costs.

  • Install lines for Claude Code, Codex and Cursor, and a 30-second check that the bootstrap fired.
  • Which version each marketplace serves, and what 6.4.1 vs 6.4.2 changes in your plans.
  • The eight-stage workflow and the artifact each stage leaves, so you review a design and a plan instead of every diff.
  • Four copy-paste prompts, per-invocation token costs, and a failure table with recovery steps.
FactValue (read 2026-09-26 unless noted)Source
Upstream version6.4.2, released 2026-09-25plugin.json, release notes
Anthropic’s claude-plugins-officialpins a commit that is 6.4.1marketplace manifest
Author’s superpowers-marketplace6.4.2claude plugin install in Claude Code 2.1.283, 2026-09-26
Codex openai-curated (openai-api-curated with an API key)manifest says 6.3.0openai/plugins
Skills15, plus a SessionStart hook that injects using-superpowersclaude plugin details
Always-on context~838 tokens per sessionclaude plugin details, Claude Code 2.1.283, 2026-09-26
Popularity1,009,371 installs on the Claude plugin directory; 291.7k GitHub starsclaude.com/plugins/superpowers; GitHub, read 2026-09-26

The install count covers Claude Code installs from Anthropic’s directory only, and stars measure attention on the repository, not active use. Neither says the workflow improves your code — that is what your own pilot measures.

How Superpowers works: skills, a bootstrap and harness adapters

Section titled “How Superpowers works: skills, a bootstrap and harness adapters”

Superpowers has three layers, and the difference between them explains most installation problems.

  1. Skills. Fifteen SKILL.md procedures: brainstorming, using-git-worktrees, writing-plans, subagent-driven-development, executing-plans, dispatching-parallel-agents, test-driven-development, systematic-debugging, requesting-code-review, receiving-code-review, verification-before-completion, finishing-a-development-branch, writing-skills, using-superpowers and diagnosing-superpowers.
  2. The bootstrap. using-superpowers tells the agent to check for an applicable skill before it answers, asks a question or touches the repository. The plugin injects it at session start; without it, the skills sit on disk and fire only when the model happens to pick them.
  3. Harness adapters. Manifests and tool mappings for each host. Claude Code and Cursor use a session-start hook; Codex declares no hooks and relies on its native skill discovery.

The precedence is explicit: your instructions — direct requests and CLAUDE.md / AGENTS.md alike — override Superpowers skills, which override the agent’s defaults. If you want to learn skills and plugins in general first, start with the ecosystem overview and the plugins section.

Install Superpowers in Claude Code, Codex and Cursor

Section titled “Install Superpowers in Claude Code, Codex and Cursor”

Install the native plugin for each agent you use. The README is explicit that each harness needs its own install.

Inside a session, pick one marketplace:

/plugin install superpowers@claude-plugins-official

That gives you Anthropic’s pinned commit (6.4.1 on 2026-09-26). To track upstream (6.4.2) instead:

/plugin marketplace add obra/superpowers-marketplace
/plugin install superpowers@superpowers-marketplace

The terminal equivalents are claude plugin marketplace add obra/superpowers-marketplace and claude plugin install superpowers@superpowers-marketplace (verified in Claude Code 2.1.283, 2026-09-26). Skills are namespaced: force a stage with /superpowers:brainstorming.

The upstream README also documents Antigravity, Devin CLI, Factory Droid, Gemini CLI, GitHub Copilot CLI, Grok Build CLI, Kimi Code, OpenCode, Pi, Qwen Code, Hermes Agent and Muse — copy those commands from the installation section, not from older articles. For Google’s agents, say “Antigravity CLI (personal accounts) / Gemini CLI (enterprise, API key)”: since 2026-06-18 Gemini CLI no longer serves free, Google AI Pro and Ultra users, who moved to Antigravity CLI. Superpowers dropped Gemini CLI in 6.1.0 and restored it in 6.2.0, so gemini extensions install https://github.com/obra/superpowers works again for Code Assist Standard or Enterprise, Google Cloud and paid API-key users.

Start a fresh session and send the README’s own smoke test:

Let's make a react todo list

A working install triggers brainstorming: the agent asks why you want the app and what it must do before it writes any code. If it starts scaffolding, the bootstrap is not loaded. In Claude Code, claude plugin list shows what is installed and claude plugin details superpowers shows the 15 skills, the SessionStart hook and the always-on token cost.

Which Superpowers version did you get, and why it matters

Section titled “Which Superpowers version did you get, and why it matters”

The three marketplaces served three versions in the same week. The differences are behavioural, not cosmetic:

You haveWhat changes in practice
6.3.0 (Codex openai-curated)No diagnosing-superpowers. executing-plans is the old stub that stops every few tasks to check in. Brainstorming already scales ceremony (spike, bounded, architectural). The package ships its bundled scripts non-executable, so the SDD helpers fail with Permission denied (fixed upstream in 6.4.1).
6.4.1 (Claude official)diagnosing-superpowers; executing-plans rebuilt as Native execution with one whole-branch review; you must review the saved plan before anything runs; plans carry a Review Focus section; TDD runs the project’s whole suite. Plans still spell out complete code.
6.4.2 (upstream, author’s marketplace)writing-plans records decisions — signatures, test names and assertions, the spec’s values, the verification command — instead of full code. It was prompted by some frontier models implementing the project while “planning”. The maintainers report plans in about a quarter of the time and a third of the tokens in their reproduction (release notes, 2026-09-25).

If your team mixes Claude Code and Codex, the same prompt can produce a code-heavy plan on one and a decision-level plan on the other. Record the version in your pilot notes, and prefer the author’s marketplace in Claude Code if you want 6.4.2 now.

Each stage leaves an artifact you can review instead of the diff. The stages are the same in Claude Code, Codex and Cursor; what differs is the version each marketplace serves (see above) and the Claude Code-only nested controller.

  1. Brainstorm and classify. brainstorming asks why you want the change, then classifies it as a spike, a bounded change or architectural work. A bounded change gets a short design in chat; architectural work gets a written spec at docs/superpowers/specs/YYYY-MM-DD-<topic>-design.md. Every path stops for your approval.
  2. Isolate. using-git-worktrees detects an existing isolated workspace or creates a worktree on a new branch, installs dependencies and runs the baseline tests. A red baseline stops the run and asks you what to do.
  3. Plan. writing-plans saves docs/superpowers/plans/YYYY-MM-DD-<feature-name>.md: per-task files, interfaces, test assertions, a verification command with its expected output, and a Review Focus list of inputs the spec implies but no test covers yet. You review the plan before execution.
  4. Choose execution. The handoff offers two modes and recommends one. Subagent-driven dispatches a fresh implementer per task and a task review (spec compliance plus code quality) after each. Native (executing-plans) implements every task in the current session and gets one whole-branch review at the end — the cheapest mode.
  5. Implement under TDD. test-driven-development requires a test that fails for the right reason, the smallest change that passes, then refactoring. Green means the project’s whole test command, not only the new test file.
  6. Review. Task reviews loop up to five fix rounds before the controller adjudicates. A final reviewer reads the integrated branch.
  7. Verify. verification-before-completion requires the actual output of the test, build, type-check or lint command before any “done”, “fixed” or “passing”.
  8. Finish. finishing-a-development-branch reruns the tests and offers to merge locally, push and open a PR, or keep the branch. Discarding work happens only on your explicit request, with typed confirmation.

Example: organization API keys from vague request to PR

Section titled “Example: organization API keys from vague request to PR”

The request is: Add organization-level API keys with rotation and audit history. A healthy run looks like this:

StageArtifact you should seeYour checkpoint
BrainstormQuestions on ownership, key visibility, revocation, permissions, retention and migration; two or three designs with trade-offs; classified as architecturalApprove behaviour and threat model
Specdocs/superpowers/specs/2026-09-26-organization-api-keys-design.md, committedRead it: it is shorter than the code
IsolationA worktree with a passing baselineDecide on any pre-existing failure
PlanSchema, API, UI, migration and test tasks; Review Focus names “revoked key used during rotation” and “audit row on failed rotation”Approve scope; add missing edge cases
ExecutionPer task: failing test, minimal code, passing suite, commit, reviewer verdictAnswer product questions only
Final reviewWhole-branch findings, fixed and re-reviewedDecide whether the branch is ready
VerificationFresh migration, test, type-check, lint and build outputCompare against acceptance criteria
FinishPR openedNormal PR review and CI

The value is not that the agent “made a plan”: the design, plan, test evidence and review findings stay inspectable after compaction or a handoff.

The last prompt needs 6.4.1 or later, where diagnosing-superpowers exists.

How you verify Superpowers output without reading every line

Section titled “How you verify Superpowers output without reading every line”

Superpowers moves your review effort from the diff to three short documents and a set of machine gates:

  • Before code: you approve the spec (behaviour, threat model, non-goals) and the plan (task list, test assertions, Review Focus). This is where a human catches “built the wrong thing”.
  • During execution: each task commit must follow a failing test, and each task has a reviewer verdict. Spot-check two or three tasks: the test should fail on the parent commit and pass on the task commit.
  • Before merge: the verification output, CI on the PR, and your repository’s required review. Add the gates Superpowers cannot provide — coverage or mutation thresholds and required status checks — so a skipped TDD step fails loudly instead of silently. See the evidence bundle a PR should carry and how to review agent pull requests.
  • Sign-off: the developer who ran the session owns the spec and plan approval; the PR reviewer (or tech lead, for architectural work) signs off on the evidence, not on line-by-line reading.

What Superpowers costs per session and per invocation

Section titled “What Superpowers costs per session and per invocation”

Claude Code’s own projection (claude plugin details, 2.1.283, 2026-09-26):

CostTokens
Always on, every session (skill listing + bootstrap)~838
brainstorming, each time it fires~6.3k
subagent-driven-development, each time it fires~11.8k, plus a fresh context per implementer and per reviewer

The always-on cost is low for its family — Compound Engineering is ~2,989 and Everything Claude Code ~41,515 on the same measurement. The spend comes from execution. To keep it down:

  • Choose Native execution for plans of small, same-shape tasks; it uses one context plus one final review.
  • In Claude Code, ask for the SDD controller to run as a nested subagent on a mid-tier model; the 6.4.1 release notes (2026-09-18) report about half the cost and wall-clock time, measured by the maintainers. It is opt-in.
  • Let brainstorming classify small work as bounded rather than forcing the full spec path.

When to use Superpowers and when to skip it

Section titled “When to use Superpowers and when to skip it”

Use the full workflow when requirements are ambiguous, the change spans several layers, behaviour can be tested, or the work will be handed across sessions or people.

Skip it, or use single skills, for one-line or mechanical fixes, throwaway spikes, repositories without a test runner (TDD and verification then produce ceremony without a signal), or a process that already has stronger gates. systematic-debugging and verification-before-completion are worth having on their own; the testing and debugging skills page covers them without the framework.

If you need a durable, living specification rather than a per-feature design doc, compare spec-driven frameworks such as Spec Kit and OpenSpec. For the whole field, see the frameworks comparison.

What breaks with Superpowers and how to recover

Section titled “What breaks with Superpowers and how to recover”
SymptomCauseRecovery
The agent jumps straight to codeBootstrap not loaded: a portable npx skills add copy, or the hook blockedInstall the native plugin; run the smoke test; check claude plugin details superpowers for the SessionStart hook
Skills appear twiceInstalled from two marketplaces, or plugin plus npx skills addRemove one copy per agent
Plans are pages of code on one tool, short on anotherVersion drift: 6.4.1 or 6.3.0 vs 6.4.2Align versions; in Claude Code install from superpowers-marketplace
Execution stops every few tasks to check inCodex 6.3.0 executing-plansExpect it until openai-curated updates, or choose Subagent-driven
Codex: skill scripts fail with Permission deniedopenai-curated 6.3.0 package strips executable bits (fixed in 6.4.1)Invoke the script through its interpreter (bash scripts/<name>.sh) or wait for openai-curated to ship 6.4.1+
Codex asks to trust a SessionStart hookA pre-6.1.1 packageUpdate the plugin; Codex versions from 6.1.1 declare no hooks
Windows: bootstrap silently never loadsHook shell quoting before 6.2.0Update to 6.2.0+, use Claude Code 2.1.81 or later, and install Git for Windows
Two plans in one checkout overwrite each other’s state.superpowers/sdd/ shared before 6.2.0 (same basename before 6.4.1)Update; still give each concurrent session its own worktree
“Tests pass” with no outputVerification skippedPaste the evidence prompt above; treat the run as failed until output appears
TDD skippedModel non-complianceTreat as a failed run; add a CI or review gate instead of stronger wording
Worktree removal would lose filesUncommitted work in the worktree6.3.0+ stops and names the files; commit or stash, never --force
A skill fired or stayed silent unexpectedlyTrigger mismatchRun the diagnosis prompt above (6.4.1+) and file the scrubbed bundle upstream

Security checks before you adopt Superpowers

Section titled “Security checks before you adopt Superpowers”

Skills are instructions an agent with shell and network access acts on, so review them like code. Follow skill supply-chain security, then check these specifics:

  • The Claude Code, Cursor and several other adapters run a SessionStart hook; read hooks/ before you install, and re-read it on updates.
  • The optional brainstorming Visual Companion is a local Node server bound to 127.0.0.1 by default. Binding it to 0.0.0.0 widens the trust boundary.
  • The companion loads a Prime Radiant logo URL that includes the Superpowers version. Set SUPERPOWERS_DISABLE_TELEMETRY=true to stop it; it also honours DISABLE_TELEMETRY and CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC.
  • An issue opened on 2026-09-10 (#2277) reports that systematic-debugging ships pressure-test fixtures written as imperative scenarios, which get installed into runtime directories. Check whether your version still has them.
  • receiving-code-review treats review comments as claims to verify. External PR comments are still untrusted input: never let one trigger secrets access or unrelated commands.

The maintainers run superpowers-evals, a lab that drives real agent CLIs through triggering, TDD, review and verification scenarios. It measures workflow compliance; its few no-plugin comparisons cover single skills (6.4.1 found the old executing-plans stub no better than no plugin), not the whole bundle, so it does not show that the bundle beats an agent without it. No independent, controlled comparison of the whole framework was available on 2026-09-26, so run your own.

  1. Pick one pilot. A medium feature with clear acceptance criteria — not a toy, not the riskiest migration.
  2. Record a baseline on comparable work without the plugin: first-pass acceptance, human corrections, review defects, wall-clock time and tokens.
  3. Pin and record the Superpowers version, marketplace, agent version and model for every participant, so drift does not look like a result.
  4. Set policy in CLAUDE.md / AGENTS.md: who approves specs and plans, when TDD exceptions are allowed (visual work, generated files), and which commands always need confirmation.
  5. Run each concurrent session in its own worktree, with no production secrets in the environment. See Git worktrees for parallel agents.
  6. Measure with and without the plugin. In Claude Code, claude plugin eval runs eval cases against the plugin and reports the score delta against a no-plugin baseline. Put your cases (case.yaml, or prompt.md plus graders/*.md) in a directory below the plugin — a local checkout works as the target — and pass its name with --eval-dir; the no-plugin baseline arm (--ablation with-without) is on by default (Claude Code 2.1.285, checked 2026-09-26).
  7. Keep, trim or remove skills after the pilot, and re-run it after every model, agent or plugin major update.