Discipline packs: Pocock skills, gstack, agent-skills, Ponytail and ECC
Discipline packs are skill bundles that change how a coding agent works instead of giving it new tools. Matt Pocock’s skills, gstack, addyosmani/agent-skills and Everything Claude Code (ECC) decide what to build and in which order. Ponytail decides how much code to write. The pairing that holds up is one of the first four plus Ponytail, never two process packs.
You asked the agent for a date filter on the invoice list. It interviewed you, wrote a spec, made tickets and wrote tests first, exactly as the pack promised. The pull request still adds two npm packages, an interface with one implementation, and a config flag nobody will ever set. The process pack made sure the agent built the right thing; nothing in it asked whether the agent built too much of it.
This page is for the developer choosing a pack and the tech lead standardising one for a team. It compares the five most-starred packs on what they cost and where they collide, then shows the pairing that covers both questions.
What you get from pairing a process pack with Ponytail
Section titled “What you get from pairing a process pack with Ponytail”- A comparison of five discipline packs on the question each answers, install route, always-on context cost and collision risk.
- Verified install commands for Claude Code, Codex and Cursor, including the two packs that do not install the way their names suggest.
- A worked
/ponytail-reviewrun on an over-built diff, and what to do with its delete list. - A seven-step workflow that pairs one “what to build” pack with Ponytail, with tests and a correctness review as the gates.
- Three copy-paste prompts, and a table of the failures that stacking packs produces.
How do the five discipline packs compare?
Section titled “How do the five discipline packs compare?”| Pack | Question it answers | Entry points | Always-on cost (Claude Code) | Popularity (GitHub, 2026-09-26) | Pick it when | Avoid it when |
|---|---|---|---|---|---|---|
Matt Pocock’s skills (mattpocock/skills) | Do you and the agent agree on what to build? | grill-with-docs → to-spec → to-tickets → implement → code-review | ~1,609 tokens (25 skills) | 269.8k stars | Misalignment is your main failure; you want small, editable skills | You need a durable spec trail approved by product owners |
gstack (garrytan/gstack) | Which role should look at this next? | /office-hours, /plan-eng-review, /review, /qa, /ship, /retro | Not a plugin; measure with its own gstack-context-bill | 134.2k stars | You want product, review, QA and release roles with built-in browser QA, mostly in Claude Code | You cannot accept hooks in ~/.claude/settings.json |
| addyosmani/agent-skills | Which lifecycle phase are you in? | /spec, /plan, /build, /test, /review, /ship and three more | ~3,620 tokens (34 skills, 4 agents) | 99.1k stars | You want one command per phase and skills that trigger on the kind of work | You already run another pack’s /review or /ship |
Everything Claude Code (affaan-m/ECC) | Can one harness cover everything? | 386 skills, 68 agents, 7 hooks | ~41,515 tokens | 267.6k stars | You will prune it to a handful of parts | Context-sensitive work, or you will not prune |
Ponytail (DietrichGebert/ponytail) | How much code does this need? | /ponytail [lite|full|ultra|off], /ponytail-review, /ponytail-audit | ~983 tokens (6 skills, 3 hooks) | 146.1k stars | Agents over-build: extra dependencies, one-implementation abstractions | You want it to decide scope; it only decides size |
Two readings of this table matter more than the rows. First, the first four packs overlap: each has its own idea of planning and review, so two of them give you two /review commands and two opinions about process. Second, Ponytail overlaps with none of them. It has no planning stage and no documents; it constrains how the agent writes whatever the other pack decided on.
Superpowers (~838 always-on tokens) belongs in the same “what to build” family and has its own page; everything below about pairing applies to it too.
Install Matt Pocock’s skills as a framework
Section titled “Install Matt Pocock’s skills as a framework”Pocock’s skills are small and composable, and several of them are one-line delegates: grill-me only calls grilling, and grill-with-docs calls grilling and domain-modeling. As a framework, the chain is: grill until you agree, turn the conversation into a spec, split it into tickets, implement each ticket test-first, and review against the spec. The Grill Me page covers the interview step in depth.
claude plugin install mattpocock-skillsThe plugin is in the official marketplace, so nothing needs adding first. Plugin skills are namespaced: type /mattpocock-skills:grill-with-docs, /mattpocock-skills:to-spec and so on.
npx skills add mattpocock/skills --skill grill-with-docs grilling domain-modeling \ to-spec to-tickets implement tdd code-review setup-matt-pocock-skills -a codexThe README lists no native Codex plugin yet. Mention a skill with $, for example $grill-with-docs.
npx skills add mattpocock/skills --skill grill-with-docs grilling domain-modeling \ to-spec to-tickets implement tdd code-review setup-matt-pocock-skills -a cursorInvoke with /grill-with-docs, or ask for the skill by name if it does not appear in the menu.
Then run setup-matt-pocock-skills once per repository. It asks which issue tracker to-spec and to-tickets should write to (GitHub, Linear or local files), which triage labels you use, and where docs go. Install through the plugin or through npx skills, not both, or every skill appears twice.
Install gstack without the /review collision
Section titled “Install gstack without the /review collision”gstack is a Git repository with a setup script, not a plugin marketplace: /plugin marketplace add garrytan/gstack installs nothing. It needs Git and Bun.
git clone --single-branch --depth 1 https://github.com/garrytan/gstack.git ~/.claude/skills/gstackcd ~/.claude/skills/gstack && ./setup --prefix--prefix names every skill /gstack-<name>. Without it, gstack’s /review claims the name that Claude Code 2.1.283 already uses as the alias of its bundled /code-review. For Codex and Cursor, clone to ~/gstack and run ./setup --host codex or ./setup --host cursor. Setup also registers a default-on Stop hook in ~/.claude/settings.json; ./setup --no-timeline-stop-hook skips it. The gstack page runs one feature through every role.
Install addyosmani/agent-skills and count its skills correctly
Section titled “Install addyosmani/agent-skills and count its skills correctly”/plugin marketplace add addyosmani/agent-skills/plugin install agent-skills@addy-agent-skillsThe plugin’s /review and /ship collide with bare names you may already use. Type the namespaced form, /agent-skills:review, to be sure which reviewer runs.
codex plugin marketplace add addyosmani/agent-skillscodex plugin add agent-skills@agent-skillsThe marketplace registers as agent-skills in Codex (verified with codex plugin list, 0.157.1), not addy-agent-skills as in Claude Code. The README invokes skills with @, for example @spec-driven-development.
npx skills add addyosmani/agent-skills -a cursorThe README recommends workflow skills under .cursor/skills/ and only short policies in .cursor/rules/*.mdc, never whole skills pasted into rules.
The README says “25 skills” (24 lifecycle skills plus the using-agent-skills meta-skill); claude plugin details reports 34 skills, because the plugin also ships the nine lifecycle commands as skills. Quote the number for the route you describe. One more trap: older npx skills add installs copied only the skills/ folder, so skills that cite references/ files failed. If a skill cites a missing references/ file, reinstall from the current repository or use the plugin route.
Install Ponytail and choose its level
Section titled “Install Ponytail and choose its level”Ponytail makes the agent stop at the first rung of a ladder that holds before it writes code:
- Does this need to exist? If not, skip it.
- Is it already in this codebase? Reuse it.
- Does the standard library do it?
- Does a native platform feature do it (
<input type="date">, a database constraint, CSS)? - Does an already-installed dependency do it?
- Can it be one line?
- Only then: the minimum that works.
The README is explicit that trust-boundary validation, data-loss handling, security and accessibility are never cut, and the skill tells the agent to build anything explicitly requested without re-arguing. That rule is the hinge of the pairing below: whatever your spec requires counts as requested.
Send these as two separate prompts:
/plugin marketplace add DietrichGebert/ponytail/plugin install ponytail@ponytailSwitch level with /ponytail lite, /ponytail full (default), /ponytail ultra or /ponytail off. Review a diff with /ponytail-review; the fully namespaced name is /ponytail:ponytail-review.
codex plugin marketplace add DietrichGebert/ponytailcodex plugin add ponytail@ponytailThen run codex, open /hooks, review and trust Ponytail’s lifecycle hooks, and start a new thread. In Codex the commands are skills: @ponytail-review.
git clone https://github.com/DietrichGebert/ponytailnode ponytail/scripts/cursor-hooks.js installThis merges two hooks into ~/.cursor/hooks.json (--project writes the project file instead). Switch level by sending /ponytail ultra as a plain message. The hooks inject the ruleset and handle level switching, but ship no review command; for the review, we also install the skill with npx skills add DietrichGebert/ponytail --skill ponytail-review -a cursor. This route is our recommendation, not the README’s; it works because the skill folder exists at skills/ponytail-review. Cursor subagents run without the ruleset, and cloud agents never fire sessionStart.
| Level | What the agent does |
|---|---|
lite | Builds what you asked and names the lazier alternative in one line; you choose |
full (default) | Enforces the ladder: standard library and native features first, shortest diff |
ultra | Deletes before it adds, ships the one-liner and challenges the rest of the requirement |
off | Disabled for the session |
Set the default for every session with PONYTAIL_DEFAULT_MODE or a defaultMode field in ~/.config/ponytail/config.json. The hooks re-inject the ruleset on every prompt and into subagents; PONYTAIL_SUBAGENT_MATCHER (a regex against the subagent type) limits which subagents get it.
Measured cost. On Claude Code 2.1.283, claude plugin details ponytail@ponytail reports ~983 always-on tokens for the whole plugin (6 skills, 3 hooks). A skill’s full body loads only when that skill fires. The three hooks (SessionStart, SubagentStart, UserPromptSubmit) are harness-side and cost no model context, but they run node, so node must be on the non-interactive shell’s PATH. Nix and nvm users hit this first: the skills still work, and the always-on activation silently stays off.
Keep ECC only if you prune it
Section titled “Keep ECC only if you prune it”ECC is the largest bundle measured: 386 skills, 68 agents and 7 hooks, projected at about 41,515 always-on tokens in Claude Code. That is roughly 26 times Pocock’s set and 42 times Ponytail. Its README warns that re-uploads “may contain malware” and lists the only official names: the repository affaan-m/ECC, npm ecc-universal and ecc-agentshield, and the plugin ecc@ecc. The README’s npx ecc-universal@2.2.2 setup returns E404 because npm’s latest is 2.2.1 (checked 2026-09-26). If you still want it, install from the official sources only. In Claude Code:
/plugin marketplace add https://github.com/affaan-m/ECC/plugin install ecc@eccFor other agents, run the guided setup with the version npm actually has: npx ecc-universal@2.2.1 setup. Do not stop at the install: the ECC page has a keep-and-prune table that cuts it to ten components.
Example: run /ponytail-review on an over-built diff
Section titled “Example: run /ponytail-review on an over-built diff”The project is an invoicing app on Next.js with TypeScript. The spec for the invoice-list date filter says: filter by an inclusive from/to date range, reject to before from on the server, keep the range in the URL. The agent’s pull request passes its tests and adds:
react-datepickeranddate-fnstopackage.json;- a
DateRangeProviderinterface with a singleDefaultDateRangeProviderimplementation; - a 30-line
formatIsoDatehelper; - a
filters.enableTimezonesconfig flag that defaults tofalseand is read nowhere else.
On the feature branch, run the review (Claude Code: /ponytail-review; Codex: @ponytail-review). The skill returns one line per finding, tagged delete:, stdlib:, native:, yagni: or shrink:, and ends with a net line count. For this diff, a run in that format looks like this (illustrative, not a recorded transcript):
package.json:L41-42: native: react-datepicker + date-fns for two date fields. <input type="date"> twice, 0 deps.src/lib/date-range.ts:L1-24: yagni: DateRangeProvider with one implementation. Inline the two values into the query.src/lib/date-range.ts:L26-55: native: formatIsoDate re-creates what <input type="date"> already returns (YYYY-MM-DD). Delete it.src/config/filters.ts:L3-9: delete: enableTimezones flag nobody sets. Nothing replaces it.net: -118 lines possible.Three things about this output decide what you do next:
- It lists; it does not fix.
/ponytail-reviewapplies nothing. You or the agent apply the cuts in a separate step. - It is blind to correctness by design. The skill puts correctness bugs, security holes and performance explicitly out of scope. The server-side check that rejects
tobeforefromis a trust-boundary validation Ponytail will not flag and must not cut; a correctness review has to confirm it exists. - Your tests are the judge. Every cut must leave the spec’s tests green. If a cut breaks one, the cut was wrong, not the test.
Workflow: pair one “what to build” pack with Ponytail
Section titled “Workflow: pair one “what to build” pack with Ponytail”The pairing splits the two questions a review keeps asking. The process pack answers “is this the right thing, and is it proven?”. Ponytail answers “is this the least code that does it?”. The tabs show where the commands differ; the steps are the same in every tool.
-
Pick exactly one process pack. Pocock’s skills if misalignment is your main failure, gstack if you want roles and browser QA in Claude Code, agent-skills if you want one command per phase, Superpowers if you want test-driven discipline that fires on its own. Uninstall a second one before you start; do not run two.
-
Install Ponytail at
fulland measure the session. Run/contextin a fresh Claude Code session before and after both installs, andclaude plugin details <plugin>for each plugin (for gstack, which is not a plugin, run its owngstack-context-bill). Stop if the total surprises you: the cost is paid on every request. -
Align before the ladder runs. Run the pack’s interview (
grill-with-docs,/office-hoursor/spec). Ponytail’s first rung, “does this need to exist?”, is only safe once scope is written down, because anything the spec asks for counts as explicitly requested. -
Write acceptance criteria as tests. Use the pack’s spec and ticket step (
to-specandto-tickets, or/plan), and make every criterion name a test. See acceptance criteria that tests can check. -
Implement test-first with Ponytail active. The pack drives TDD; Ponytail shapes each change. Its own minimum is one runnable check for non-trivial logic, and it never flags that check for deletion. The spec’s tests are requested work, so it keeps them too.
-
Review in two passes. First correctness, then size:
/code-review high/ponytail-reviewTerminal window codex review --base mainThen, in a Codex session on the branch:
@ponytail-review.Run your usual correctness review, then ask Agent chat: “Use the ponytail-review skill on the diff against main.”
-
Gate on evidence, not on reading. CI runs tests, type check and lint on the trimmed branch. A human signs off twice: on the spec before implementation, and on the review findings before merge. The pull request carries the test results, both review outputs and the net line change, as the evidence bundle page describes.
How do you prove the pairing pays off on your repository?
Section titled “How do you prove the pairing pays off on your repository?”Do not adopt a pack on its star count or its vendor benchmark. Ponytail’s README reports a mean of 54% fewer lines of code across 12 feature tasks on Claude Haiku 4.5 (n=4, benchmark write-up dated 2026-06-18). That is vendor-reported, on one open-source repository, and near zero where the code was already minimal. Your repository is the only benchmark that decides.
Run a two-week pilot on five to ten real tickets with the pair installed, and compare against the tickets before it:
- Tests stay the oracle. Every merged ticket has its acceptance tests green in CI. A pilot ticket that needed a test deleted to pass counts as a failure.
- Diff size and new dependencies. Record
git diff --statand the number of packages added per ticket. - Review load. Count correctness findings per pull request; a pack that shrinks diffs but raises correctness findings is not helping.
- Context cost. Record
/contextat session start; the pair should stay within a few thousand tokens.
What breaks when you stack discipline packs?
Section titled “What breaks when you stack discipline packs?”| Symptom | Cause | Recovery |
|---|---|---|
| Ponytail never activates, no error | The hooks run node, which is missing from the non-interactive PATH (nvm, Nix) | Make node visible to non-interactive shells (for zsh, export PATH in ~/.zshenv; or symlink node into /usr/local/bin), restart, and check that the startup message shows the current level |
| Ponytail skips a feature the ticket asked for | Scope lived in chat, not in the spec, so rung 1 treated it as speculative | Write it into the spec’s acceptance criteria, or say “explicitly requested” in the prompt; use lite while scope is still moving |
| The agent drops a test to make the diff smaller | Ponytail’s one-check minimum read as a ceiling | Add the precedence rule from the prompt above to AGENTS.md; tests named in the spec are required |
| Two planning flows start at once | Two process packs installed, both auto-triggering | Keep one; claude plugin disable <plugin> the other and confirm with /context |
| Session context jumps by tens of thousands of tokens | ECC installed whole | Prune to the parts you use, or uninstall it |
| Codex ignores Ponytail after install | Its hooks were never trusted | Open /hooks in Codex, trust them, start a new thread |
| Uninstall leaves a status line or mode flag behind | Ponytail writes state outside the plugin folder | Run node scripts/uninstall.js from a clone before /plugin remove ponytail |