Skip to content

Discipline packs: Pocock skills, gstack, agent-skills, Ponytail and ECC

Discipline packs are skill bundles that change how a coding agent works instead of giving it new tools. Matt Pocock’s skills, gstack, addyosmani/agent-skills and Everything Claude Code (ECC) decide what to build and in which order. Ponytail decides how much code to write. The pairing that holds up is one of the first four plus Ponytail, never two process packs.

You asked the agent for a date filter on the invoice list. It interviewed you, wrote a spec, made tickets and wrote tests first, exactly as the pack promised. The pull request still adds two npm packages, an interface with one implementation, and a config flag nobody will ever set. The process pack made sure the agent built the right thing; nothing in it asked whether the agent built too much of it.

This page is for the developer choosing a pack and the tech lead standardising one for a team. It compares the five most-starred packs on what they cost and where they collide, then shows the pairing that covers both questions.

What you get from pairing a process pack with Ponytail

Section titled “What you get from pairing a process pack with Ponytail”
  • A comparison of five discipline packs on the question each answers, install route, always-on context cost and collision risk.
  • Verified install commands for Claude Code, Codex and Cursor, including the two packs that do not install the way their names suggest.
  • A worked /ponytail-review run on an over-built diff, and what to do with its delete list.
  • A seven-step workflow that pairs one “what to build” pack with Ponytail, with tests and a correctness review as the gates.
  • Three copy-paste prompts, and a table of the failures that stacking packs produces.
PackQuestion it answersEntry pointsAlways-on cost (Claude Code)Popularity (GitHub, 2026-09-26)Pick it whenAvoid it when
Matt Pocock’s skills (mattpocock/skills)Do you and the agent agree on what to build?grill-with-docs → to-spec → to-tickets → implement → code-review~1,609 tokens (25 skills)269.8k starsMisalignment is your main failure; you want small, editable skillsYou need a durable spec trail approved by product owners
gstack (garrytan/gstack)Which role should look at this next?/office-hours, /plan-eng-review, /review, /qa, /ship, /retroNot a plugin; measure with its own gstack-context-bill134.2k starsYou want product, review, QA and release roles with built-in browser QA, mostly in Claude CodeYou cannot accept hooks in ~/.claude/settings.json
addyosmani/agent-skillsWhich lifecycle phase are you in?/spec, /plan, /build, /test, /review, /ship and three more~3,620 tokens (34 skills, 4 agents)99.1k starsYou want one command per phase and skills that trigger on the kind of workYou already run another pack’s /review or /ship
Everything Claude Code (affaan-m/ECC)Can one harness cover everything?386 skills, 68 agents, 7 hooks~41,515 tokens267.6k starsYou will prune it to a handful of partsContext-sensitive work, or you will not prune
Ponytail (DietrichGebert/ponytail)How much code does this need?/ponytail [lite|full|ultra|off], /ponytail-review, /ponytail-audit~983 tokens (6 skills, 3 hooks)146.1k starsAgents over-build: extra dependencies, one-implementation abstractionsYou want it to decide scope; it only decides size

Two readings of this table matter more than the rows. First, the first four packs overlap: each has its own idea of planning and review, so two of them give you two /review commands and two opinions about process. Second, Ponytail overlaps with none of them. It has no planning stage and no documents; it constrains how the agent writes whatever the other pack decided on.

Superpowers (~838 always-on tokens) belongs in the same “what to build” family and has its own page; everything below about pairing applies to it too.

Install Matt Pocock’s skills as a framework

Section titled “Install Matt Pocock’s skills as a framework”

Pocock’s skills are small and composable, and several of them are one-line delegates: grill-me only calls grilling, and grill-with-docs calls grilling and domain-modeling. As a framework, the chain is: grill until you agree, turn the conversation into a spec, split it into tickets, implement each ticket test-first, and review against the spec. The Grill Me page covers the interview step in depth.

Terminal window
claude plugin install mattpocock-skills

The plugin is in the official marketplace, so nothing needs adding first. Plugin skills are namespaced: type /mattpocock-skills:grill-with-docs, /mattpocock-skills:to-spec and so on.

Then run setup-matt-pocock-skills once per repository. It asks which issue tracker to-spec and to-tickets should write to (GitHub, Linear or local files), which triage labels you use, and where docs go. Install through the plugin or through npx skills, not both, or every skill appears twice.

Install gstack without the /review collision

Section titled “Install gstack without the /review collision”

gstack is a Git repository with a setup script, not a plugin marketplace: /plugin marketplace add garrytan/gstack installs nothing. It needs Git and Bun.

Terminal window
git clone --single-branch --depth 1 https://github.com/garrytan/gstack.git ~/.claude/skills/gstack
cd ~/.claude/skills/gstack && ./setup --prefix

--prefix names every skill /gstack-<name>. Without it, gstack’s /review claims the name that Claude Code 2.1.283 already uses as the alias of its bundled /code-review. For Codex and Cursor, clone to ~/gstack and run ./setup --host codex or ./setup --host cursor. Setup also registers a default-on Stop hook in ~/.claude/settings.json; ./setup --no-timeline-stop-hook skips it. The gstack page runs one feature through every role.

Install addyosmani/agent-skills and count its skills correctly

Section titled “Install addyosmani/agent-skills and count its skills correctly”
/plugin marketplace add addyosmani/agent-skills
/plugin install agent-skills@addy-agent-skills

The plugin’s /review and /ship collide with bare names you may already use. Type the namespaced form, /agent-skills:review, to be sure which reviewer runs.

The README says “25 skills” (24 lifecycle skills plus the using-agent-skills meta-skill); claude plugin details reports 34 skills, because the plugin also ships the nine lifecycle commands as skills. Quote the number for the route you describe. One more trap: older npx skills add installs copied only the skills/ folder, so skills that cite references/ files failed. If a skill cites a missing references/ file, reinstall from the current repository or use the plugin route.

Ponytail makes the agent stop at the first rung of a ladder that holds before it writes code:

  1. Does this need to exist? If not, skip it.
  2. Is it already in this codebase? Reuse it.
  3. Does the standard library do it?
  4. Does a native platform feature do it (<input type="date">, a database constraint, CSS)?
  5. Does an already-installed dependency do it?
  6. Can it be one line?
  7. Only then: the minimum that works.

The README is explicit that trust-boundary validation, data-loss handling, security and accessibility are never cut, and the skill tells the agent to build anything explicitly requested without re-arguing. That rule is the hinge of the pairing below: whatever your spec requires counts as requested.

Send these as two separate prompts:

/plugin marketplace add DietrichGebert/ponytail
/plugin install ponytail@ponytail

Switch level with /ponytail lite, /ponytail full (default), /ponytail ultra or /ponytail off. Review a diff with /ponytail-review; the fully namespaced name is /ponytail:ponytail-review.

LevelWhat the agent does
liteBuilds what you asked and names the lazier alternative in one line; you choose
full (default)Enforces the ladder: standard library and native features first, shortest diff
ultraDeletes before it adds, ships the one-liner and challenges the rest of the requirement
offDisabled for the session

Set the default for every session with PONYTAIL_DEFAULT_MODE or a defaultMode field in ~/.config/ponytail/config.json. The hooks re-inject the ruleset on every prompt and into subagents; PONYTAIL_SUBAGENT_MATCHER (a regex against the subagent type) limits which subagents get it.

Measured cost. On Claude Code 2.1.283, claude plugin details ponytail@ponytail reports ~983 always-on tokens for the whole plugin (6 skills, 3 hooks). A skill’s full body loads only when that skill fires. The three hooks (SessionStart, SubagentStart, UserPromptSubmit) are harness-side and cost no model context, but they run node, so node must be on the non-interactive shell’s PATH. Nix and nvm users hit this first: the skills still work, and the always-on activation silently stays off.

ECC is the largest bundle measured: 386 skills, 68 agents and 7 hooks, projected at about 41,515 always-on tokens in Claude Code. That is roughly 26 times Pocock’s set and 42 times Ponytail. Its README warns that re-uploads “may contain malware” and lists the only official names: the repository affaan-m/ECC, npm ecc-universal and ecc-agentshield, and the plugin ecc@ecc. The README’s npx ecc-universal@2.2.2 setup returns E404 because npm’s latest is 2.2.1 (checked 2026-09-26). If you still want it, install from the official sources only. In Claude Code:

/plugin marketplace add https://github.com/affaan-m/ECC
/plugin install ecc@ecc

For other agents, run the guided setup with the version npm actually has: npx ecc-universal@2.2.1 setup. Do not stop at the install: the ECC page has a keep-and-prune table that cuts it to ten components.

Example: run /ponytail-review on an over-built diff

Section titled “Example: run /ponytail-review on an over-built diff”

The project is an invoicing app on Next.js with TypeScript. The spec for the invoice-list date filter says: filter by an inclusive from/to date range, reject to before from on the server, keep the range in the URL. The agent’s pull request passes its tests and adds:

  • react-datepicker and date-fns to package.json;
  • a DateRangeProvider interface with a single DefaultDateRangeProvider implementation;
  • a 30-line formatIsoDate helper;
  • a filters.enableTimezones config flag that defaults to false and is read nowhere else.

On the feature branch, run the review (Claude Code: /ponytail-review; Codex: @ponytail-review). The skill returns one line per finding, tagged delete:, stdlib:, native:, yagni: or shrink:, and ends with a net line count. For this diff, a run in that format looks like this (illustrative, not a recorded transcript):

package.json:L41-42: native: react-datepicker + date-fns for two date fields. <input type="date"> twice, 0 deps.
src/lib/date-range.ts:L1-24: yagni: DateRangeProvider with one implementation. Inline the two values into the query.
src/lib/date-range.ts:L26-55: native: formatIsoDate re-creates what <input type="date"> already returns (YYYY-MM-DD). Delete it.
src/config/filters.ts:L3-9: delete: enableTimezones flag nobody sets. Nothing replaces it.
net: -118 lines possible.

Three things about this output decide what you do next:

  1. It lists; it does not fix. /ponytail-review applies nothing. You or the agent apply the cuts in a separate step.
  2. It is blind to correctness by design. The skill puts correctness bugs, security holes and performance explicitly out of scope. The server-side check that rejects to before from is a trust-boundary validation Ponytail will not flag and must not cut; a correctness review has to confirm it exists.
  3. Your tests are the judge. Every cut must leave the spec’s tests green. If a cut breaks one, the cut was wrong, not the test.

Workflow: pair one “what to build” pack with Ponytail

Section titled “Workflow: pair one “what to build” pack with Ponytail”

The pairing splits the two questions a review keeps asking. The process pack answers “is this the right thing, and is it proven?”. Ponytail answers “is this the least code that does it?”. The tabs show where the commands differ; the steps are the same in every tool.

  1. Pick exactly one process pack. Pocock’s skills if misalignment is your main failure, gstack if you want roles and browser QA in Claude Code, agent-skills if you want one command per phase, Superpowers if you want test-driven discipline that fires on its own. Uninstall a second one before you start; do not run two.

  2. Install Ponytail at full and measure the session. Run /context in a fresh Claude Code session before and after both installs, and claude plugin details <plugin> for each plugin (for gstack, which is not a plugin, run its own gstack-context-bill). Stop if the total surprises you: the cost is paid on every request.

  3. Align before the ladder runs. Run the pack’s interview (grill-with-docs, /office-hours or /spec). Ponytail’s first rung, “does this need to exist?”, is only safe once scope is written down, because anything the spec asks for counts as explicitly requested.

  4. Write acceptance criteria as tests. Use the pack’s spec and ticket step (to-spec and to-tickets, or /plan), and make every criterion name a test. See acceptance criteria that tests can check.

  5. Implement test-first with Ponytail active. The pack drives TDD; Ponytail shapes each change. Its own minimum is one runnable check for non-trivial logic, and it never flags that check for deletion. The spec’s tests are requested work, so it keeps them too.

  6. Review in two passes. First correctness, then size:

    /code-review high
    /ponytail-review
  7. Gate on evidence, not on reading. CI runs tests, type check and lint on the trimmed branch. A human signs off twice: on the spec before implementation, and on the review findings before merge. The pull request carries the test results, both review outputs and the net line change, as the evidence bundle page describes.

How do you prove the pairing pays off on your repository?

Section titled “How do you prove the pairing pays off on your repository?”

Do not adopt a pack on its star count or its vendor benchmark. Ponytail’s README reports a mean of 54% fewer lines of code across 12 feature tasks on Claude Haiku 4.5 (n=4, benchmark write-up dated 2026-06-18). That is vendor-reported, on one open-source repository, and near zero where the code was already minimal. Your repository is the only benchmark that decides.

Run a two-week pilot on five to ten real tickets with the pair installed, and compare against the tickets before it:

  • Tests stay the oracle. Every merged ticket has its acceptance tests green in CI. A pilot ticket that needed a test deleted to pass counts as a failure.
  • Diff size and new dependencies. Record git diff --stat and the number of packages added per ticket.
  • Review load. Count correctness findings per pull request; a pack that shrinks diffs but raises correctness findings is not helping.
  • Context cost. Record /context at session start; the pair should stay within a few thousand tokens.

What breaks when you stack discipline packs?

Section titled “What breaks when you stack discipline packs?”
SymptomCauseRecovery
Ponytail never activates, no errorThe hooks run node, which is missing from the non-interactive PATH (nvm, Nix)Make node visible to non-interactive shells (for zsh, export PATH in ~/.zshenv; or symlink node into /usr/local/bin), restart, and check that the startup message shows the current level
Ponytail skips a feature the ticket asked forScope lived in chat, not in the spec, so rung 1 treated it as speculativeWrite it into the spec’s acceptance criteria, or say “explicitly requested” in the prompt; use lite while scope is still moving
The agent drops a test to make the diff smallerPonytail’s one-check minimum read as a ceilingAdd the precedence rule from the prompt above to AGENTS.md; tests named in the spec are required
Two planning flows start at onceTwo process packs installed, both auto-triggeringKeep one; claude plugin disable <plugin> the other and confirm with /context
Session context jumps by tens of thousands of tokensECC installed wholePrune to the parts you use, or uninstall it
Codex ignores Ponytail after installIts hooks were never trustedOpen /hooks in Codex, trust them, start a new thread
Uninstall leaves a status line or mode flag behindPonytail writes state outside the plugin folderRun node scripts/uninstall.js from a clone before /plugin remove ponytail