gstack: a role-based skill stack for Claude Code
gstack is Garry Tan’s MIT-licensed set of Claude Code skills that splits feature work into roles invoked as slash commands: /office-hours and /plan-ceo-review for product, /plan-eng-review for the engineering manager, /review for a staff engineer, /qa for QA, and /ship for release. The same skills also install into Codex and Cursor. Its planning documents live outside your repository by default.
You open Claude Code with a one-line ticket, “Add CSV export to invoices”, and get back a working button in 10 minutes. Then the questions start. Which date range does the export use? What happens with 80,000 rows? Did anyone check the timezone of issued_at? Nobody asked, because a blank prompt plays no roles: no product owner, no tech lead, no QA.
This page is for the developer who wants those roles played every time without hiring them or reading every generated line. It runs one feature through gstack from idea to pull request, shows the file each role leaves behind, and marks where you still have to step in.
What running one feature through gstack gives you
Section titled “What running one feature through gstack gives you”- A working install for Claude Code, Codex or Cursor, and the one setup flag that avoids a command-name collision.
- One feature, CSV export for an invoicing app, taken from idea to an open pull request in seven commands.
- A table that maps each gstack role onto the artifact chain, including the two links gstack does not produce at all.
- Four copy-paste prompts: a scoped
/office-hoursopener, a constrained engineering review, promoting gstack’s private notes into committed files, and a pre-ship evidence check. - The gates that prove the result without you reading the diff, and the traps that old tutorials miss.
Which roles does gstack give you?
Section titled “Which roles does gstack give you?”gstack presents its skills as specialists who run in sprint order: Think, Plan, Build, Review, Test, Ship, Reflect. Each skill reads what the previous one wrote. The README lists more than 30 commands; these nine carry a normal feature:
| Role | Command | What it does | What it leaves behind |
|---|---|---|---|
| Product partner | /office-hours | Asks forcing questions about the problem, challenges your premises, and proposes two or three approaches with effort estimates | A design doc in ~/.gstack/projects/<repo>/ |
| Founder or CEO | /plan-ceo-review | Reviews scope in one of four modes: expansion, selective expansion, hold scope, or reduction | Scope decisions in ~/.gstack/projects/ |
| Engineering manager | /plan-eng-review | Locks architecture, data flow, failure modes, and the test matrix, with diagrams | A reviewed plan and a test plan in ~/.gstack/projects/ |
| Review pipeline | /autoplan | Runs the CEO, design, developer-experience, and engineering reviews in sequence, and asks you only about taste decisions | The same files, with fewer questions |
| Staff engineer | /review | Hunts for bugs that pass CI: N+1 queries, races, trust boundaries, missing enum branches. Fixes the obvious ones and asks about the rest | In-place fixes (listed as [AUTO-FIXED] lines) and a review record |
| QA lead | /qa | Reads the diff, opens the affected pages in a real browser, fixes what breaks, and writes a regression test per fix | A report in .gstack/qa-reports/ and new tests |
| Release engineer | /ship | Syncs the main branch, runs tests, audits coverage, updates docs, pushes, and opens the PR | The pull request with a coverage summary |
| Debugger | /investigate | Root-cause debugging that stops after three failed fixes | Findings in the session |
| Engineering manager | /retro | Weekly retrospective from commit history, with test-ratio trends | A JSON snapshot in .context/retros/ |
Only one of these is a hard gate by default: the engineering review. gstack’s Review Readiness Dashboard marks it Required; the CEO and design reviews are informational.
Install gstack for Claude Code, Codex or Cursor
Section titled “Install gstack for Claude Code, Codex or Cursor”gstack needs Git and Bun 1.0 or later, plus Node.js on Windows. It is a Git repository with a setup script, not a plugin marketplace, so /plugin marketplace add garrytan/gstack installs nothing.
-
Clone and run setup for your agent (terminal):
Terminal window git clone --single-branch --depth 1 https://github.com/garrytan/gstack.git ~/.claude/skills/gstackcd ~/.claude/skills/gstack && ./setup --prefix--prefixinstalls every skill as/gstack-<name>(/gstack-review,/gstack-qa). The default is short names (/review,/qa), and on Claude Code 2.1.283/reviewis already the alias of the bundled/code-review. With the prefix you always know which reviewer you invoked. Switch back later with./setup --no-prefix. The rest of this page uses the short names for readability; addgstack-if you installed with the prefix.Restart Claude Code. The skills are user-level, so they appear in every project.
Terminal window git clone --single-branch --depth 1 https://github.com/garrytan/gstack.git ~/gstackcd ~/gstack && ./setup --host codexSetup generates Codex-format skills into
${CODEX_HOME:-~/.codex}/skills/gstack-*/. Restart Codex, then mention a skill with$, for example$gstack-office-hours, or pick it from/skills. The/codexsecond-opinion skill does not exist on Codex; the counterpart isgstack-claude-code, which needs Claude Code installed and signed in. If Codex prints “Skipped loading skill(s) due to invalid SKILL.md”, the generated descriptions are stale: runcd ~/gstack && git pull && ./setup --host codex.Terminal window git clone --single-branch --depth 1 https://github.com/garrytan/gstack.git ~/gstackcd ~/gstack && ./setup --host cursorSetup writes the skills to
~/.cursor/skills/gstack-*/. How Cursor lists and triggers them was not re-checked for this page, because cursor.com was unreachable on 2026-09-26; ask the agent for a skill by name (“use the gstack-qa skill”) if it does not appear in the menu. -
Decide on the hooks before your first session. Setup writes hooks into
~/.claude/settings.json: a default-onStophook that closes session-timeline entries,AskUserQuestionhooks only if you accept them when setup asks, and aSessionStartauto-update check in team mode. Hooks run outside the permission prompt. To skip the timeline hook, rerun with./setup --no-timeline-stop-hook; to see what is registered, run~/.claude/skills/gstack/bin/gstack-settings-hook list-sources. On first use gstack also asks about telemetry, which stays off unless you say yes. -
Measure what gstack adds to every session. The skill descriptions load at startup; the skill bodies load when you call them. gstack ships an offline estimator (terminal):
Terminal window ~/.claude/skills/gstack/bin/gstack-context-bill ~/.claude/skillsIt prints the always-on cost and the per-invocation cost of each skill. Compare
/contextin a fresh session before and after install. On 1.91.2.0 theoffice-hours,reviewandshipskill files are each 70–85 KB, so a single invocation is a large prompt. For scale: run against the 1.91.2.0 source tree on 2026-09-26, the estimator reported 63 skills with about 25.7 KB of frontmatter, roughly 6.6k tokens always on (offline estimate;--exactneeds an Anthropic API key). Your installed figure depends on which hosts and skills setup linked.claude plugin detailsdoes not apply, because gstack installs as skills, not as a plugin.
For the general mechanics of skill folders and scopes, see installing and managing skills.
Run one feature through gstack, role by role
Section titled “Run one feature through gstack, role by role”The project is an invoicing app on Next.js App Router with Postgres through Prisma, unit tests in Vitest and browser tests in Playwright. The ticket: finance wants a CSV export of invoices for a date range. Work on a feature branch; /qa and /review compare against the main branch.
-
Think:
/office-hours. Start with the problem, not the button./office-hourshas two modes. Startup mode, meant for founders and intrapreneurs, asks six forcing questions about demand, the status quo, and the narrowest wedge. Builder mode, meant for side projects and hackathons, is generative rather than interrogative. For a feature that paying customers asked for, ask for startup mode: its questions are the ones that test whether “a CSV export” is the real need. It pushes back on your framing, lists premises for you to accept or reject, and ends with a design doc that the later skills read.Accept the design doc only when it names the users, the outcome, and what is out of scope. You are the product owner here: this is the one step gstack cannot sign off for you.
-
Plan:
/plan-eng-review. Switch Claude Code to plan mode (Shift+Tab, or/plan), let it draft the plan, then run the engineering review. In plan mode the skill reviews the active plan automatically and says so in one line. It forces the decisions a blank prompt skips: streaming against buffering, the timezone boundary, and what a partial failure does. It writes a test plan that/qapicks up later. Use/autoplaninstead when you want the CEO, design, and engineering reviews in one pass and only the taste decisions surfaced. -
Build: exit plan mode and approve the plan. Claude Code implements it. gstack adds no build role of its own here; the plan is the contract.
-
Review:
/review(or/gstack-review). It applies mechanical fixes and lists each one as an[AUTO-FIXED]line, and asks about anything ambiguous, such as a race or a trust boundary. For a CSV export, the findings worth having are formula injection in cells that start with=,+,-or@, and a missing index on(account_id, issued_at). If the review reports neither, ask about both by name. For a second model,/codexsends the same diff to Codex CLI if you have it installed and signed in. -
Test:
/qa http://localhost:3000/invoices. On a feature branch,/qais diff-aware: it reads the branch diff, opens the affected pages in a browser (the Aside browser, if it is installed and open on macOS 15+, otherwise the Chromium that setup builds), exercises them, fixes what breaks, and writes a regression test for each fix. Use/qa-onlywhen you want the report without code changes. -
Ship:
/ship. It syncs the main branch, runs your tests, builds a coverage map of the diff, writes tests for gaps, updates docs through/document-release, pushes, and opens the PR with a line such asTests: 42 → 47 (+5 new). Before it opens the PR, it checks the Review Readiness Dashboard. If the engineering review is missing it asks, but it does not block you. -
Reflect:
/retroat the end of the week. It reads commit history and flags a test ratio below 20%. Treat its numbers as signals for a conversation, not as a performance metric.
Deployment stays with your pipeline. gstack also has /land-and-deploy (merge, wait for CI and deploy, check production health) and /canary (post-deploy monitoring). If your team deploys only through CI, leave them out and let the PR merge trigger the pipeline.
Where does each gstack role land on the artifact chain?
Section titled “Where does each gstack role land on the artifact chain?”The artifact chain asks for a committed file at each stage, accepted by a named person. gstack produces most of the content, but it stores the planning half in ~/.gstack/projects/ on your machine, not in the repository. Teammates, reviewers, and CI never see it unless you move it.
| Chain stage | Artifact | gstack role that produces it | Where gstack puts it | What you add |
|---|---|---|---|---|
| Plan | intent.md | /office-hours, /plan-ceo-review | Design doc in ~/.gstack/projects/ | Commit it; the product owner accepts it |
| Design | spec.md | /plan-eng-review (acceptance criteria, test matrix) | Test plan in ~/.gstack/projects/ | Commit it; turn every criterion into a named test |
| Build | plan.md and the diff | Claude Code’s plan, reviewed by /plan-eng-review | Plan file, then commits | Commit the approved plan beside the change |
| Test | Test evidence | /qa, plus the coverage audit in /ship | .gstack/qa-reports/, new test files | CI must run the new tests; attach the report to the PR |
| Review | Review findings | /review, /codex | In-place fixes (listed as [AUTO-FIXED] lines) and a review record | A code owner reads the findings, not the whole diff |
| Maintain | Incident record, next intent | /investigate, /retro, /canary | Session output, .context/retros/ | Write the incident record yourself |
The two links gstack does not produce at all are the committed intent/spec in the repository and the incident record; everything else needs only promoting.
How do you prove the gstack output without reading every line?
Section titled “How do you prove the gstack output without reading every line?”gstack’s claim is that the roles catch what you would otherwise catch by reading. Make each claim checkable:
-
Acceptance criteria as tests. After promoting the test plan, every criterion names a test.
/shipadds tests for coverage gaps, and/qaadds a regression test per bug it fixes. CI, not the agent, runs them. See acceptance criteria that tests can check. -
A gate that stops the turn. gstack’s opt-in
gstack-verify-gateis aStophook that blocks Claude Code from ending a turn until your verify command passes. Declare the command inCLAUDE.md, trust it once per repository, then register the hook (terminal, from the repository root):Terminal window echo '<!-- gstack:verify: npm test -->' >> CLAUDE.md~/.claude/skills/gstack/bin/gstack-verify-gate --trust~/.claude/skills/gstack/bin/gstack-settings-hook add-event --event Stop \--command ~/.claude/skills/gstack/bin/gstack-verify-gate --source verify-gateEditing the declared command cancels the trust until you run
--trustagain. After three blocked attempts it lets the turn end with a warning, so it does not loop forever. -
Fresh reviews only. gstack grades each diff review CURRENT, STALE, or UNVERIFIED against a fingerprint of the working tree. A review of code that changed afterwards does not count, and
/shipreads the same grade. -
Test evidence bound to the code it ran against.
gstack-evidence run --label unit -- npx vitest runwraps your test command, passes its exit code through, and records which working-tree fingerprint it ran on. Itschecksubcommand grades each label FRESH, STALE, or MISSING, and/shipcites fresh evidence instead of rerunning the suite. A green run on yesterday’s code does not count. -
Two models, one diff.
/reviewruns on Claude and/codexon Codex, so they can miss different things. Read a finding that both report first. -
A named sign-off. You accept the design doc and the reviewed plan. A code owner merges on green CI with the review findings resolved. The evidence bundle page describes what that PR should carry.
What breaks when you adopt gstack?
Section titled “What breaks when you adopt gstack?”| Symptom | Cause | Recovery |
|---|---|---|
| Skills do not appear in Claude Code | Setup did not finish, or the session predates it | cd ~/.claude/skills/gstack && ./setup, then restart. If they still do not appear, check that ~/.claude/skills/gstack-* or ~/.claude/skills/<name> links exist (ls ~/.claude/skills) and rerun ./setup. If the links exist and Claude still says it cannot see the skills, add the README’s ## gstack section to the project’s CLAUDE.md: it lists the skills so Claude routes to them, it does not register them |
| Setup ends with “Not registered (a skill you own already uses the name)” | You have your own qa/ or ship/ skill | Rename yours, or switch to ./setup --prefix so the names stop colliding |
/qa or /browse fails on Linux or Windows | The bundled Chromium did not install or cannot launch | cd ~/.claude/skills/gstack && bun install && bun run build. On Ubuntu 24.04+ with blocked user namespaces, set GSTACK_CHROMIUM_NO_SANDBOX=1 |
| A teammate’s session behaves differently after lunch | Team mode auto-updates gstack at session start, at most once per hour | For a shared repository, agree on who upgrades and when, read the CHANGELOG first, and run /gstack-upgrade deliberately |
/office-hours turns a small feature into a product pitch | A vague opener leaves it free to hunt for a bigger product, and builder mode actively looks for the “whoa” version | Name the users, the constraint, and what is out of scope, as in the prompt above; then run /plan-ceo-review in hold-scope or reduction mode |
When is gstack the right framework?
Section titled “When is gstack the right framework?”Pick gstack when you want opinionated roles with a lot of built-in browser QA and release automation, and you mostly work in Claude Code. Prefer Superpowers when you want a smaller, test-driven discipline that fires on its own, and GSD when your problem is context loss across a long multi-phase build. Two role packs at once give you two /review commands and two opinions about process, so pick one. The frameworks comparison sets them side by side.