Skip to content

GSD: context-engineered phases for solo builders

GSD Core, formerly Get Shit Done, is an open-source framework that drives Claude Code, Codex or Cursor through a discuss, plan, execute, verify and ship loop for each roadmap phase. Research, planning and coding run in fresh subagents while state lives in .planning/. As a Claude Code plugin it adds about 10,700 always-on tokens per session.

You are building a side product alone. The first evening with the agent goes well; by the third, the session has compacted twice, the agent has forgotten why you chose SQLite, and a “finished” endpoint turns out to be a stub. You do not have a reviewer, and you do not want to read 4,000 lines of generated code to find out what works.

This page runs one real project through GSD, shows which files and gates prove each phase works, and says what the framework costs.

What you get from running GSD on one project

Section titled “What you get from running GSD on one project”
  • A working install for Claude Code, Codex or Cursor, and the command spelling each one uses.
  • One uptime monitor taken from a one-page brief to an open pull request, with the file each step writes, so you check a 40-line plan instead of a 400-line diff.
  • Three copy-paste prompts, the verification layers GSD runs, and a CI job that makes the pull request prove itself.
  • The token bill, and the five settings that change it.

GSD’s premise is that a long session degrades as its context fills (context rot), so the main session only orchestrates. Each heavy job goes to a subagent that starts with a clean context, writes its result to a file in .planning/, and exits. The next subagent reads the file, not the conversation. For the underlying mechanics, see how context windows fill and degrade.

Each roadmap phase goes through the same five steps:

StepCommand (Claude Code)What runsFile it writesYour job
Discuss/gsd-discuss-phase 1Adaptive questions about how to build the phase01-CONTEXT.md, 01-DISCUSSION-LOG.mdAnswer. Every answer becomes an input to the planner
Plan/gsd-plan-phase 1Optional researcher, then a planner, then a plan-checker loop01-RESEARCH.md, 01-01-PLAN.md…, 01-VALIDATION.mdRead the plans, not the code
Execute/gsd-execute-phase 1Executor subagents in parallel waves, one per plan, each committing its work; then a verifier01-01-SUMMARY.md…, 01-VERIFICATION.mdNothing, unless a checkpoint asks
Verify/gsd-verify-work 1A walkthrough of the user-visible deliverables, one checkpoint at a time01-UAT.md, fix plans if anything failsTry each behaviour; type pass or describe what is wrong
Ship/gsd-ship 1Preflight gates, push, gh pr create with a body built from the artifactsThe pull requestMerge when CI is green

In Codex type $gsd-discuss-phase 1 and so on; with the plugin, /gsd-core:discuss-phase 1.

After the first phase of the project in the next section, .planning/ looks like this:

  • Directory.planning/
    • PROJECT.md what you are building and why
    • REQUIREMENTS.md one ID per capability, for example CHK-01
    • ROADMAP.md phases, each with a goal and success criteria
    • STATE.md where you are; survives /clear and a closed laptop
    • config.json mode, model profile, gates
    • Directoryresearch/
      • …
    • Directoryphases/
      • Directory01-check-runner/
        • 01-CONTEXT.md your decisions from discuss
        • 01-RESEARCH.md
        • 01-01-PLAN.md a task with a verify command and a done condition
        • 01-02-PLAN.md
        • 01-01-SUMMARY.md what the executor built and committed
        • 01-02-SUMMARY.md
        • 01-VERIFICATION.md requirement coverage from the verifier
        • 01-UAT.md your walkthrough results
        • 01-SECURITY.md written by /gsd-secure-phase

The files are the point: small, committed by default (planning.commit_docs: true), and each answers a question you would otherwise answer by reading code.

Install GSD Core for Claude Code, Codex or Cursor

Section titled “Install GSD Core for Claude Code, Codex or Cursor”

The npm package is @opengsd/gsd-core. Versions 1.14.0 and 1.15.0 declare Node.js 24 or later and npm 10 or later in their engines field. Ship also needs the GitHub CLI (gh), authenticated.

  1. Run the installer from your project directory (terminal). With no flags it asks for the runtime and the scope:

    Terminal window
    npx @opengsd/gsd-core@latest

    Or name them directly:

    Terminal window
    npx @opengsd/gsd-core@latest --claude --local # this project only, in ./.claude/
    npx @opengsd/gsd-core@latest --claude --global # every project, in ~/.claude/

    Restart Claude Code. Commands appear as /gsd-new-project, /gsd-plan-phase and so on. The installer also changes your settings.json:

    • Hooks: an update check, PreToolUse guards (prompt injection, secret-file reads, destructive writes and more) and a context monitor.
    • Allow rules for Bash(npx gsd-core *) and for Read and Edit on .planning/*.
    • Removed deny rules: it deletes the exact rules Read(.env), Read(.env.*) and Read(.secrets), even ones you wrote, because its secret-read hook replaces them. Re-add them if other tools rely on them.

    The alternative is the native plugin:

    /plugin marketplace add open-gsd/gsd-core
    /plugin install gsd-core@gsd-core

    Plugin commands are namespaced /gsd-core:plan-phase, and the plugin still needs the gsd-tools binary from the npm package on your PATH. Use one path, not both.

  2. Check that the commands resolved. In the agent, run the help command (/gsd-help in Claude Code, $gsd-help in Codex, the gsd-help entry in Cursor’s / menu). If it is missing, see the recovery table at the end of this page.

  3. Decide your permission posture before the first execute. GSD’s allow rules cover its own files, but executor subagents also run git add, git commit and your tests. On Claude Code’s latest channel (v2.1.283), interactive sessions start in auto mode on supported models unless your settings disable it. In Manual mode, add allow rules such as Bash(git add *) and Bash(git commit *) to .claude/settings.local.json instead of the --dangerously-skip-permissions flag GSD’s user guide suggests. See permissions and sandboxing for agents.

Run one project through GSD, phase by phase

Section titled “Run one project through GSD, phase by phase”

The project is pingboard, a self-hosted uptime monitor: a Fastify API in TypeScript that checks URLs on a schedule, stores results in SQLite, and alerts by email when a check fails twice in a row. Three phases: the check runner and storage, the HTTP API, then alerting. This section takes phase 1 all the way to a pull request.

  1. Write the brief GSD will ingest. /gsd-new-project --auto @file.md extracts the project from a document instead of interviewing you. Have the agent draft it, then edit it yourself: every later subagent inherits it.

  2. Create the project (agent prompt):

    /gsd-new-project --auto @docs/prd.md

    After a few config questions, --auto runs research, requirements and the roadmap without stopping. Because nothing pauses for your approval, open ROADMAP.md before you go on: each phase needs a goal and success criteria you can observe. Nobody can verify “the scheduler is robust”; the verifier and you can both verify “a check that times out after 5 s is stored with status timeout”.

  3. Clear the session and discuss phase 1:

    /clear
    /gsd-discuss-phase 1

    GSD asks about implementation choices: interval granularity, what counts as a failure, how to store timings. Answer concretely; --assumptions lists what the agent would assume instead, a fast way to find the questions that matter. Answers land in 01-CONTEXT.md.

  4. Plan phase 1 with a failing test per behaviour:

    /gsd-plan-phase 1 --tdd

    A planner turns 01-CONTEXT.md into small plans, and a plan-checker loops with it until each plan meets the goal. --tdd makes each behaviour-adding task start with a failing test; --skip-research saves a researcher on a familiar stack.

  5. Review the plans before anything runs. Each PLAN.md holds tasks with the files they touch, action steps, a <verify> command and a done condition. This is the cheapest place to catch a wrong decision.

    On NO-GO, fix 01-CONTEXT.md or re-run /gsd-plan-phase 1 with the gap named in your prompt. Do not hand-edit the plans and hope the executor agrees.

  6. Execute:

    /clear
    /gsd-execute-phase 1

    Independent plans run in parallel waves (up to 3 executors by default), each in a fresh context, each task committed on its own. A verifier then checks the code against the phase’s requirements and writes 01-VERIFICATION.md. A checkpoint stops the run only for a human decision, such as an irreversible schema change.

  7. Walk the deliverables:

    /gsd-verify-work 1

    GSD turns each SUMMARY.md into user-visible checks, one at a time: “Adding https://example.com with a 30 s interval stores a result row within 35 s.” Type pass or describe what is wrong. On a failure it writes fix plans for /gsd-execute-phase 1 --gaps-only.

  8. Close the security gate and ship:

    /gsd-secure-phase 1
    /gsd-ship 1

    Ship runs its gates (verification passed, zero open threats, a clean tree; see the next section), then pushes and opens a pull request built from the phase artifacts.

Repeat steps 3 to 8 for phases 2 and 3; /gsd-progress tells you where you are. For an existing codebase, start with /gsd-onboard. For a one-off change, /gsd-quick "fix the timeout parsing" keeps atomic commits and state but skips the optional agents.

How do you know a GSD phase actually works?

Section titled “How do you know a GSD phase actually works?”

GSD automates roadmap, planning, execution, requirement coverage and the pull request. Your share is judgement: observable criteria, reading plans, UAT and merging. It stacks five checks, none of which requires you to read the diff:

  1. Plan-checker. Every plan is checked against the phase goal before execution; do not turn it off with --skip-verify.
  2. Per-task verify commands. Each task’s <verify> command runs after the executor writes the code. With --tdd, the first run is red by design.
  3. Verifier. 01-VERIFICATION.md maps every requirement in the phase to evidence. The verifier fingerprints the files its verdict covered, so editing one afterwards marks the result stale.
  4. UAT. /gsd-verify-work is the one check that needs you: you exercise the behaviour and GSD records the result in 01-UAT.md.
  5. Ship gates. In 1.14.0 and 1.15.0, verification must be passed (not stale), and with the default workflow.security_enforcement: true the phase’s SECURITY.md must report threats_open: 0. The working tree must be clean.

All five run inside the agent that wrote the code. Add one that does not: CI on the pull request, from a clean checkout, with no secrets and read-only access:

.github/workflows/pr-gate.yml
name: pr-gate
on: pull_request
permissions:
contents: read
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
- uses: actions/setup-node@v7
with:
node-version: 24
- run: npm ci
- run: npx tsc --noEmit
- run: npx vitest run

For more independence, /gsd-code-review 1 --depth=deep reviews the phase’s changed files before UAT, and an agent review on the pull request adds a second model’s view. Sign-off stays with you. On a team, the pull request body GSD writes is a ready-made evidence bundle for the reviewer.

GSD spends context in three places, and each needs a different fix.

Always-on cost. On Claude Code 2.1.283, claude plugin details gsd-core@gsd-core projects about 10,700 tokens in every session: 144 skills, 64 agents and 7 hooks, measured on plugin 1.14.0. Superpowers projects about 838 and Everything Claude Code about 41,515. Much of it is duplication: the plugin ships both plan-phase and gsd-plan-phase, while the npm installer writes one form of each (72 skills). Measure your own install:

Terminal window
claude plugin details gsd-core@gsd-core # terminal, plugin path only

On the npm path, run /gsd-surface status (enabled skills with a token summary) or compare /context in a fresh session before and after installing. codex-cli 0.157.1 has no equivalent of details (codex plugin offers only add, list, marketplace and remove); compare the context meter there.

To cut the listing without losing the loop, switch off skill clusters you do not use, for example /gsd-surface disable ui or /gsd-surface disable ai_eval; it re-stages the skills in place, with no reinstall. The installer’s --minimal flag is the bigger cut (its help text estimates about 700 tokens instead of about 12,000), but in 1.14.0 and 1.15.0 neither the core nor the standard profile includes ship or secure-phase, so step 8 of this walkthrough disappears with it.

Per-invocation cost. A skill’s body is small (about 1.8k tokens for gsd-plan-phase), but it @-includes a workflow file into your main session: about 96 KB for workflows/plan-phase.md and 91 KB for execute-phase.md in 1.15.0. Hence the /clear between steps.

Per-phase cost. Every step spawns subagents (researcher, planner, plan-checker, one executor per plan, verifier, debuggers), each starting from an agent definition of about 47 KB in 1.15.0 plus the .planning/ files it reads; plan and execute declare effort: max in Claude Code. The default balanced profile puts the planner on opus and most agents on sonnet, which Claude Code resolves to the models on the models hub; the other tiers are in GSD’s gsd-core/references/model-profiles.md. On Codex, the 1.15.0 agent .toml files pin no model, so they run on your session’s model unless you set GSD’s model_overrides.

The settings that change the bill, in the order to try them:

LeverCommandWhat you give up
Disable unused skill clusters/gsd-surface disable uiCommands you would not have used on this project
Skip research on familiar stacks/gsd-plan-phase 1 --skip-research, or workflow.research: falseLibrary and pattern research before planning
Cheaper model profile/gsd-config --profile budgetPlanner and checker quality on hard phases
Coarser plans/gsd-plan-phase 1 --granularity coarseFewer, larger plans; more risk of stubs
Use /gsd-quick for small fixes/gsd-quick "…"The full discuss, plan and verify loop

Measure the real spend with /usage after a phase (Claude Code and Codex both have it), and track it per phase over a milestone with the approach in agent cost tracking.

Should you let /gsd-autonomous run the remaining phases?

Section titled “Should you let /gsd-autonomous run the remaining phases?”

/gsd-autonomous runs discuss, plan and execute for every remaining phase without stopping, and --from, --to and --only bound the range. It runs the verifier, but when a phase needs human verification it records verification_deferred_human in STATE.md and moves on. So an overnight run ends with code that passed automated checks and a list of /gsd-verify-work sessions you still owe.

Run it only when the success criteria are observable, your tests fail on wrong behaviour (see oracle strength), and the agent runs in a disposable environment with no production credentials. --no-reversibility-gates on plan-phase removes the human checkpoint before irreversible decisions; do not combine it with an unattended run on anything that holds real data. For single-goal runs without GSD’s ceremony, the built-in /goal command is lighter, and hours of autonomy compares GSD with Ralph loops.

Your situationBetter fitWhy
A one-evening feature in an existing repoSuperpowers or plain plan modeGSD’s roadmap and phase files cost more than the feature
A team that reviews specs, not plansSpec KitIts artifacts are feature specs a team can review; GSD’s are one builder’s plans
Brownfield changes that must leave a living specOpenSpecOpenSpec merges each change into a current spec; GSD archives phases
One well-defined goal with a strong test oracle/goalOne command, no framework, no 10,700-token listing
An organisation that needs a supplier reviewWait, or pin and auditGSD Core is a community fork after a governance incident
A multi-phase greenfield build by one personGSDDurable state, fresh contexts and verification gates are what it was built for

The frameworks comparison puts GSD next to every other framework on ceremony, artifacts and context cost.

What breaks when you run GSD, and how do you recover?

Section titled “What breaks when you run GSD, and how do you recover?”
SymptomCauseRecovery
/gsd-* commands missing after installThe agent was not restarted, or you installed the plugin and typed the npm spellingRestart; check the / menu for /gsd-, /gsd-core: or $gsd-
Plugin commands run but fail on their backing logicgsd-tools is not on PATH, or node is not available to the hooksInstall the npm package as well, and confirm node --version works in the agent’s shell
An executor stops with “Permission denied” on BashManual mode with no allow rules for git and your test runnerAdd allow rules for git add, git commit and your test command to .claude/settings.local.json
Execution produces stubs or half-built filesPlans too large for one executorRe-plan with --granularity fine; GSD’s guidance is two or three tasks per plan
The main session gets slow and forgetfulThe orchestrating session itself filled up/gsd-health --context (warns from 60% and suggests /gsd-thread), or /clear then /gsd-resume-work
“FATAL: worktree base mismatch — HEAD is …, expected …”Your branch is ahead of the default branch, and executor worktrees fork from origin/HEADGSD falls back to sequential execution on its own; see its worktree base-mismatch guide to restore parallel runs
Codex agents fail on an unknown modelAgent .toml files from an older install pin an Anthropic alias such as sonnetReinstall with --codex --global, which writes the files without a model pin; gsd-tools validate agents flags any pin left behind
/gsd-ship blocks with PHASE_VERIFICATION_INCOMPLETEVerification is stale, failed or deferred/gsd-verify-work 1, then ship again
/gsd-ship blocks with SECURITY_SHIP_GATE_NO_REVIEWNo phase SECURITY.md/gsd-secure-phase 1, resolve open threats, ship again

A subagent can report failure even though its commits landed. Check git log --oneline -10 before you re-run anything: the commits are the ground truth.