Skip to content

Hours of Autonomy: GSD Core, Ralph Loops, /goal and oh-my-claudecode

Autonomous loops keep a coding agent working for hours unattended: the built-in /goal command in Claude Code and Codex, the Ralph technique (the ralph-loop plugin, snarktank’s ralph.sh, frankbria’s ralph), GSD Core’s /gsd-autonomous, and the oh-my-claudecode and oh-my-codex harnesses. Each one is only as trustworthy as the test oracle that decides when it stops.

You have a clear, boring job: 140 legacy JavaScript files to move to TypeScript, or a backlog of 12 small stories with tests already written. You want to start it at 7 p.m. and review a branch at 9 a.m. The last time someone on your team tried, the agent ran on a laptop with --dangerously-skip-permissions, “finished” by loosening two assertions, and spent a week’s usage budget on one night.

This page is for the developer who runs that job and the tech lead who decides where unattended runs are allowed: pick a loop, run it in a sandbox with a stop condition the agent cannot fake, and review evidence in 15 minutes the next morning.

All of these tools stop the agent from ending its turn until a condition holds. They differ in whether each iteration gets a fresh context, who checks the stop condition, and how much they add to every session’s context.

Your jobPickFresh context per iteration?Who decides “done”Always-on context (Claude Code)
One goal, one session, a command that proves itBuilt-in /goal (Claude Code, Codex)No, one sessionClaude Code: a small fast model after each turn. Codex: Codex itself, checking the goal after each turn; the loop also stops on a usage limit, a goal token budget, or the same blocker in three consecutive turnsNone: built in
The same prompt re-fed inside one Claude Code or Cursor sessionralph-loop plugin (Anthropic’s for Claude Code, Cursor’s own for Cursor)No, one sessionExact match of a <promise> string the agent prints~84 tokens
A backlog of small stories, each fitting one contextsnarktank ralph.shYes, a new agent per iterationEvery story in prd.json has passes: true~195 tokens (skills plugin)
Long unattended runs that must survive API rate limitsfrankbria ralphYesCompletion indicators plus an explicit EXIT_SIGNAL: trueNot a plugin
A multi-phase greenfield build with planning artifactsGSD Core /gsd-autonomousYes, fresh subagentsA verifier per phase; it pauses for you when a phase needs human validation~10,700 tokens
A multi-agent team inside Claude Codeoh-my-claudecode /autopilot, /ralphPartly (subagents)Its evidence-based verify loops (tests, build, lint, type check)~3,915 tokens
The same, inside Codexoh-my-codex $ultragoalPartlyCheckpoint tracking per goalNot measurable: codex plugin in codex-cli 0.157.1 has only add, list, marketplace and remove

Two rules decide most choices. If the job fits in one session and “done” is a command, use /goal: it is built in and maintained by the vendor. If it is bigger than one context window, pick a loop that restarts the agent each iteration and keeps state on disk.

Popularity as of 2026-09-26 (GitHub stars, frameworks research dossier): oh-my-claudecode 39.4k, oh-my-codex 33.4k, snarktank/ralph 21.9k, GSD Core 9.9k (its archived original 64.5k), frankbria/ralph-claude-code 9.6k. The ralph-loop plugin showed 196,527 installs in Anthropic’s official marketplace on the same date.

/goal is built into Claude Code since v2.1.139 and into Codex since CLI 0.128.0, on by default since 0.133.0. The checker model and the pause and resume commands are on the /goal command page.

/goal Every file in src/legacy/ is TypeScript, `npx tsc --noEmit` exits 0 and `npm test` passes. Do not edit anything under tests/.

Headless, for a sandbox or CI, /goal works with print mode:

Terminal window
claude -p "/goal CHANGELOG.md has an entry for every PR merged this week"

/goal is implemented as a session-scoped Stop hook, so it does not run where disableAllHooks is set, where managed settings allow only managed hooks, or in a workspace whose trust dialog you have not accepted.

Ralph loops: the technique and three packagings

Section titled “Ralph loops: the technique and three packagings”

Geoffrey Huntley’s Ralph technique runs an agent on the same prompt in a loop, keeps state in files and git, and relies on tests, types and lints as “backpressure”. The minimal form in ghuntley/how-to-ralph-wiggum is while :; do cat PROMPT.md | claude ; done, which has no cap and no oracle; the packagings add both.

Anthropic’s ralph-loop plugin re-feeds the prompt through a Stop hook inside the current session, so the conversation persists between iterations too.

/plugin install ralph-loop@claude-plugins-official
/ralph-loop:ralph-loop "Migrate src/legacy/*.js to TypeScript. All tests pass, tsc exits 0. Output <promise>DONE</promise> when complete." --completion-promise "DONE" --max-iterations 30
/ralph-loop:cancel-ralph

The same prompt comes back after each attempt to stop, until the agent prints <promise>DONE</promise> or iteration 30 ends. /ralph-loop:cancel-ralph stops the loop early. State lives in .claude/ralph-loop.local.md. The setup script says --max-iterations defaults to unlimited, and the promise is an exact string match.

snarktank/ralph (Ryan Carson) starts a new agent with a clean context every iteration and works through prd.json one story at a time:

/plugin marketplace add snarktank/ralph
/plugin install ralph-skills@ralph-marketplace

Write a PRD with the prd skill, convert it to prd.json with the ralph skill, copy ralph.sh and the CLAUDE.md prompt template into scripts/ralph/, then run in a sandbox:

Terminal window
./scripts/ralph/ralph.sh --tool claude 20

Each iteration implements the next story with passes: false, runs the checks, sets passes: true, appends to progress.txt and commits. When every story passes, the agent prints <promise>COMPLETE</promise> and the script exits. The default is 10 iterations and the default tool is Amp; --tool accepts only amp or claude.

frankbria/ralph-claude-code adds a call-rate limit (100 per hour by default), a circuit breaker and a two-condition exit gate:

Terminal window
git clone https://github.com/frankbria/ralph-claude-code.git && cd ralph-claude-code && ./install.sh
cd ../my-project && ralph-enable && ralph --monitor

ralph-enable writes .ralph/ (prompt, plan and .ralphrc); ralph --calls 50 lowers the hourly limit.

GSD Core, the community fork of Get Shit Done, runs discuss, plan, execute, verify and ship per roadmap phase, with state in .planning/. /gsd-autonomous drives every remaining phase (--from N, --to N and --only N bound the range), but its workflow file says it “pauses only for explicit user decisions (grey area acceptance, blockers, validation requests)”. Two of those stop an overnight run in practice:

  • A phase whose verifier returns human_needed. If you continue without validating, GSD writes verification_deferred_human with a verify-work resume command into STATE.md and stops autonomous mode.
  • A one-way-door decision. The planner inserts a human checkpoint before any task rated one-way, such as a data migration. /gsd-plan-phase --no-reversibility-gates suppresses it; keep the gate on, because a morning question is cheaper than an irreversible migration nobody approved.

So run /gsd-discuss-phase for the phases you hand it before you leave. Install with npx @opengsd/gsd-core@latest --claude --global (or --codex, --cursor), pinned to a version you reviewed; the GSD page has the full walkthrough.

Both are multi-agent harnesses by Yeachan Heo that add hooks, agent teams and persistent modes.

/plugin marketplace add https://github.com/Yeachan-Heo/oh-my-claudecode
/plugin install oh-my-claudecode
/omc-setup

The CLI route is npm i -g oh-my-claude-sisyphus@latest then omc setup. Clarify requirements with /deep-interview "tenant-scoped API keys", then run /autopilot "add tenant-scoped API keys with rotation", or /ralph to keep going until the work is verified. Type cancelomc in the chat (a prompt trigger, not a slash command) to stop active OMC modes. Version 5.5.0 adds 64 skills, 19 agents and 11 hooks.

Keep the README’s rule of one primary loop authority per session: /goal and /ralph together make two stop conditions fight over the same turn. The slash commands need a live session, so the README advises against relying on them in CI.

Run a sandboxed overnight build with a test oracle

Section titled “Run a sandboxed overnight build with a test oracle”

The loop is the easy part. What makes an overnight run safe to merge is a stop condition the agent cannot move, a machine it cannot damage, and a review of evidence instead of lines: oracle → sandbox → capped loop → morning review.

  1. Write the task and the oracle before any code. The oracle is one script that exits 0 only when the job is done, and the agent must not be able to edit it. For the TypeScript migration:

    scripts/oracle.sh
    #!/usr/bin/env bash
    set -euo pipefail
    # 1. Nothing left to migrate
    test -z "$(find src/legacy -name '*.js' -print -quit)"
    # 2. Types and tests
    npx tsc --noEmit
    npm test -- --run

    This assumes Vitest; replace npm test -- --run with your runner’s non-watch command (for example npx jest --ci). Run it now and confirm it fails: an oracle that passes before the work starts proves nothing. If your tests would pass on a stub, fix them first; see oracle strength and acceptance criteria an agent can be held to.

  2. Freeze what the agent must not touch. Commit PROMPT.md and the oracle on main before the run, so the morning diff against main in step 5 starts from the same point; that commit is BASE. After every iteration the loop compares the protected paths (tests/, scripts/oracle.sh, package.json, package-lock.json, tsconfig.json, vitest.config.ts, .github/ and PROMPT.md) against BASE and stops if any changed. That catches the most common way loops cheat: weakening a test, excluding it in the runner config or loosening the compiler so the oracle passes. Keep loop.sh and its log outside the repository’s working tree (for example /tmp/loop.sh), so they never end up in the agent’s commits or the returned bundle. That is hygiene, not protection: with permissions bypassed the agent runs as the same OS user as the loop and can rewrite both. The in-sandbox check is an early stop. The checks that count run outside the sandbox in step 5: a protected-path diff that only reads git objects, and the oracle in CI from a clean checkout, never on your own machine.

  3. Start a disposable sandbox with restricted egress. Use the E2B harness from the sandboxes guide: it uploads your repository as a git bundle, allows outbound traffic only to the model API and the package registry, holds no GitHub token, and returns a bundle of new commits. To make it an overnight run, upload loop.sh to /tmp/ (the harness already stages its files there, and /tmp is writable by the sandbox’s default user), call it instead of the single agent command, and raise timeoutMs. The loop logs to /tmp/loop.log, outside the repository; download it alongside the bundle, and treat both the log and the loop’s exit code as hints, because the agent could have edited them. E2B documents a 1-hour maximum on Hobby and 24 hours on Pro, so an overnight run needs Pro or a local sandbox from the same guide.

  4. Run a capped loop that checks the oracle itself. It restarts the agent each iteration, caps spend and iterations, and never trusts the agent’s own “done”:

    /tmp/loop.sh (sandbox only, outside the repository)
    #!/usr/bin/env bash
    set -u
    BASE=$(git rev-parse HEAD); MAX=${MAX:-25}
    PROTECTED="tests scripts/oracle.sh package.json package-lock.json tsconfig.json vitest.config.ts .github PROMPT.md"
    for i in $(seq 1 "$MAX"); do
    echo "=== iteration $i $(date -u +%H:%M) ===" >> /tmp/loop.log
    claude -p --dangerously-skip-permissions --max-budget-usd 3 --output-format json < PROMPT.md >> /tmp/loop.log 2>&1
    if ! git diff --quiet "$BASE" -- $PROTECTED || [ -n "$(git status --porcelain -- $PROTECTED)" ]; then
    echo "STOP: protected files changed" >> /tmp/loop.log; exit 3
    fi
    grep -q '^BLOCKED:' MIGRATION.md 2>/dev/null && { echo "STOP: agent blocked" >> /tmp/loop.log; exit 4; }
    if bash scripts/oracle.sh >> /tmp/loop.log 2>&1; then echo "DONE at iteration $i" >> /tmp/loop.log; exit 0; fi
    done
    echo "STOP: iteration cap" >> /tmp/loop.log; exit 2

    --max-budget-usd caps API spend per call and works only with --print (claude --help, 2.1.283). With MAX=25 and 3 USD per iteration, the night’s API spend is bounded at about 75 USD. --output-format json writes each call’s total_cost_usd into /tmp/loop.log. Pass ANTHROPIC_API_KEY into the sandbox from your secret store; that API bill is what --max-budget-usd caps.

    The exit code is a first hint, not proof, because the agent can edit loop.sh: 0 done, 2 cap reached, 3 protected files touched, 4 the agent asked for help.

  5. Review the morning branch by evidence. Download the bundle and /tmp/loop.log from the sandbox (here to /tmp/out.bundle and /tmp/loop.log). Fetch the bundle into a branch and inspect it without running it: the commands below only read git objects and execute none of the agent’s code. If the protected-path diff is empty, push the branch and open a pull request, so CI, an agent PR review and a human see it like any other change.

    Terminal window
    git fetch /tmp/out.bundle agent/ts-migration:agent/ts-migration # creates the branch, does not check it out
    git diff --stat main...agent/ts-migration # scope: only src/ and MIGRATION.md?
    git diff main...agent/ts-migration -- tests scripts package.json package-lock.json tsconfig.json vitest.config.ts .github # must be empty
    git log --oneline main..agent/ts-migration | wc -l # roughly one commit per module
    git push origin agent/ts-migration # only if the diff above is empty

    The oracle that counts runs in CI on that pull request, from a clean checkout of agent/ts-migration, in a job that holds no secrets (permissions: contents: read, no secrets.*). GitHub runs the workflow file from the pull request’s own branch, so this is your workflow and your oracle only because the protected-path diff above showed .github/ and scripts/ unchanged since BASE. Do not run npm test or the oracle from this branch on your laptop: both execute code the agent wrote, possibly after a prompt injection, next to your SSH keys and tokens. If you want a result before CI, run it in a fresh disposable sandbox with no credentials.

Sign-off stays with a person, and the pull request goes through the same CI and review as human work. Attach the downloaded loop.log, MIGRATION.md and the CI oracle result to the pull request as its evidence bundle.

How do you prove an overnight run did the job?

Section titled “How do you prove an overnight run did the job?”

A loop’s own “done” is the weakest evidence you have. Stack checks, and count only the ones that run outside the sandbox as proof:

CheckWhat it catchesWho runs it
Oracle fails before the runAn oracle that proves nothingYou, before starting
Protected-file diff after every iteration, then again on the fetched branchWeakened or deleted tests, an edited oracleloop.sh as an early stop (the agent can edit it), then you, outside the sandbox
Oracle run by the loop, then by CI on the pull request“All tests pass” claims that are not trueloop.sh as a hint, then your CI, in a job with no secrets
Search for type escapes and skipped testsGreen results bought with any or .skipReview prompt above, or a lint rule
CI from a clean checkoutAnything that only worked inside the sandboxYour CI
A second model’s reviewWrong behaviour the tests do not coverAn agent PR reviewer

If any of those needs you to read the whole diff, the oracle is too weak; the fix is a better test, not a longer review. Evidence, not diffs explains the shift.

Cost has two parts. The always-on part is what a plugin adds to every session, measured with claude plugin details on Claude Code 2.1.283 (the last column of the decision table); /goal adds nothing. Measure your own:

Terminal window
claude plugin install ralph-loop@claude-plugins-official
claude plugin details ralph-loop@claude-plugins-official

The per-run part scales with iterations, and of all these tools only Claude Code’s print mode caps dollars. Set three limits before you start: an iteration cap, a per-call budget (--max-budget-usd in Claude Code print mode) and a wall-clock timeout on the sandbox. /usage in your own session does not see a sandbox run, so read real spend from the log or from the API provider’s console (for Codex, the console only):

Terminal window
# Claude Code runs only: Codex does not write total_cost_usd
grep -o '"total_cost_usd":[0-9.]*' /tmp/loop.log | awk -F: '{s+=$2} END {print s}'

Stay on the tool’s default model and tune effort first; the models hub lists current defaults.

A tech lead can adopt this checklist as the policy for any run nobody watches:

  • Where: a disposable sandbox with egress limited to the model API and package registries; never a laptop or a runner with deploy secrets.
  • Stop condition: a committed oracle that fails before the run. No oracle, no unattended run.
  • Caps: iterations, per-call budget and sandbox timeout, written in the run script.
  • Protected paths: tests, the oracle, dependency manifests, compiler and CI config, checked by the run script after every iteration.
  • Output: a branch and a log, never a push to main or a deploy.
  • Tools: built-in /goal first; third-party loops pinned and reviewed like any dependency.
  • Owner: one named person approves the oracle before the run and the evidence after it.

What breaks in autonomous loops, and how do you recover?

Section titled “What breaks in autonomous loops, and how do you recover?”
SymptomCauseRecovery
The loop is still running in the morning and the usage budget is goneNo --max-iterations, or a promise the agent never prints/ralph-loop:cancel-ralph (Claude Code plugin), the cancel-ralph skill (Cursor plugin), cancelomc typed in the chat (OMC prompt trigger, not a slash command), /goal clear (Codex); next time set the cap and a per-call budget
The oracle is green but the feature is wrongTests weakened, skipped or never able to failRestore tests/ from BASE, strengthen the oracle, rerun from the last good commit
The agent printed <promise>DONE</promise> earlyThe promise is a string match, not a checkRun the oracle in the harness, not the agent; treat the promise as a hint
/ralph-loop:ralph-loop stops after one turnHooks disabled (--bare, disableAllHooks) or an untrusted workspaceRun without --bare, accept the trust dialog, check the Stop hook is listed
Each iteration redoes or undoes the last oneOne long session drifted, or no progress fileUse a fresh-context loop with a one-line progress note per iteration
Iterations produce no commitsStories too big for one contextSplit the work; snarktank’s README requires each story to fit in one context window
The loop stops on rate limits at 2 a.m.API limits hit mid-runfrankbria’s ralph --calls limit, or fewer iterations with smaller stories
GSD’s overnight run stopped at phase 2 waiting for an answerA grey area, a blocker, a human_needed verification or a one-way-door checkpointAnswer it, or run the verify-work command recorded in STATE.md; next time run /gsd-discuss-phase before leaving