Hours of Autonomy: GSD Core, Ralph Loops, /goal and oh-my-claudecode
Autonomous loops keep a coding agent working for hours unattended: the built-in /goal command in Claude Code and Codex, the Ralph technique (the ralph-loop plugin, snarktank’s ralph.sh, frankbria’s ralph), GSD Core’s /gsd-autonomous, and the oh-my-claudecode and oh-my-codex harnesses. Each one is only as trustworthy as the test oracle that decides when it stops.
You have a clear, boring job: 140 legacy JavaScript files to move to TypeScript, or a backlog of 12 small stories with tests already written. You want to start it at 7 p.m. and review a branch at 9 a.m. The last time someone on your team tried, the agent ran on a laptop with --dangerously-skip-permissions, “finished” by loosening two assertions, and spent a week’s usage budget on one night.
This page is for the developer who runs that job and the tech lead who decides where unattended runs are allowed: pick a loop, run it in a sandbox with a stop condition the agent cannot fake, and review evidence in 15 minutes the next morning.
Which autonomous loop fits your job?
Section titled “Which autonomous loop fits your job?”All of these tools stop the agent from ending its turn until a condition holds. They differ in whether each iteration gets a fresh context, who checks the stop condition, and how much they add to every session’s context.
| Your job | Pick | Fresh context per iteration? | Who decides “done” | Always-on context (Claude Code) |
|---|---|---|---|---|
| One goal, one session, a command that proves it | Built-in /goal (Claude Code, Codex) | No, one session | Claude Code: a small fast model after each turn. Codex: Codex itself, checking the goal after each turn; the loop also stops on a usage limit, a goal token budget, or the same blocker in three consecutive turns | None: built in |
| The same prompt re-fed inside one Claude Code or Cursor session | ralph-loop plugin (Anthropic’s for Claude Code, Cursor’s own for Cursor) | No, one session | Exact match of a <promise> string the agent prints | ~84 tokens |
| A backlog of small stories, each fitting one context | snarktank ralph.sh | Yes, a new agent per iteration | Every story in prd.json has passes: true | ~195 tokens (skills plugin) |
| Long unattended runs that must survive API rate limits | frankbria ralph | Yes | Completion indicators plus an explicit EXIT_SIGNAL: true | Not a plugin |
| A multi-phase greenfield build with planning artifacts | GSD Core /gsd-autonomous | Yes, fresh subagents | A verifier per phase; it pauses for you when a phase needs human validation | ~10,700 tokens |
| A multi-agent team inside Claude Code | oh-my-claudecode /autopilot, /ralph | Partly (subagents) | Its evidence-based verify loops (tests, build, lint, type check) | ~3,915 tokens |
| The same, inside Codex | oh-my-codex $ultragoal | Partly | Checkpoint tracking per goal | Not measurable: codex plugin in codex-cli 0.157.1 has only add, list, marketplace and remove |
Two rules decide most choices. If the job fits in one session and “done” is a command, use /goal: it is built in and maintained by the vendor. If it is bigger than one context window, pick a loop that restarts the agent each iteration and keeps state on disk.
Popularity as of 2026-09-26 (GitHub stars, frameworks research dossier): oh-my-claudecode 39.4k, oh-my-codex 33.4k, snarktank/ralph 21.9k, GSD Core 9.9k (its archived original 64.5k), frankbria/ralph-claude-code 9.6k. The ralph-loop plugin showed 196,527 installs in Anthropic’s official marketplace on the same date.
Install and start each loop
Section titled “Install and start each loop”Built-in /goal
Section titled “Built-in /goal”/goal is built into Claude Code since v2.1.139 and into Codex since CLI 0.128.0, on by default since 0.133.0. The checker model and the pause and resume commands are on the /goal command page.
/goal Every file in src/legacy/ is TypeScript, `npx tsc --noEmit` exits 0 and `npm test` passes. Do not edit anything under tests/.Headless, for a sandbox or CI, /goal works with print mode:
claude -p "/goal CHANGELOG.md has an entry for every PR merged this week"/goal is implemented as a session-scoped Stop hook, so it does not run where disableAllHooks is set, where managed settings allow only managed hooks, or in a workspace whose trust dialog you have not accepted.
/goal Every file in src/legacy/ is TypeScript; done when npx tsc --noEmit exits 0 and npm test passes. Never edit tests/.Check that codex features list shows goals stable true. Control the run with /goal pause, /goal resume, /goal edit and /goal clear; goals persist across --resume.
Cursor’s CLI reference listed /goal [objective] as “Rolling out” on 2026-08-28, with no documented evaluator: the agent doing the work also decides when it is done. Confirm it in your build’s /help and put the oracle below, run by CI, behind it. Cursor’s Cloud Agents run unattended in vendor VMs; see background and cloud agents.
Ralph loops: the technique and three packagings
Section titled “Ralph loops: the technique and three packagings”Geoffrey Huntley’s Ralph technique runs an agent on the same prompt in a loop, keeps state in files and git, and relies on tests, types and lints as “backpressure”. The minimal form in ghuntley/how-to-ralph-wiggum is while :; do cat PROMPT.md | claude ; done, which has no cap and no oracle; the packagings add both.
Anthropic’s ralph-loop plugin re-feeds the prompt through a Stop hook inside the current session, so the conversation persists between iterations too.
/plugin install ralph-loop@claude-plugins-official/ralph-loop:ralph-loop "Migrate src/legacy/*.js to TypeScript. All tests pass, tsc exits 0. Output <promise>DONE</promise> when complete." --completion-promise "DONE" --max-iterations 30/ralph-loop:cancel-ralphThe same prompt comes back after each attempt to stop, until the agent prints <promise>DONE</promise> or iteration 30 ends. /ralph-loop:cancel-ralph stops the loop early. State lives in .claude/ralph-loop.local.md. The setup script says --max-iterations defaults to unlimited, and the promise is an exact string match.
snarktank/ralph (Ryan Carson) starts a new agent with a clean context every iteration and works through prd.json one story at a time:
/plugin marketplace add snarktank/ralph/plugin install ralph-skills@ralph-marketplaceWrite a PRD with the prd skill, convert it to prd.json with the ralph skill, copy ralph.sh and the CLAUDE.md prompt template into scripts/ralph/, then run in a sandbox:
./scripts/ralph/ralph.sh --tool claude 20Each iteration implements the next story with passes: false, runs the checks, sets passes: true, appends to progress.txt and commits. When every story passes, the agent prints <promise>COMPLETE</promise> and the script exits. The default is 10 iterations and the default tool is Amp; --tool accepts only amp or claude.
frankbria/ralph-claude-code adds a call-rate limit (100 per hour by default), a circuit breaker and a two-condition exit gate:
git clone https://github.com/frankbria/ralph-claude-code.git && cd ralph-claude-code && ./install.shcd ../my-project && ralph-enable && ralph --monitorralph-enable writes .ralph/ (prompt, plan and .ralphrc); ralph --calls 50 lowers the hourly limit.
The packagings above are Claude Code only (snarktank’s script also drives Amp); in Codex, run the technique as the capped codex exec loop in step 4 of the overnight workflow, inside its sandbox.
Cursor’s own ralph-loop plugin lives in Cursor’s first-party cursor/plugins repository (README checked on GitHub, 2026-09-26). Install it in Agent chat:
/add-plugin ralph-loopThen start the loop in the same chat:
Start a ralph loop: "Migrate src/legacy/*.js to TypeScript. All tests pass, tsc exits 0. Output <promise>DONE</promise> when complete." --completion-promise "DONE" --max-iterations 30Two hooks drive it: an afterAgentResponse hook watches each response for the <promise> tag, and a stop hook sends the original prompt back as a follow-up message until the promise appears or the cap is reached. Stop it early with the plugin’s cancel-ralph skill, which removes the state file. The README says to always pass --max-iterations to prevent runaway loops.
GSD Core’s autonomous mode
Section titled “GSD Core’s autonomous mode”GSD Core, the community fork of Get Shit Done, runs discuss, plan, execute, verify and ship per roadmap phase, with state in .planning/. /gsd-autonomous drives every remaining phase (--from N, --to N and --only N bound the range), but its workflow file says it “pauses only for explicit user decisions (grey area acceptance, blockers, validation requests)”. Two of those stop an overnight run in practice:
- A phase whose verifier returns
human_needed. If you continue without validating, GSD writesverification_deferred_humanwith averify-workresume command intoSTATE.mdand stops autonomous mode. - A one-way-door decision. The planner inserts a human checkpoint before any task rated
one-way, such as a data migration./gsd-plan-phase --no-reversibility-gatessuppresses it; keep the gate on, because a morning question is cheaper than an irreversible migration nobody approved.
So run /gsd-discuss-phase for the phases you hand it before you leave. Install with npx @opengsd/gsd-core@latest --claude --global (or --codex, --cursor), pinned to a version you reviewed; the GSD page has the full walkthrough.
oh-my-claudecode and oh-my-codex
Section titled “oh-my-claudecode and oh-my-codex”Both are multi-agent harnesses by Yeachan Heo that add hooks, agent teams and persistent modes.
/plugin marketplace add https://github.com/Yeachan-Heo/oh-my-claudecode/plugin install oh-my-claudecode/omc-setupThe CLI route is npm i -g oh-my-claude-sisyphus@latest then omc setup. Clarify requirements with /deep-interview "tenant-scoped API keys", then run /autopilot "add tenant-scoped API keys with rotation", or /ralph to keep going until the work is verified. Type cancelomc in the chat (a prompt trigger, not a slash command) to stop active OMC modes. Version 5.5.0 adds 64 skills, 19 agents and 11 hooks.
Keep the README’s rule of one primary loop authority per session: /goal and /ralph together make two stop conditions fight over the same turn. The slash commands need a live session, so the README advises against relying on them in CI.
npm install -g oh-my-codexomx setup --scope project --merge-agentsomx doctorThe workflow is $deep-interview → $ralplan → $ultragoal (multi-goal execution with checkpoints). The README’s launch line, omx --madmax --xhigh, uses --madmax, OMX shorthand for Codex’s --dangerously-bypass-approvals-and-sandbox: run it only in a sandbox, with one named worktree per concurrent session (omx --worktree=feat/task).
Neither harness installs into Cursor. oh-my-claudecode can hand implementation work to Cursor from a Claude Code session: with Cursor’s CLI installed and signed in, inside a tmux session, /team 1:cursor "…" (or omc team 1:cursor "…" in the terminal) runs a Cursor worker, and final approval stays with the lead session.
Run a sandboxed overnight build with a test oracle
Section titled “Run a sandboxed overnight build with a test oracle”The loop is the easy part. What makes an overnight run safe to merge is a stop condition the agent cannot move, a machine it cannot damage, and a review of evidence instead of lines: oracle → sandbox → capped loop → morning review.
-
Write the task and the oracle before any code. The oracle is one script that exits 0 only when the job is done, and the agent must not be able to edit it. For the TypeScript migration:
scripts/oracle.sh #!/usr/bin/env bashset -euo pipefail# 1. Nothing left to migratetest -z "$(find src/legacy -name '*.js' -print -quit)"# 2. Types and testsnpx tsc --noEmitnpm test -- --runThis assumes Vitest; replace
npm test -- --runwith your runner’s non-watch command (for examplenpx jest --ci). Run it now and confirm it fails: an oracle that passes before the work starts proves nothing. If your tests would pass on a stub, fix them first; see oracle strength and acceptance criteria an agent can be held to. -
Freeze what the agent must not touch. Commit
PROMPT.mdand the oracle onmainbefore the run, so the morning diff againstmainin step 5 starts from the same point; that commit isBASE. After every iteration the loop compares the protected paths (tests/,scripts/oracle.sh,package.json,package-lock.json,tsconfig.json,vitest.config.ts,.github/andPROMPT.md) againstBASEand stops if any changed. That catches the most common way loops cheat: weakening a test, excluding it in the runner config or loosening the compiler so the oracle passes. Keeploop.shand its log outside the repository’s working tree (for example/tmp/loop.sh), so they never end up in the agent’s commits or the returned bundle. That is hygiene, not protection: with permissions bypassed the agent runs as the same OS user as the loop and can rewrite both. The in-sandbox check is an early stop. The checks that count run outside the sandbox in step 5: a protected-path diff that only reads git objects, and the oracle in CI from a clean checkout, never on your own machine. -
Start a disposable sandbox with restricted egress. Use the E2B harness from the sandboxes guide: it uploads your repository as a git bundle, allows outbound traffic only to the model API and the package registry, holds no GitHub token, and returns a bundle of new commits. To make it an overnight run, upload
loop.shto/tmp/(the harness already stages its files there, and/tmpis writable by the sandbox’s default user), call it instead of the single agent command, and raisetimeoutMs. The loop logs to/tmp/loop.log, outside the repository; download it alongside the bundle, and treat both the log and the loop’s exit code as hints, because the agent could have edited them. E2B documents a 1-hour maximum on Hobby and 24 hours on Pro, so an overnight run needs Pro or a local sandbox from the same guide. -
Run a capped loop that checks the oracle itself. It restarts the agent each iteration, caps spend and iterations, and never trusts the agent’s own “done”:
/tmp/loop.sh (sandbox only, outside the repository) #!/usr/bin/env bashset -uBASE=$(git rev-parse HEAD); MAX=${MAX:-25}PROTECTED="tests scripts/oracle.sh package.json package-lock.json tsconfig.json vitest.config.ts .github PROMPT.md"for i in $(seq 1 "$MAX"); doecho "=== iteration $i $(date -u +%H:%M) ===" >> /tmp/loop.logclaude -p --dangerously-skip-permissions --max-budget-usd 3 --output-format json < PROMPT.md >> /tmp/loop.log 2>&1if ! git diff --quiet "$BASE" -- $PROTECTED || [ -n "$(git status --porcelain -- $PROTECTED)" ]; thenecho "STOP: protected files changed" >> /tmp/loop.log; exit 3figrep -q '^BLOCKED:' MIGRATION.md 2>/dev/null && { echo "STOP: agent blocked" >> /tmp/loop.log; exit 4; }if bash scripts/oracle.sh >> /tmp/loop.log 2>&1; then echo "DONE at iteration $i" >> /tmp/loop.log; exit 0; fidoneecho "STOP: iteration cap" >> /tmp/loop.log; exit 2--max-budget-usdcaps API spend per call and works only with--print(claude --help, 2.1.283). WithMAX=25and 3 USD per iteration, the night’s API spend is bounded at about 75 USD.--output-format jsonwrites each call’stotal_cost_usdinto/tmp/loop.log. PassANTHROPIC_API_KEYinto the sandbox from your secret store; that API bill is what--max-budget-usdcaps.Use the same
loop.shwith one line changed:Terminal window codex exec --dangerously-bypass-approvals-and-sandbox - < PROMPT.md >> /tmp/loop.log 2>&1codex-cli 0.157.1 has no dollar cap flag, so the iteration cap is your spend limit. Log in inside the sandbox with
printenv OPENAI_API_KEY | codex login --with-api-key, with the key passed from your secret store, never typed on the command line.The Cursor CLI’s headless command was not re-verified for this page. Run the loop with Claude Code or Codex, or give the job to a Cursor Cloud Agent and keep the oracle as a required CI check on its pull request.
The exit code is a first hint, not proof, because the agent can edit
loop.sh: 0 done, 2 cap reached, 3 protected files touched, 4 the agent asked for help. -
Review the morning branch by evidence. Download the bundle and
/tmp/loop.logfrom the sandbox (here to/tmp/out.bundleand/tmp/loop.log). Fetch the bundle into a branch and inspect it without running it: the commands below only read git objects and execute none of the agent’s code. If the protected-path diff is empty, push the branch and open a pull request, so CI, an agent PR review and a human see it like any other change.Terminal window git fetch /tmp/out.bundle agent/ts-migration:agent/ts-migration # creates the branch, does not check it outgit diff --stat main...agent/ts-migration # scope: only src/ and MIGRATION.md?git diff main...agent/ts-migration -- tests scripts package.json package-lock.json tsconfig.json vitest.config.ts .github # must be emptygit log --oneline main..agent/ts-migration | wc -l # roughly one commit per modulegit push origin agent/ts-migration # only if the diff above is emptyThe oracle that counts runs in CI on that pull request, from a clean checkout of
agent/ts-migration, in a job that holds no secrets (permissions: contents: read, nosecrets.*). GitHub runs the workflow file from the pull request’s own branch, so this is your workflow and your oracle only because the protected-path diff above showed.github/andscripts/unchanged sinceBASE. Do not runnpm testor the oracle from this branch on your laptop: both execute code the agent wrote, possibly after a prompt injection, next to your SSH keys and tokens. If you want a result before CI, run it in a fresh disposable sandbox with no credentials.
Sign-off stays with a person, and the pull request goes through the same CI and review as human work. Attach the downloaded loop.log, MIGRATION.md and the CI oracle result to the pull request as its evidence bundle.
How do you prove an overnight run did the job?
Section titled “How do you prove an overnight run did the job?”A loop’s own “done” is the weakest evidence you have. Stack checks, and count only the ones that run outside the sandbox as proof:
| Check | What it catches | Who runs it |
|---|---|---|
| Oracle fails before the run | An oracle that proves nothing | You, before starting |
| Protected-file diff after every iteration, then again on the fetched branch | Weakened or deleted tests, an edited oracle | loop.sh as an early stop (the agent can edit it), then you, outside the sandbox |
| Oracle run by the loop, then by CI on the pull request | “All tests pass” claims that are not true | loop.sh as a hint, then your CI, in a job with no secrets |
| Search for type escapes and skipped tests | Green results bought with any or .skip | Review prompt above, or a lint rule |
| CI from a clean checkout | Anything that only worked inside the sandbox | Your CI |
| A second model’s review | Wrong behaviour the tests do not cover | An agent PR reviewer |
If any of those needs you to read the whole diff, the oracle is too weak; the fix is a better test, not a longer review. Evidence, not diffs explains the shift.
What does an overnight loop cost?
Section titled “What does an overnight loop cost?”Cost has two parts. The always-on part is what a plugin adds to every session, measured with claude plugin details on Claude Code 2.1.283 (the last column of the decision table); /goal adds nothing. Measure your own:
claude plugin install ralph-loop@claude-plugins-officialclaude plugin details ralph-loop@claude-plugins-officialThe per-run part scales with iterations, and of all these tools only Claude Code’s print mode caps dollars. Set three limits before you start: an iteration cap, a per-call budget (--max-budget-usd in Claude Code print mode) and a wall-clock timeout on the sandbox. /usage in your own session does not see a sandbox run, so read real spend from the log or from the API provider’s console (for Codex, the console only):
# Claude Code runs only: Codex does not write total_cost_usdgrep -o '"total_cost_usd":[0-9.]*' /tmp/loop.log | awk -F: '{s+=$2} END {print s}'Stay on the tool’s default model and tune effort first; the models hub lists current defaults.
Team rules for unattended loops
Section titled “Team rules for unattended loops”A tech lead can adopt this checklist as the policy for any run nobody watches:
- Where: a disposable sandbox with egress limited to the model API and package registries; never a laptop or a runner with deploy secrets.
- Stop condition: a committed oracle that fails before the run. No oracle, no unattended run.
- Caps: iterations, per-call budget and sandbox timeout, written in the run script.
- Protected paths: tests, the oracle, dependency manifests, compiler and CI config, checked by the run script after every iteration.
- Output: a branch and a log, never a push to
mainor a deploy. - Tools: built-in
/goalfirst; third-party loops pinned and reviewed like any dependency. - Owner: one named person approves the oracle before the run and the evidence after it.
What breaks in autonomous loops, and how do you recover?
Section titled “What breaks in autonomous loops, and how do you recover?”| Symptom | Cause | Recovery |
|---|---|---|
| The loop is still running in the morning and the usage budget is gone | No --max-iterations, or a promise the agent never prints | /ralph-loop:cancel-ralph (Claude Code plugin), the cancel-ralph skill (Cursor plugin), cancelomc typed in the chat (OMC prompt trigger, not a slash command), /goal clear (Codex); next time set the cap and a per-call budget |
| The oracle is green but the feature is wrong | Tests weakened, skipped or never able to fail | Restore tests/ from BASE, strengthen the oracle, rerun from the last good commit |
The agent printed <promise>DONE</promise> early | The promise is a string match, not a check | Run the oracle in the harness, not the agent; treat the promise as a hint |
/ralph-loop:ralph-loop stops after one turn | Hooks disabled (--bare, disableAllHooks) or an untrusted workspace | Run without --bare, accept the trust dialog, check the Stop hook is listed |
| Each iteration redoes or undoes the last one | One long session drifted, or no progress file | Use a fresh-context loop with a one-line progress note per iteration |
| Iterations produce no commits | Stories too big for one context | Split the work; snarktank’s README requires each story to fit in one context window |
| The loop stops on rate limits at 2 a.m. | API limits hit mid-run | frankbria’s ralph --calls limit, or fewer iterations with smaller stories |
| GSD’s overnight run stopped at phase 2 waiting for an answer | A grey area, a blocker, a human_needed verification or a one-way-door checkpoint | Answer it, or run the verify-work command recorded in STATE.md; next time run /gsd-discuss-phase before leaving |