Skills for Context and Token Discipline: caveman, handoff, karpathy-guidelines
Context and token skills are Agent Skills that control what a coding agent writes, reads, and loses between sessions. caveman shortens the agent’s prose, handoff writes a portable brief that a fresh session resumes from, and karpathy-guidelines keeps each diff to the request. None of them proves the work is correct; tests and CI still do.
You are three hours into a webhook-retry migration. The context meter reads 70%, the agent narrates every tool call in four paragraphs, and the next compaction will squeeze the reasons behind six decisions into one line of summary. Tomorrow’s session will re-derive those decisions, or quietly contradict them. This page is for developers who run long agent sessions, and for tech leads who want one team-wide way to carry work from one session to the next.
What you get from context and token skills
Section titled “What you get from context and token skills”- A map of five skills and one plugin: what each saves, what it costs, and when it is the wrong tool
- A decision table for the phase boundary: compact, clear, fork, or hand off
- Install commands for Claude Code, Codex, and Cursor, checked against skills CLI 1.7.0 on 2026-09-26
- A worked handoff at 70% context, with the prompts that write, audit, and resume it
- A multi-session workflow in which commits and tests, not the conversation, carry the state
Which context and token skills solve which problem?
Section titled “Which context and token skills solve which problem?”Each skill attacks a different leak. Pick by the symptom you see, not by the install count.
| Skill (repository) | What it changes | Fires | Pick it when |
|---|---|---|---|
caveman (JuliusBrussee/caveman) | The agent answers in short fragments. Code, commands, paths, and exact errors stay verbatim; security warnings and irreversible-action confirmations return in full sentences | On /caveman, “be brief” or “less tokens”; the Claude Code plugin turns it on at every session start | Prose, not code, fills the conversation |
caveman-compress (same repository) | Rewrites a memory file such as CLAUDE.md in compressed form and keeps a readable backup | On /caveman-compress <file> | A memory file loaded into every session has grown long |
handoff (mattpocock/skills) | Writes a handoff document that references specs, plans, and commits instead of copying them, adds a “suggested skills” section, and redacts secrets | Only when you type /handoff (disable-model-invocation: true) | The work must outlive the session: overnight, another harness, a colleague, or a forked side task |
karpathy-guidelines (multica-ai/andrej-karpathy-skills) | Four rules: think before coding, simplicity first, surgical changes, goal-driven execution with a verify step per plan item | When the agent writes, reviews, or refactors code | Diffs sprawl, and every sprawling diff costs context to read and time to verify |
| Agent Skills for Context Engineering (muratcankoylan) | 17 skills on context design, such as context-degradation, context-compression, filesystem-context, and multi-agent-patterns | When you design or debug an agent harness | You build agents or subagent setups, not only use them |
planning-with-files (OthmanAdi) | Keeps task_plan.md, findings.md, and progress.md on disk and re-reads the plan as work proceeds | Plugin commands and hooks | A task runs for days and needs a durable plan; covered in persistent memory plugins |
Popularity, as of 2026-09-26: skills.sh all-time installs were handoff 869,575 and caveman 539,286. Both counts come from the third-party LinklyAI/best-skills scrape of that date (a secondary source; re-read the live numbers on skills.sh). GitHub stars on the same date: multica-ai/andrej-karpathy-skills 215.2k, JuliusBrussee/caveman 107.9k, OthmanAdi/planning-with-files 27.1k, and muratcankoylan/Agent-Skills-for-Context-Engineering about 18k. Stars measure attention on a repository, not use of one skill.
Should you compact, clear, fork, or hand off?
Section titled “Should you compact, clear, fork, or hand off?”Every option except “keep going” turns the conversation into a summary. The difference is where the summary lives and who can read it. At most boundaries /compact is enough; reach for a handoff file only when the work has to travel.
| Move | What survives | Use it when |
|---|---|---|
| Keep going | Everything | The task finishes well inside the window |
/compact [focus] | A summary inside this session, steered by your focus text | Same tool, same directory, same task, and you stay in the loop |
/clear | Nothing from the conversation; project memory files reload | The finished phase is disposable |
/fork, /branch (Claude Code), /fork or codex fork (Codex) | An exact copy of the context | A side experiment on the same machine and tool |
/handoff | A file you can read, correct, move, and share | The next session starts later, in another tool, or with another person |
On-disk plan (planning-with-files or your own plan file) | The plan and progress, updated as work happens | The task spans days, whatever you do at each boundary |
Claude Code 2.1.283 has no bundled /handoff command (checked against its command reference on 2026-09-26), so /handoff always means the skill.
How do you install the context and token skills?
Section titled “How do you install the context and token skills?”The portable route is the skills CLI (npm skills 1.7.0). It installs into each agent’s own skills folder: .claude/skills/ for Claude Code and .agents/skills/ for Codex. It records every install in skills-lock.json. Install at project scope so the set is reviewed and committed like code.
-
Install the skills in the repository root.
Terminal window # handoff: portable copy, invoked as /handoffnpx skills add mattpocock/skills --skill handoff -a claude-code -y# caveman: the output style and the memory-file compressor onlynpx skills add JuliusBrussee/caveman --skill caveman caveman-compress -a claude-code -y# karpathy-guidelines as a plugin (the README's recommended route)claude plugin marketplace add multica-ai/andrej-karpathy-skillsclaude plugin install andrej-karpathy-skills@karpathy-skills# context engineering: name the skills; --full-depth is requirednpx skills add muratcankoylan/Agent-Skills-for-Context-Engineering --full-depth \--skill context-degradation context-compression -a claude-code -yAlternatives:
claude plugin install mattpocock-skills --scope projectinstalls all 25 Pocock skills, namespaced as/mattpocock-skills:handoff.claude plugin marketplace add JuliusBrussee/caveman && claude plugin install caveman@cavemaninstalls the caveman plugin, whoseSessionStartandUserPromptSubmithooks run Node scripts and switch caveman on in every session.Terminal window npx skills add mattpocock/skills --skill handoff -a codex -ynpx skills add JuliusBrussee/caveman --skill caveman caveman-compress -a codex -ynpx skills add multica-ai/andrej-karpathy-skills --skill karpathy-guidelines -a codex -ynpx skills add muratcankoylan/Agent-Skills-for-Context-Engineering --full-depth \--skill context-degradation context-compression -a codex -yCodex reads
.agents/skills/; type/skillsto list what loaded and$handoffto invoke a skill by name. The handoff skill’sagents/openai.yamlsetsallow_implicit_invocation: false, so Codex runs it only when you name it. The context-engineering repository has no Codex plugin marketplace manifest, so the skills CLI is the route there.Terminal window npx skills add mattpocock/skills --skill handoff -a cursor -ynpx skills add JuliusBrussee/caveman --skill caveman caveman-compress -a cursor -ynpx skills add multica-ai/andrej-karpathy-skills --skill karpathy-guidelines -a cursor -ynpx skills add muratcankoylan/Agent-Skills-for-Context-Engineering --full-depth \--skill context-degradation context-compression -a cursor -yThe karpathy repository also ships a committed Cursor rule,
.cursor/rules/karpathy-guidelines.mdc, which you can copy into your own.cursor/rules/instead of installing the skill. Cursor’s skills documentation could not be re-checked on 2026-09-26; if a skill does not appear, ask the agent to use it by name. -
Check what landed.
npx skills listshows the installed skills. In Claude Code,claude plugin details andrej-karpathy-skills@karpathy-skillsprints a plugin’s inventory and projected context cost, and/contextbefore and after an install shows the difference in the current session. -
Read the files, then commit them. Skills run with the agent’s permissions, and
caveman-compressships Python scripts that call Claude. Commit.agents/skills/,.claude/skills/, andskills-lock.json(or.claude/settings.jsonfor plugins), and review changes to them in pull requests. See skill supply-chain security for the review checklist.
What does a handoff at 70% context look like?
Section titled “What does a handoff at 70% context look like?”The example below is an illustration of the steps, not a recorded run. The repository, commits, and file names are invented for the example.
Why 70%: at that point the agent still writes the handoff from the full conversation, not from an earlier compaction summary, and there is room left for the handoff itself. It is a rule of thumb, not a vendor threshold. Claude Code on a model with a native 1M window compacts at about 967K tokens by default, so waiting for auto-compaction means waiting a long time with a lot of stale context.
-
Put the meter where you can see it. In Claude Code, add a status line to
~/.claude/settings.json; the example comes from Claude Code’s status line documentation:{"statusLine": {"type": "command","command": "jq -r '\"[\\(.model.display_name)] \\(.context_window.used_percentage // 0)% context\"'"}}used_percentagemeasures against the model’s full window. If you setCLAUDE_CODE_AUTO_COMPACT_WINDOW, the percentage no longer tells you when compaction runs./contextshows the full breakdown. In Codex, add context usage to the footer with/statuslineand hand off at the same 70% used;/statusshows token usage. In Cursor, hand off at each phase boundary instead of at a number. -
Reach a green checkpoint first. Finish the current slice, run the tests, and commit. A handoff written in the middle of an edit describes a state that no commit holds.
-
Write the handoff. Pass the next session’s job as the argument, so the skill keeps the reasoning that bears on it.
-
Move the file somewhere durable. The skill’s
SKILL.mdtells the agent to save to the operating system’s temp directory, not the workspace. A reboot, a container reset, or a sandbox that discards/tmpdeletes it. Copy it into a git-ignored folder in the repository:Terminal window mkdir -p .handoffs && grep -qx '.handoffs/' .gitignore || echo '.handoffs/' >> .gitignorecp /tmp/handoff-webhook-retry.md .handoffs/2026-09-26-webhook-retry.mdReplace
/tmp/handoff-webhook-retry.mdwith the path the agent reported. -
Audit the handoff before you clear. The next agent treats the document as a contract and does not re-check it. A belief written as a fact (“the replay endpoint is done”) becomes a false premise.
-
Start fresh and resume. Clear the session:
/clear webhook-retry-day1in Claude Code (the name labels the old conversation in the/resumepicker),/newin Codex, or a new Agent chat in Cursor. Then point the new session at the file; do not paste the summary into a shell command, where backticks and$(...)get mangled.
A good handoff for this task is short, because the spec and plan already hold the settled detail:
# Handoff: webhook retry, slice 4
## Goal for the next sessionDead-letter table and replay endpoint (docs/plans/webhook-retry.md, slice 4).
## Verified- Slices 1-3 merged on branch feat/webhook-retry at 4f5e6a7.- `npm test -- webhooks` passes at 4f5e6a7 (41 tests).
## Assumed, not checked- Migration 021 (dead_letter table) is drafted in db/migrations/, not applied.
## Decisions and why- Retries stop after 6 attempts: spec docs/specs/webhook-retry.md section 3.- Replay is admin-only: ADR docs/adr/0012-replay-auth.md.
## Next steps1. Apply migration 021 locally; run the migration test.2. Write the failing test for POST /admin/webhooks/:id/replay.
## Suggested skillstdd, diagnosing-bugsHow do you run a long task across sessions without compaction loss?
Section titled “How do you run a long task across sessions without compaction loss?”Compaction loss is a symptom of state that lives only in the conversation. The fix is to move state into artifacts the next session can read and CI can check, and to treat every session boundary as a checkpoint.
-
Plan on disk. Write the spec and a sliced plan to files before building; grill-me is the Pocock skill for getting there. For multi-day work, planning-with-files keeps plan, findings, and progress files current as the agent works.
-
Build in small slices.
karpathy-guidelinesmakes the agent state assumptions, touch only what the request needs, and name a verify step for each plan item. A small diff costs less context to read now and less review time later. -
End every slice green and committed. The commit and the passing tests are the checkpoint. The conversation is not.
-
Keep the window lean. Turn on
/caveman litefor long debugging or exploration phases, where narration dominates. Run/caveman-compress CLAUDE.mdonce if your memory file has grown long, then review the diff line by line: that file steers every session. Keep focus text on compactions, for example/compact keep the decisions list, open questions, and failing test names. -
At the phase boundary, pick the move from the decision table. Same tool and same task: compact. Next session tomorrow, in another tool, or for someone else: hand off, audit, and clear.
-
Resume with a test run, not a recap. The first action of every new session is the test command from the handoff. A green run that matches the handoff lets the work continue; a mismatch stops it.
The session commands differ by tool; the loop does not.
Meter: status line context_window.used_percentage, or /context. Boundary: /compact [focus], /clear [name], /fork, /branch, /resume. Compact earlier with /autocompact 500k (100K to 1M) or CLAUDE_AUTOCOMPACT_PCT_OVERRIDE, which can only lower the threshold. Session cost: /usage.
Meter: the context-usage item in the footer (configure it with /statusline), and /status for token usage. Boundary: /compact, /new, /fork, /resume; from the shell, codex resume --last and codex fork --last. Skills are invoked as $handoff, $caveman.
Start a new Agent chat at each phase boundary and point it at the handoff and plan files. Cursor’s context indicator and summarisation behaviour could not be verified from cursor.com on 2026-09-26, so this page gives no Cursor threshold.
Where Agent Skills for Context Engineering fits
Section titled “Where Agent Skills for Context Engineering fits”The context-engineering collection is for the tech lead who designs the harness: subagent layouts, memory, tool descriptions, and evals. context-degradation names five failure patterns (lost-in-middle, poisoning, distraction, confusion, and clash). context-compression argues for optimising tokens per task, including the cost of re-fetching what a summary dropped, rather than tokens per request. Each fired skill is large (about 18–19 KB of SKILL.md), so install two or three by name, not the collection.
How do you prove the resumed work is still correct?
Section titled “How do you prove the resumed work is still correct?”A shorter conversation is not evidence. These skills change how the agent talks and what it carries; the proof stays outside the agent.
- The handoff is checked against the repository. Every “Verified” line maps to a commit and a passing test. The audit prompt above downgrades the rest before
/clear, and the resume prompt re-runs the tests before any edit. - CI is the gate. caveman never rewrites code blocks, and karpathy-guidelines only shapes the diff. The pull request still passes typecheck, lint, and tests, and ships an evidence bundle a reviewer checks without rereading the diff.
- Savings are measured on your work. The token-saving figures in the caveman README are the author’s and third parties’ as the README reports them; this page did not verify them. The README itself says agentic sessions save far less than chat, because most tokens are code and tool calls the skill never touches. Compare
/usageand/contexton similar tasks with and without the skill before you standardise it. - karpathy-guidelines shows up in the diff. Track files changed and lines changed per pull request before and after adoption. The README’s own signals are fewer unrequested changes and clarifying questions asked before implementation, not after mistakes.
Who signs off. The developer who writes the handoff audits it before clearing the session. The reviewer approves the pull request on green CI and the evidence bundle. The tech lead owns the committed skill set and changes it through review.
What do context and token skills cost in context?
Section titled “What do context and token skills cost in context?”A skill’s name and description load into every session; the body loads when the skill fires. Sizes of SKILL.md on main, 2026-09-26:
| What loads | When | Size |
|---|---|---|
handoff | When you type /handoff | 894 bytes |
caveman | When activated, then it stays in the conversation | 7,061 bytes |
caveman-compress | When you compress a file | 4,697 bytes |
karpathy-guidelines | When the agent writes, reviews, or refactors code | 2,518 bytes |
context-degradation / context-compression | When the skill fires | 19,180 / 18,214 bytes |
mattpocock-skills plugin (25 skills) | Every session | about 1,609 tokens (claude plugin details, Claude Code 2.1.283) |
planning-with-files plugin | Every session | about 1,124 tokens (same measurement) |
At roughly four bytes per token, caveman costs about 1,800 tokens of input once it is active. It pays back only when the prose it removes exceeds that, which is why it suits long, chatty sessions and not short, code-heavy ones.
What breaks when you use context and token skills?
Section titled “What breaks when you use context and token skills?”The handoff vanished. The temp directory was cleared between sessions or by a reboot. Recovery: copy the file into .handoffs/ as soon as it is written (step 4), and ask the agent for the absolute path every time.
The new session trusts a false “done”. The handoff recorded an assumption as a fact, and the next agent built on it. Recovery: run the audit prompt before /clear; in the resume prompt, require the test run and a stop on any mismatch.
caveman makes the answers unclear. At ultra, the order of steps in a multi-step instruction can blur. The skill tells the model to return to full sentences for security warnings, irreversible actions, and ambiguous multi-step sequences, but the model decides when a sequence counts as ambiguous. Recovery: use /caveman lite, or say stop caveman for design discussions and anything a human reads later.
A compressed CLAUDE.md lost a rule. caveman-compress keeps headings, code, paths, and commands, but rewrites the prose around them, and its backup lives outside the repository under $XDG_DATA_HOME/caveman-compress/backups/. Recovery: review the git diff before you commit, and restore from git, not from the backup folder.
Skills overlap and argue. caveman’s surgical-patch and karpathy-guidelines both govern edit scope; two procedures on one trigger make the agent drift between them. Recovery: keep one skill per concern and remove the other with npx skills remove.
Context still fills up fast. The skills address prose and session boundaries, not the big consumers: MCP tool schemas, large file reads, and long command output. Recovery: check /context, prune MCP servers and memory files as described in managing context windows, and delegate noisy searches to subagents.