Skip to content

Skills for Context and Token Discipline: caveman, handoff, karpathy-guidelines

Context and token skills are Agent Skills that control what a coding agent writes, reads, and loses between sessions. caveman shortens the agent’s prose, handoff writes a portable brief that a fresh session resumes from, and karpathy-guidelines keeps each diff to the request. None of them proves the work is correct; tests and CI still do.

You are three hours into a webhook-retry migration. The context meter reads 70%, the agent narrates every tool call in four paragraphs, and the next compaction will squeeze the reasons behind six decisions into one line of summary. Tomorrow’s session will re-derive those decisions, or quietly contradict them. This page is for developers who run long agent sessions, and for tech leads who want one team-wide way to carry work from one session to the next.

What you get from context and token skills

Section titled “What you get from context and token skills”
  • A map of five skills and one plugin: what each saves, what it costs, and when it is the wrong tool
  • A decision table for the phase boundary: compact, clear, fork, or hand off
  • Install commands for Claude Code, Codex, and Cursor, checked against skills CLI 1.7.0 on 2026-09-26
  • A worked handoff at 70% context, with the prompts that write, audit, and resume it
  • A multi-session workflow in which commits and tests, not the conversation, carry the state

Which context and token skills solve which problem?

Section titled “Which context and token skills solve which problem?”

Each skill attacks a different leak. Pick by the symptom you see, not by the install count.

Skill (repository)What it changesFiresPick it when
caveman (JuliusBrussee/caveman)The agent answers in short fragments. Code, commands, paths, and exact errors stay verbatim; security warnings and irreversible-action confirmations return in full sentencesOn /caveman, “be brief” or “less tokens”; the Claude Code plugin turns it on at every session startProse, not code, fills the conversation
caveman-compress (same repository)Rewrites a memory file such as CLAUDE.md in compressed form and keeps a readable backupOn /caveman-compress <file>A memory file loaded into every session has grown long
handoff (mattpocock/skills)Writes a handoff document that references specs, plans, and commits instead of copying them, adds a “suggested skills” section, and redacts secretsOnly when you type /handoff (disable-model-invocation: true)The work must outlive the session: overnight, another harness, a colleague, or a forked side task
karpathy-guidelines (multica-ai/andrej-karpathy-skills)Four rules: think before coding, simplicity first, surgical changes, goal-driven execution with a verify step per plan itemWhen the agent writes, reviews, or refactors codeDiffs sprawl, and every sprawling diff costs context to read and time to verify
Agent Skills for Context Engineering (muratcankoylan)17 skills on context design, such as context-degradation, context-compression, filesystem-context, and multi-agent-patternsWhen you design or debug an agent harnessYou build agents or subagent setups, not only use them
planning-with-files (OthmanAdi)Keeps task_plan.md, findings.md, and progress.md on disk and re-reads the plan as work proceedsPlugin commands and hooksA task runs for days and needs a durable plan; covered in persistent memory plugins

Popularity, as of 2026-09-26: skills.sh all-time installs were handoff 869,575 and caveman 539,286. Both counts come from the third-party LinklyAI/best-skills scrape of that date (a secondary source; re-read the live numbers on skills.sh). GitHub stars on the same date: multica-ai/andrej-karpathy-skills 215.2k, JuliusBrussee/caveman 107.9k, OthmanAdi/planning-with-files 27.1k, and muratcankoylan/Agent-Skills-for-Context-Engineering about 18k. Stars measure attention on a repository, not use of one skill.

Should you compact, clear, fork, or hand off?

Section titled “Should you compact, clear, fork, or hand off?”

Every option except “keep going” turns the conversation into a summary. The difference is where the summary lives and who can read it. At most boundaries /compact is enough; reach for a handoff file only when the work has to travel.

MoveWhat survivesUse it when
Keep goingEverythingThe task finishes well inside the window
/compact [focus]A summary inside this session, steered by your focus textSame tool, same directory, same task, and you stay in the loop
/clearNothing from the conversation; project memory files reloadThe finished phase is disposable
/fork, /branch (Claude Code), /fork or codex fork (Codex)An exact copy of the contextA side experiment on the same machine and tool
/handoffA file you can read, correct, move, and shareThe next session starts later, in another tool, or with another person
On-disk plan (planning-with-files or your own plan file)The plan and progress, updated as work happensThe task spans days, whatever you do at each boundary

Claude Code 2.1.283 has no bundled /handoff command (checked against its command reference on 2026-09-26), so /handoff always means the skill.

How do you install the context and token skills?

Section titled “How do you install the context and token skills?”

The portable route is the skills CLI (npm skills 1.7.0). It installs into each agent’s own skills folder: .claude/skills/ for Claude Code and .agents/skills/ for Codex. It records every install in skills-lock.json. Install at project scope so the set is reviewed and committed like code.

  1. Install the skills in the repository root.

    Terminal window
    # handoff: portable copy, invoked as /handoff
    npx skills add mattpocock/skills --skill handoff -a claude-code -y
    # caveman: the output style and the memory-file compressor only
    npx skills add JuliusBrussee/caveman --skill caveman caveman-compress -a claude-code -y
    # karpathy-guidelines as a plugin (the README's recommended route)
    claude plugin marketplace add multica-ai/andrej-karpathy-skills
    claude plugin install andrej-karpathy-skills@karpathy-skills
    # context engineering: name the skills; --full-depth is required
    npx skills add muratcankoylan/Agent-Skills-for-Context-Engineering --full-depth \
    --skill context-degradation context-compression -a claude-code -y

    Alternatives: claude plugin install mattpocock-skills --scope project installs all 25 Pocock skills, namespaced as /mattpocock-skills:handoff. claude plugin marketplace add JuliusBrussee/caveman && claude plugin install caveman@caveman installs the caveman plugin, whose SessionStart and UserPromptSubmit hooks run Node scripts and switch caveman on in every session.

  2. Check what landed. npx skills list shows the installed skills. In Claude Code, claude plugin details andrej-karpathy-skills@karpathy-skills prints a plugin’s inventory and projected context cost, and /context before and after an install shows the difference in the current session.

  3. Read the files, then commit them. Skills run with the agent’s permissions, and caveman-compress ships Python scripts that call Claude. Commit .agents/skills/, .claude/skills/, and skills-lock.json (or .claude/settings.json for plugins), and review changes to them in pull requests. See skill supply-chain security for the review checklist.

What does a handoff at 70% context look like?

Section titled “What does a handoff at 70% context look like?”

The example below is an illustration of the steps, not a recorded run. The repository, commits, and file names are invented for the example.

Why 70%: at that point the agent still writes the handoff from the full conversation, not from an earlier compaction summary, and there is room left for the handoff itself. It is a rule of thumb, not a vendor threshold. Claude Code on a model with a native 1M window compacts at about 967K tokens by default, so waiting for auto-compaction means waiting a long time with a lot of stale context.

  1. Put the meter where you can see it. In Claude Code, add a status line to ~/.claude/settings.json; the example comes from Claude Code’s status line documentation:

    {
    "statusLine": {
    "type": "command",
    "command": "jq -r '\"[\\(.model.display_name)] \\(.context_window.used_percentage // 0)% context\"'"
    }
    }

    used_percentage measures against the model’s full window. If you set CLAUDE_CODE_AUTO_COMPACT_WINDOW, the percentage no longer tells you when compaction runs. /context shows the full breakdown. In Codex, add context usage to the footer with /statusline and hand off at the same 70% used; /status shows token usage. In Cursor, hand off at each phase boundary instead of at a number.

  2. Reach a green checkpoint first. Finish the current slice, run the tests, and commit. A handoff written in the middle of an edit describes a state that no commit holds.

  3. Write the handoff. Pass the next session’s job as the argument, so the skill keeps the reasoning that bears on it.

  4. Move the file somewhere durable. The skill’s SKILL.md tells the agent to save to the operating system’s temp directory, not the workspace. A reboot, a container reset, or a sandbox that discards /tmp deletes it. Copy it into a git-ignored folder in the repository:

    Terminal window
    mkdir -p .handoffs && grep -qx '.handoffs/' .gitignore || echo '.handoffs/' >> .gitignore
    cp /tmp/handoff-webhook-retry.md .handoffs/2026-09-26-webhook-retry.md

    Replace /tmp/handoff-webhook-retry.md with the path the agent reported.

  5. Audit the handoff before you clear. The next agent treats the document as a contract and does not re-check it. A belief written as a fact (“the replay endpoint is done”) becomes a false premise.

  6. Start fresh and resume. Clear the session: /clear webhook-retry-day1 in Claude Code (the name labels the old conversation in the /resume picker), /new in Codex, or a new Agent chat in Cursor. Then point the new session at the file; do not paste the summary into a shell command, where backticks and $(...) get mangled.

A good handoff for this task is short, because the spec and plan already hold the settled detail:

# Handoff: webhook retry, slice 4
## Goal for the next session
Dead-letter table and replay endpoint (docs/plans/webhook-retry.md, slice 4).
## Verified
- Slices 1-3 merged on branch feat/webhook-retry at 4f5e6a7.
- `npm test -- webhooks` passes at 4f5e6a7 (41 tests).
## Assumed, not checked
- Migration 021 (dead_letter table) is drafted in db/migrations/, not applied.
## Decisions and why
- Retries stop after 6 attempts: spec docs/specs/webhook-retry.md section 3.
- Replay is admin-only: ADR docs/adr/0012-replay-auth.md.
## Next steps
1. Apply migration 021 locally; run the migration test.
2. Write the failing test for POST /admin/webhooks/:id/replay.
## Suggested skills
tdd, diagnosing-bugs

How do you run a long task across sessions without compaction loss?

Section titled “How do you run a long task across sessions without compaction loss?”

Compaction loss is a symptom of state that lives only in the conversation. The fix is to move state into artifacts the next session can read and CI can check, and to treat every session boundary as a checkpoint.

  1. Plan on disk. Write the spec and a sliced plan to files before building; grill-me is the Pocock skill for getting there. For multi-day work, planning-with-files keeps plan, findings, and progress files current as the agent works.

  2. Build in small slices. karpathy-guidelines makes the agent state assumptions, touch only what the request needs, and name a verify step for each plan item. A small diff costs less context to read now and less review time later.

  3. End every slice green and committed. The commit and the passing tests are the checkpoint. The conversation is not.

  4. Keep the window lean. Turn on /caveman lite for long debugging or exploration phases, where narration dominates. Run /caveman-compress CLAUDE.md once if your memory file has grown long, then review the diff line by line: that file steers every session. Keep focus text on compactions, for example /compact keep the decisions list, open questions, and failing test names.

  5. At the phase boundary, pick the move from the decision table. Same tool and same task: compact. Next session tomorrow, in another tool, or for someone else: hand off, audit, and clear.

  6. Resume with a test run, not a recap. The first action of every new session is the test command from the handoff. A green run that matches the handoff lets the work continue; a mismatch stops it.

The session commands differ by tool; the loop does not.

Meter: status line context_window.used_percentage, or /context. Boundary: /compact [focus], /clear [name], /fork, /branch, /resume. Compact earlier with /autocompact 500k (100K to 1M) or CLAUDE_AUTOCOMPACT_PCT_OVERRIDE, which can only lower the threshold. Session cost: /usage.

Where Agent Skills for Context Engineering fits

Section titled “Where Agent Skills for Context Engineering fits”

The context-engineering collection is for the tech lead who designs the harness: subagent layouts, memory, tool descriptions, and evals. context-degradation names five failure patterns (lost-in-middle, poisoning, distraction, confusion, and clash). context-compression argues for optimising tokens per task, including the cost of re-fetching what a summary dropped, rather than tokens per request. Each fired skill is large (about 18–19 KB of SKILL.md), so install two or three by name, not the collection.

How do you prove the resumed work is still correct?

Section titled “How do you prove the resumed work is still correct?”

A shorter conversation is not evidence. These skills change how the agent talks and what it carries; the proof stays outside the agent.

  • The handoff is checked against the repository. Every “Verified” line maps to a commit and a passing test. The audit prompt above downgrades the rest before /clear, and the resume prompt re-runs the tests before any edit.
  • CI is the gate. caveman never rewrites code blocks, and karpathy-guidelines only shapes the diff. The pull request still passes typecheck, lint, and tests, and ships an evidence bundle a reviewer checks without rereading the diff.
  • Savings are measured on your work. The token-saving figures in the caveman README are the author’s and third parties’ as the README reports them; this page did not verify them. The README itself says agentic sessions save far less than chat, because most tokens are code and tool calls the skill never touches. Compare /usage and /context on similar tasks with and without the skill before you standardise it.
  • karpathy-guidelines shows up in the diff. Track files changed and lines changed per pull request before and after adoption. The README’s own signals are fewer unrequested changes and clarifying questions asked before implementation, not after mistakes.

Who signs off. The developer who writes the handoff audits it before clearing the session. The reviewer approves the pull request on green CI and the evidence bundle. The tech lead owns the committed skill set and changes it through review.

What do context and token skills cost in context?

Section titled “What do context and token skills cost in context?”

A skill’s name and description load into every session; the body loads when the skill fires. Sizes of SKILL.md on main, 2026-09-26:

What loadsWhenSize
handoffWhen you type /handoff894 bytes
cavemanWhen activated, then it stays in the conversation7,061 bytes
caveman-compressWhen you compress a file4,697 bytes
karpathy-guidelinesWhen the agent writes, reviews, or refactors code2,518 bytes
context-degradation / context-compressionWhen the skill fires19,180 / 18,214 bytes
mattpocock-skills plugin (25 skills)Every sessionabout 1,609 tokens (claude plugin details, Claude Code 2.1.283)
planning-with-files pluginEvery sessionabout 1,124 tokens (same measurement)

At roughly four bytes per token, caveman costs about 1,800 tokens of input once it is active. It pays back only when the prose it removes exceeds that, which is why it suits long, chatty sessions and not short, code-heavy ones.

What breaks when you use context and token skills?

Section titled “What breaks when you use context and token skills?”

The handoff vanished. The temp directory was cleared between sessions or by a reboot. Recovery: copy the file into .handoffs/ as soon as it is written (step 4), and ask the agent for the absolute path every time.

The new session trusts a false “done”. The handoff recorded an assumption as a fact, and the next agent built on it. Recovery: run the audit prompt before /clear; in the resume prompt, require the test run and a stop on any mismatch.

caveman makes the answers unclear. At ultra, the order of steps in a multi-step instruction can blur. The skill tells the model to return to full sentences for security warnings, irreversible actions, and ambiguous multi-step sequences, but the model decides when a sequence counts as ambiguous. Recovery: use /caveman lite, or say stop caveman for design discussions and anything a human reads later.

A compressed CLAUDE.md lost a rule. caveman-compress keeps headings, code, paths, and commands, but rewrites the prose around them, and its backup lives outside the repository under $XDG_DATA_HOME/caveman-compress/backups/. Recovery: review the git diff before you commit, and restore from git, not from the backup folder.

Skills overlap and argue. caveman’s surgical-patch and karpathy-guidelines both govern edit scope; two procedures on one trigger make the agent drift between them. Recovery: keep one skill per concern and remove the other with npx skills remove.

Context still fills up fast. The skills address prose and session boundaries, not the big consumers: MCP tool schemas, large file reads, and long command output. Recovery: check /context, prune MCP servers and memory files as described in managing context windows, and delegate noisy searches to subagents.

Where to go next with context and token discipline

Section titled “Where to go next with context and token discipline”