Skip to content

Token management in Cursor

Token management in Cursor means cutting the context the agent does not use: rules that load on every request, files attached “just in case”, long chat histories, and repeated exploration. A cut counts only when a fixed set of benchmark tasks still passes its tests and the measured tokens and model cost per task go down.

Your team’s included usage runs out in the third week of the month. Nobody changed how they work, but the repository now carries a dozen rules set to always apply, an AGENTS.md and an old .cursorrules that repeat them, and chats that run for 40 messages until Cursor compacts them and the agent re-reads files it already read.

This page is for the developer whose usage runs out early and the tech lead who wants the saving committed for the whole team and proven with numbers. For the tool-neutral economics behind it, read the cost of context.

  • A baseline of tokens and model cost per task, measured with @cursor/sdk rather than estimated.
  • A .cursor/hooks.json hook that logs every context compaction, so you can see which chats overflow and how often.
  • A rule audit prompt that converts always-on rules into file-scoped rules, agent-decided rules, or skills.
  • A handoff prompt that replaces a long chat with a short note and a fresh one.
  • A before/after check that proves the saving did not cost quality, and a named person who accepts it.

Every agent request sends the model more than your prompt. Five things fill the context, and you control most of them.

ConsumerWhat fills itWhat you control
Instruction filesRules with alwaysApply: true, a root AGENTS.md, .cursorrules, and (with third-party config loading on) CLAUDE.md and CLAUDE.local.mdRule scope and length, stray instruction files
Attached contextFiles and folders you @-mention, pasted logs, imagesHow much you attach
Conversation historyEvery earlier message, tool call and tool result in the chatWhen you start a new chat
Agent explorationFiles the agent reads and searches it runs to find codePrompt precision, .cursorignore
Cursor’s harnessThe system prompt and tool definitions, including tools from MCP servers you enableOnly which MCP servers are enabled

The instruction-file behavior on this page comes from the agent runtime shipped in @cursor/sdk 1.0.32 (checked 2026-09-26). The Cursor IDE is expected to use the same rule service, but that could not be re-verified on cursor.com, which was blocked. The runtime watches .cursor/rules/**/*.mdc, AGENTS.md, CLAUDE.md, CLAUDE.local.md and .cursorrules. It loads a root .cursorrules as an always-apply rule, so a forgotten legacy file costs tokens on every request. It treats AGENTS.md as instructions unconditionally: the root file on every request, and nested AGENTS.md files in subdirectories as well. It treats CLAUDE.md and CLAUDE.local.md as instructions only when third-party configuration loading is enabled. If your team also uses Claude Code, check that setting before you assume a CLAUDE.md is free in Cursor.

The model decides two more things: how large the context window is and what each token costs. Both vary by model and change often, so they live on the models hub, not here.

A saving you cannot measure is an opinion. Pick three tasks your team does every week, freeze them as prompts, and measure what each costs today.

  1. Freeze three benchmark tasks. Store each prompt in bench/, for example bench/task-01.md: “Add input validation to POST /api/orders and a test that rejects a negative quantity”. Each task needs an acceptance check you can run, usually a named test file.

  2. Install the SDK as a dev dependency. It needs Node.js 22.13 or later. In your terminal:

    Terminal window
    npm install --save-dev @cursor/sdk@1.0.32 tsx
    export CURSOR_API_KEY="$(op read 'op://Dev/cursor-api-key/credential')"

    The second line reads the key with the 1Password CLI; use whichever secret store your team has. Never paste the key into a command line or a file.

  3. Run each task with the measuring script below, on a throwaway branch or worktree, because the local agent edits files and runs commands. Run the acceptance check after each task and record pass or fail.

  4. Record the numbers in a table: task, model, inputTokens, cacheReadTokens, outputTokens, rawCostCents, chargedCents, acceptance result. That table is your baseline.

bench/measure.mts
// Run from the repository root: npx tsx bench/measure.mts MODEL_ID bench/task-01.md
import { readFileSync } from 'node:fs';
import { Agent } from '@cursor/sdk';
const [modelId, taskFile] = process.argv.slice(2);
const agent = await Agent.create({
apiKey: process.env.CURSOR_API_KEY,
model: { id: modelId },
local: { cwd: process.cwd(), settingSources: ['project'] },
});
const run = await agent.send(readFileSync(taskFile, 'utf8'));
const result = await run.wait();
const { cost } = await agent.getUsage();
console.log(JSON.stringify({
task: taskFile,
model: modelId,
status: result.status,
...result.usage,
rawCostCents: cost?.rawCostCents ?? null,
chargedCents: cost?.chargedCents ?? null,
}));
agent.close();

MODEL_ID is a model ID from Cursor.models.list(); a local agent requires one. The .mts extension makes the file an ES module, which the top-level await needs. settingSources: ['project'] loads your project rules, so rule changes show up in the numbers.

In @cursor/sdk 1.0.32, rawCostCents is the undiscounted model token cost and 0 for request-priced usage, and chargedCents is what Cursor actually bills, including discounts and Cursor’s token fee. chargedCents is 0 for plan-included, BYOK and credit-grant usage, while result.usage counts tokens either way. Cost is eventually consistent, so read getUsage() again later if cost is missing. The Cursor SDK guide covers keys and cloud runs.

When a chat fills the model’s context, Cursor compacts it: older messages are summarized, and detail is lost. A preCompact hook fires just before that happens. It only observes; it cannot stop the compaction. In the same SDK runtime, the hook receives JSON on stdin with trigger (auto or manual), model, context_usage_percent, context_tokens, context_window_size, message_count, messages_to_compact and is_first_compaction, and runs with CURSOR_PROJECT_DIR set to the workspace.

.cursor/hooks.json
{
"version": 1,
"hooks": {
"preCompact": [
{ "command": "sh \"$CURSOR_PROJECT_DIR/.cursor/hooks/log-compaction.sh\"" }
]
}
}
.cursor/hooks/log-compaction.sh
#!/bin/sh
# Append one JSON line per compaction, then let the compaction proceed. Requires jq.
jq -c '{at: (now | todate), trigger, model, context_usage_percent, context_tokens,
context_window_size, message_count, messages_to_compact, is_first_compaction}' \
>> "$CURSOR_PROJECT_DIR/.cursor/compaction-log.jsonl"
echo '{}'

Add .cursor/compaction-log.jsonl to .gitignore. After a week, count the lines. Many auto compactions with a high message_count mean your team keeps chats open across tasks, which the handoff prompt below fixes. For hook events and blocking behavior, see automation workflows in Cursor.

Rules are the cheapest fix, because one change applies to every request by every person. The same runtime classifies a rule by its frontmatter in this order:

FrontmatterHow the rule loadsUse it for
alwaysApply: trueOn every requestA short project overview, hard safety rules
globs: setWhen the agent works on matching filesLanguage style, API conventions for one directory
Only description: setWhen the agent judges the description relevantFeature guides, migration playbooks
None of theseOnly when you @-mention the ruleRarely used checklists

A glob-scoped rule for API routes looks like this:

.cursor/rules/api-routes.mdc
---
description: Conventions for HTTP route handlers in src/api
globs: src/api/**/*.ts
alwaysApply: false
---
Follow the handler pattern in @src/api/users.ts: validate with zod at the top,
return typed errors from src/api/errors.ts, and never read process.env in a handler.

The rule points to an example file instead of pasting 200 lines of it, so the agent reads the example only when the rule loads. A long procedure that is not tied to files, such as a release checklist, fits better as an Agent Skill in .cursor/skills/<name>/SKILL.md. Under the Agent Skills format, the agent sees each skill’s name and description and reads the body only when it uses the skill. For rule design in depth, see project rules in Cursor.

Apply the changes on a branch, rerun the three benchmark tasks, and keep the change only if every acceptance check still passes.

Conversation history is the consumer that grows fastest, because every tool result stays in the chat. Start a new chat when you finish a task, switch to another area of the codebase, or see the agent re-read files it already opened. Do not wait for compaction: the summary it writes is Cursor’s choice of what to keep, not yours.

Carry what matters across with a handoff note that you control.

Then open a new chat and send: “Read @docs/handoff/current.md and continue with the step under Open.” The new chat starts with 40 lines instead of the whole history. Add docs/handoff/ to .gitignore unless your team wants the notes in review.

Attaching eight files “to be safe” pays for eight files on every turn of the chat. Give the agent one entry point and a pattern to copy, and let it read further only if it needs to:

# Over-attached
Add a delete endpoint @src/routes/users.ts @src/routes/posts.ts @src/routes/comments.ts
@src/models/user.ts @src/models/post.ts @src/middleware/auth.ts @src/lib/db.ts
# Pointed
Add DELETE /api/users/:id in @src/routes/users.ts, following the existing PATCH handler
in that file. Soft-delete by setting deleted_at. Add a test next to users.test.ts that
checks a second delete returns 404.

The pointed prompt is also more precise, which cuts exploration: a vague prompt makes the agent search the repository to find out what you meant. Keep generated code, snapshots and vendored files out of reach with .cursorignore, which tuning Cursor for large codebases sets up. For @-mentions and other context controls, see context management in Cursor.

MCP servers cost tokens the same way: every enabled server adds its tool definitions to every request. Disable the MCP servers the task does not need (for a project server, remove it from .cursor/mcp.json on the benchmark branch), then rerun one benchmark task and compare inputTokens with and without them.

Choose the model with your benchmark, not by habit

Section titled “Choose the model with your benchmark, not by habit”

A cheaper model that needs three attempts costs more than a stronger one that passes the first time. The rule this site follows: start on the tool’s default model, and switch only when your own measurements say so. Cursor’s model list, defaults and prices change often and could not be re-verified for this page, so check them in Cursor’s model picker and on the models hub.

To compare two models, run the same three benchmark tasks with each MODEL_ID and compare raw cost per passing task, not price per token. A model that fails an acceptance check scores zero for that task, however cheap it was. Model selection in Cursor covers the decision criteria.

Cloud Agents run in isolated VMs with a full development environment, so a vague task pays for exploration on a machine you are not watching. Plan first, locally, with Plan Mode, which writes an implementation plan before any code. Then hand the Cloud Agent the plan and one acceptance check per step. For cloud agents, getUsage() returns a per-run breakdown in runs, so each run’s tokens and cost go into the same baseline table. See Cloud Agents and Automations.

How do you prove a token saving did not cost quality?

Section titled “How do you prove a token saving did not cost quality?”

Fewer tokens is only a win if the work still passes. Accept a change to rules, habits or model only with this evidence:

EvidenceWhere it comes fromPass condition
Acceptance checks on the three benchmark tasksYour test suite, run after each taskAll pass, as before
Tokens per taskresult.usage from bench/measure.mtsLower than the baseline
Raw cost per passing taskgetUsage().cost.rawCostCentsLower than the baseline
Compactions per week.cursor/compaction-log.jsonlFewer auto compactions
Reverted or reworked agent changesYour pull request history over two weeksNo increase

The tech lead signs off on changes to shared rules and team model defaults, and commits the benchmark table next to the rule change so the next person can rerun it. Each developer owns their own chat habits and reads their own compaction log. Numbers for the whole team, and joining cost to pull requests, belong to agent observability.

What goes wrong when you cut Cursor tokens?

Section titled “What goes wrong when you cut Cursor tokens?”
FailureRecovery
Output quality drops after a rule changeThe agent lost context it needed. Diff the rule change, restore the one rule tied to the failing task, and rerun the benchmark.
A glob-scoped rule never loadsThe glob does not match the files the agent edits. Test the pattern with git ls-files 'src/api/**/*.ts' and fix it.
An agent-decided rule is ignoredIts description does not say when it applies. Rewrite it as a trigger: “Use when adding or changing a database migration”.
Tokens stay high after you trimmed .cursor/rules/A root .cursorrules or root AGENTS.md still loads on every request. Rerun the audit prompt and merge or delete the stray file.
The new chat repeats old mistakesThe handoff note left out a decision. Add the decision and its reason under Decisions, then restart.
chargedCents stays at 0The usage is plan-included, BYOK or a credit grant. Compare rawCostCents and tokens instead.
rawCostCents is 0The model is request-priced, so there is no token cost to report. Compare inputTokens and outputTokens instead.
cost is missing from the outputBilling data lags the run. Run getUsage() again a few minutes later.
The compaction log stays emptyjq is missing or the hook path is wrong. Run the script by hand: echo '{}' | CURSOR_PROJECT_DIR=$PWD sh .cursor/hooks/log-compaction.sh.
A cheaper model passes the benchmark but fails in real workThree tasks were too easy. Add the hardest task from last sprint to bench/ and compare again.