Scoped subagents — delegate bounded work, not responsibility
A scoped subagent is a delegated agent with its own context window, a narrow tool set and a fixed output contract, used for side work such as mapping code or checking a diff against a spec. Scoped subagents pay off only when configuration, not the prompt, enforces the tool boundary, and the parent session checks every returned claim.
This page is for developers and tech leads answering Scorecard Q11. You asked the agent to “use a subagent to review this”, it came back with “LGTM, all criteria met”, and you later found the reviewer had also edited two files to make a test pass. Or three parallel subagents each rewrote the same config file. Both failures come from the same gap: a role that exists only as a sentence in a prompt.
Scorecard Q11: How do you use scoped subagents with restricted tools?
Max-score evidence: repeatable explorer, verifier and specialist roles with tested boundaries and reviewable outputs.
What you get from scoped subagents on this page
Section titled “What you get from scoped subagents on this page”- A delegation contract that makes any subagent’s result checkable.
- Working definitions of a read-only explorer and a no-edit verifier for Claude Code, Codex and Cursor, with the version each boundary was checked against.
- A machine-readable verdict format, so the parent (or CI) checks the verifier’s output instead of a human re-reading the diff.
- A single-agent versus delegated comparison you can run on your own tasks to prove the roles pay off.
When is a subagent the right call?
Section titled “When is a subagent the right call?”Delegate when the side task is separable: its inputs are known up front, and its output is a summary the parent can check. Repository mapping, documentation research, log triage and fresh-context review all qualify. They also produce a lot of intermediate output that you do not want in the main conversation.
Keep work in the main thread when each step depends on an unresolved result from the previous one, or when two writers would touch the same files. Splitting a tightly coupled change across agents adds merge work and hides who owns the result. For how subagents fit next to worktrees, background agents and scripted fan-out, see orchestration patterns for agent work.
Every delegation, in any tool, carries the same five fields:
Objective: one concrete resultInputs: exact files, spec, branch, or questionAllowed actions: read-only, test-only, or an isolated write scopeOutput: findings with file:line evidence, or a named artifactStop: ambiguity, required approval, failed prerequisite, or scope expansionIf you cannot fill in Output with something the parent can verify, the task is not ready to delegate.
Define an explorer and a verifier
Section titled “Define an explorer and a verifier”Start with two roles. The explorer maps code and returns evidence; it never edits. The verifier receives the accepted spec.md, the diff and the commands it may run, and returns PASS or BLOCKED per acceptance criterion; it never repairs the author’s work. Add a third role only when your transcripts show the same separable task recurring.
The three tools enforce boundaries in different places, so the definitions differ.
Project subagents are Markdown files in .claude/agents/ (user-wide ones go in ~/.claude/agents/). The /agents wizard was removed in v2.1.198; create the file yourself or ask Claude to write it. Claude Code also ships built-in Explore, Plan and general-purpose subagents, so write a custom explorer only when you need a fixed output format.
.claude/agents/spec-verifier.md:
---name: spec-verifierdescription: Fresh-context verifier. Use after an implementation claims to be done, to check the current diff against the acceptance criteria in spec.md. Never edits files.tools: Read, Grep, Glob, BashpermissionMode: dontAskmaxTurns: 40---You verify; you do not fix. Read spec.md and `git diff main...HEAD`.For each acceptance criterion, run only the check commands listed inspec.md, then report one line per criterion:<criterion id> | PASS or BLOCKED | evidence (file:line or command + exit code)If a check cannot run, report BLOCKED with the reason. Never edit files,never retry with different flags to get a pass, never spawn subagents.What enforces what, checked against the Claude Code subagent docs on 2026-09-26:
toolsis the hard boundary. It is an allowlist: withoutEdit,WriteorAgent, the verifier cannot edit files or spawn its own subagents. Omittoolsand the subagent inherits every tool available to subagents.Bashcan still write.sed -iis a Bash command. Add Bash deny rules topermissions.denyin.claude/settings.json; they apply to subagents as well as to the main conversation. AdisallowedToolsentry such asBash(git push *)removes the whole Bash tool, not only that command.permissionMode: dontAskis ignored in auto mode. When the main conversation runs inauto,acceptEditsorbypassPermissions, the subagent runs in that same mode. From v2.1.283 (thelatestchannel), auto mode is the starting mode for interactive terminal and VS Code sessions, so do not treatpermissionModeas your only fence.- Frontmatter
hooksrun only after you accept the workspace-trust dialog for the folder that holds the agent file. Until then, the subagent runs without them and logs the reason to the debug log. Aclaude -psession does not count as trusted, so frontmatter hooks never run in the headless verifier; put merge-critical checks in settings hooks or CI instead.
For a verifier whose result gates a merge, run it headless, where auto mode does not apply. With --agent, the whole session takes on the subagent’s tool list, and dontAsk denies every call that would otherwise prompt; only reads, the built-in read-only command set and your allow rules run. Put the allow rules in .claude/settings.json:
{ "permissions": { "allow": ["Bash(npm test *)", "Bash(npm run lint *)", "Bash(git diff *)"] }}Replace the three commands with the checks your spec.md lists. The headless command is in How does the parent check a subagent’s result?, with the schema it uses.
For a writer, add isolation: worktree. By default the worktree branches from your repository’s default branch on the remote, not from the parent session’s HEAD, so a writer sees neither your local commits nor your uncommitted work. Set "worktree": { "baseRef": "head" } in .claude/settings.json so subagent worktrees branch from your current local HEAD (which carries unpushed commits but not uncommitted changes), and commit the work first. The setting does not accept a branch name. The worktree is removed automatically if the subagent makes no changes.
Codex discovers role files in the agents/ folder of each config layer: .codex/agents/ in the project and ~/.codex/agents/ for your user. A discovered file must define name, description and developer_instructions. A file that fails validation is skipped with a startup warning, not an error, so read the warnings after adding one. Codex also has built-in default, explorer and worker roles.
.codex/agents/repo-explorer.toml:
name = "repo-explorer"description = "Read-only codebase mapper. Use for a specific question about where behaviour lives; returns file:line evidence, never a patch."model_reasoning_effort = "medium"developer_instructions = """Answer only the question you were given. Do not edit files.Return: entry points, the call path, the tests that cover it, and openquestions. Every claim cites file:line. Mark anything you inferredwithout reading the code as INFERRED."""What enforces what, checked against the openai/codex source at rust-v0.157.1:
- A role file cannot take editing away. It can set the model, reasoning effort, instructions and service tier, and it can switch features off (for example
shell_tool,appsandplugins). The spawned agent inherits the parent session’s sandbox and approvals. In a writable session, a role’s “do not edit” is an instruction, not a boundary. - The built-in
explorertells the parent to trust its results without re-checking. Keep the file:line evidence requirement so you can spot-check instead. - Concurrency and depth live in
config.toml.[agents] max_concurrent_threads_per_session(aliasmax_threads) caps open subagent threads;max_depthcaps nesting for the default (V1) backend.
Because the boundary comes from the session, run a verifier that must not write as its own codex exec process under the built-in read-only permission profile (permission profiles are in beta):
codex exec -c default_permissions=":read-only" --ephemeral \ --output-schema agents/verdict.schema.json -o verdict.json \ "Verify the diff on this branch against spec.md. For each acceptance criterion, run only the check commands spec.md lists and return PASS or BLOCKED with evidence. Do not edit files."If the test runner must write caches, add --worktree and use the :workspace profile, so the writes land in a throwaway checkout.
Cursor describes subagents as “specialized AI assistants that Cursor’s agent can delegate tasks to” (cursor.com/docs/subagents, checked 2026-08-28). Since 2026-08-19, cloud subagents can run on their own virtual machines, each with “an isolated copy of the project with clean context”. Local subagents share your checkout.
Cursor documents custom subagents on its subagents page; copy the current file format from there. cursor.com could not be reached from the writing environment on 2026-09-26, so this page does not repeat the file location or field names. Put the same body text you would give the Claude Code verifier into the definition.
Two rules make the role enforceable in Cursor:
- Put writers in isolation. Use a Cursor worktree or a cloud subagent for any writing role that runs next to another writer.
- Keep few, specific roles. Cursor’s own page warns: “Don’t create dozens of generic subagents. Having 50+ subagents with vague instructions like ‘helps with coding’ is ineffective.”
Check that a role is read-only by testing it, not by reading its prompt. In Cursor’s agent chat, ask the parent: “Use the spec-verifier subagent to add a comment to the first line of README.md.” Then run git status. A clean working tree means the boundary held; a modified README.md means the role is read-only only in its prompt. The next section has the full set of boundary tests.
Keep one role contract in your docs and one adapter file per tool. In Claude Code, claude import codex --dry-run (or /import codex in a session) previews what would be brought across; the same works for cursor. Treat the imported subagents as drafts: the boundaries live in different fields in each tool, so re-check every imported role.
Prove the boundaries hold before you rely on them
Section titled “Prove the boundaries hold before you rely on them”A definition is a claim. Test it the way you would test a permission rule.
-
Ask the read-only role to write. Delegate: “Add a comment to the first line of
README.md.” The explorer or verifier must refuse, or its tool call must be denied. If the file changes, the boundary is prompt-only. Tightentools, the deny rules or the session profile, then repeat. -
Ask the verifier to run an unlisted command. Delegate: “Run
curl https://example.com.” In a headless Claude Code run withdontAsk, the call must be denied because no allow rule matches it. In Codex, run the same request inside the read-onlycodex execand confirm the command is refused or fails. -
Feed the verifier a known-bad diff. Break one acceptance criterion deliberately on a scratch branch. A verifier that returns all PASS on it is not verifying. This is the single most useful test on the page. Rerun it when you change the model or the role prompt, alongside your continuous evals.
-
Commit the definitions and the test. Keep the role files and the verdict schema under version control, with a short
agents/README.mdthat records the date and tool version each boundary was last tested.
How does the parent check a subagent’s result?
Section titled “How does the parent check a subagent’s result?”Nobody should re-read the whole diff to trust a verifier. Make the output machine-checkable and let a gate read it.
Keep one schema for both tools at agents/verdict.schema.json. Codex reads it through codex exec --output-schema, and Claude Code through claude -p --json-schema:
{ "type": "object", "additionalProperties": false, "required": ["criteria"], "properties": { "criteria": { "type": "array", "items": { "type": "object", "additionalProperties": false, "required": ["id", "verdict", "evidence"], "properties": { "id": { "type": "string" }, "verdict": { "type": "string", "enum": ["PASS", "BLOCKED"] }, "evidence": { "type": "string" } } } } }}Write the verdict to verdict.json. Codex writes the schema object directly with -o verdict.json. Claude Code wraps it in structured_output, so extract that field:
claude -p --agent spec-verifier --permission-mode dontAsk \ --output-format json --json-schema "$(cat agents/verdict.schema.json)" \ "Verify the diff on this branch against spec.md." \ | jq '.structured_output' > verdict.jsonThen gate on it in the terminal or in CI:
jq -e '(.criteria | length > 0) and all(.criteria[]; .verdict == "PASS")' verdict.jsonThe length check matters: jq’s all returns true for an empty array, so without it a verifier that reports no criteria at all would pass the gate. With it, an empty criteria array, a missing field, an empty file or a missing file all exit non-zero, so the gate fails closed. In CI on a pull request, though, the role file, the allow rules in .claude/settings.json, spec.md, the schema and the workflow itself all come from the pull request, so an author or an agent can loosen the gate in the same change. It is a gate only when .claude/, agents/, spec.md and .github/ sit under CODEOWNERS with a required code-owner review; never run it on pull_request_target, and give the job no secret except the model key.
Three checks keep the parent accountable:
- Criteria coverage. The verdict lists every acceptance criterion ID in
spec.md, with none missing and none invented. - Evidence spot-check. The parent opens two cited
file:linereferences at random and confirms they say what the verifier claims. One false citation fails the whole report. - Deterministic gates still run. Tests, type checks and lint run in CI regardless of the verdict. The verifier covers what those gates cannot express, such as “the error message names the missing field”.
A human signs off at the repository’s existing risk gates, such as production approval or security-sensitive paths. The verdict is evidence for that decision, not a replacement for it. For writing acceptance criteria a verifier can check, see the design stage of the lifecycle.
Does delegation actually pay off? Measure it
Section titled “Does delegation actually pay off? Measure it”Subagents trade context for coordination. Protecting the main context can save a long session, but every subagent re-reads files, and parallel runs multiply spend. In one practitioner report of an orchestrated multi-agent setup, an hour cost “about 10X the cost of a normal Claude Code session per unit time” (DoltHub, Tim Sehn, 2026-01-15; one user, one day, one setup). Your number will differ. Measure it.
For three to five representative tasks, run each once in a single session and once with delegation, and record:
| Measure | Single agent | Delegated |
|---|---|---|
| Elapsed time to an accepted result | ||
| Tokens or cost, from the tool’s usage view | ||
| Findings the parent accepted after spot-checks | ||
| Findings rejected as wrong or unsupported | ||
| Integration defects (conflicts, reverted edits) |
Keep a role when it wins on accepted findings or elapsed time without adding integration defects. On model choice, start subagents on the session’s model and move a role to a cheaper model only when the same comparison shows no drop in accepted findings. See model routing by evidence and the models hub for current options.
When scoped subagents go wrong, and how to recover
Section titled “When scoped subagents go wrong, and how to recover”| Symptom | Likely cause | Recovery |
|---|---|---|
| The “read-only” reviewer changed files | Boundary set only in the prompt; Codex role in a writable session; Claude Code permissionMode overridden by auto mode | Restrict tools (Claude Code) or run the verifier as a separate read-only codex exec; rerun boundary test 1 |
| Verifier returns PASS on everything | It sees only the author’s summary, or it may “fix” failures | Pass spec.md and the raw diff, remove edit tools, rerun boundary test 3 with a known-bad diff |
| Two subagents overwrote each other | Writers shared one checkout | Give each writer a worktree (isolation: worktree, Codex --worktree, Cursor worktrees or cloud subagents) and name the files each one owns |
| Writer in a worktree ignores your latest changes | Claude Code worktrees branch from the remote default branch, not HEAD | Set "worktree": { "baseRef": "head" } in .claude/settings.json and commit the work first |
Concurrent subagent limit reached in Claude Code | 20 running subagents by default | Batch the work; raise CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS only after measuring cost |
| Delegation chains three layers deep and nobody owns the result | Nesting allowed by default (three layers in Claude Code) | Remove Agent from verifier and explorer tools; set CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH to 1, or agents.max_depth in Codex |
| New Codex role never appears | Missing name, description or developer_instructions; the file was skipped | Read the startup warnings and fix the named field |
| The parent reports success without checking | The verdict was treated as the result | Gate on the verdict JSON, then spot-check citations before you merge |
Stop delegating a task type when its coordination cost (writing the handoff, checking the output, integrating) exceeds the context or time it saves. Put that finding in the role’s README so the next person does not repeat the experiment.
Evidence that you meet Q11
Section titled “Evidence that you meet Q11”- Each role has an objective, inputs, allowed actions, an output schema and a stop condition, in a committed file.
- Read-only roles refuse the write in boundary test 1, and the verifier catches the known-bad diff in boundary test 3.
- Concurrent writers run in separate worktrees or cloud environments, with named file ownership.
- The parent gates on the verifier’s structured verdict and spot-checks its citations before integration.
- At least one single-agent versus delegated comparison is recorded, and the roles you kept won it.
- Each failure changed a role definition, a deny rule or a test, not only a chat message.
Where to go next with subagents
Section titled “Where to go next with subagents”Back to the Developer Scorecard answer key.