Skip to content

Scoped subagents — delegate bounded work, not responsibility

A scoped subagent is a delegated agent with its own context window, a narrow tool set and a fixed output contract, used for side work such as mapping code or checking a diff against a spec. Scoped subagents pay off only when configuration, not the prompt, enforces the tool boundary, and the parent session checks every returned claim.

This page is for developers and tech leads answering Scorecard Q11. You asked the agent to “use a subagent to review this”, it came back with “LGTM, all criteria met”, and you later found the reviewer had also edited two files to make a test pass. Or three parallel subagents each rewrote the same config file. Both failures come from the same gap: a role that exists only as a sentence in a prompt.

Scorecard Q11: How do you use scoped subagents with restricted tools?

Max-score evidence: repeatable explorer, verifier and specialist roles with tested boundaries and reviewable outputs.

What you get from scoped subagents on this page

Section titled “What you get from scoped subagents on this page”
  • A delegation contract that makes any subagent’s result checkable.
  • Working definitions of a read-only explorer and a no-edit verifier for Claude Code, Codex and Cursor, with the version each boundary was checked against.
  • A machine-readable verdict format, so the parent (or CI) checks the verifier’s output instead of a human re-reading the diff.
  • A single-agent versus delegated comparison you can run on your own tasks to prove the roles pay off.

Delegate when the side task is separable: its inputs are known up front, and its output is a summary the parent can check. Repository mapping, documentation research, log triage and fresh-context review all qualify. They also produce a lot of intermediate output that you do not want in the main conversation.

Keep work in the main thread when each step depends on an unresolved result from the previous one, or when two writers would touch the same files. Splitting a tightly coupled change across agents adds merge work and hides who owns the result. For how subagents fit next to worktrees, background agents and scripted fan-out, see orchestration patterns for agent work.

Every delegation, in any tool, carries the same five fields:

Objective: one concrete result
Inputs: exact files, spec, branch, or question
Allowed actions: read-only, test-only, or an isolated write scope
Output: findings with file:line evidence, or a named artifact
Stop: ambiguity, required approval, failed prerequisite, or scope expansion

If you cannot fill in Output with something the parent can verify, the task is not ready to delegate.

Start with two roles. The explorer maps code and returns evidence; it never edits. The verifier receives the accepted spec.md, the diff and the commands it may run, and returns PASS or BLOCKED per acceptance criterion; it never repairs the author’s work. Add a third role only when your transcripts show the same separable task recurring.

The three tools enforce boundaries in different places, so the definitions differ.

Project subagents are Markdown files in .claude/agents/ (user-wide ones go in ~/.claude/agents/). The /agents wizard was removed in v2.1.198; create the file yourself or ask Claude to write it. Claude Code also ships built-in Explore, Plan and general-purpose subagents, so write a custom explorer only when you need a fixed output format.

.claude/agents/spec-verifier.md:

---
name: spec-verifier
description: Fresh-context verifier. Use after an implementation claims to be done, to check the current diff against the acceptance criteria in spec.md. Never edits files.
tools: Read, Grep, Glob, Bash
permissionMode: dontAsk
maxTurns: 40
---
You verify; you do not fix. Read spec.md and `git diff main...HEAD`.
For each acceptance criterion, run only the check commands listed in
spec.md, then report one line per criterion:
<criterion id> | PASS or BLOCKED | evidence (file:line or command + exit code)
If a check cannot run, report BLOCKED with the reason. Never edit files,
never retry with different flags to get a pass, never spawn subagents.

What enforces what, checked against the Claude Code subagent docs on 2026-09-26:

  • tools is the hard boundary. It is an allowlist: without Edit, Write or Agent, the verifier cannot edit files or spawn its own subagents. Omit tools and the subagent inherits every tool available to subagents.
  • Bash can still write. sed -i is a Bash command. Add Bash deny rules to permissions.deny in .claude/settings.json; they apply to subagents as well as to the main conversation. A disallowedTools entry such as Bash(git push *) removes the whole Bash tool, not only that command.
  • permissionMode: dontAsk is ignored in auto mode. When the main conversation runs in auto, acceptEdits or bypassPermissions, the subagent runs in that same mode. From v2.1.283 (the latest channel), auto mode is the starting mode for interactive terminal and VS Code sessions, so do not treat permissionMode as your only fence.
  • Frontmatter hooks run only after you accept the workspace-trust dialog for the folder that holds the agent file. Until then, the subagent runs without them and logs the reason to the debug log. A claude -p session does not count as trusted, so frontmatter hooks never run in the headless verifier; put merge-critical checks in settings hooks or CI instead.

For a verifier whose result gates a merge, run it headless, where auto mode does not apply. With --agent, the whole session takes on the subagent’s tool list, and dontAsk denies every call that would otherwise prompt; only reads, the built-in read-only command set and your allow rules run. Put the allow rules in .claude/settings.json:

{
"permissions": {
"allow": ["Bash(npm test *)", "Bash(npm run lint *)", "Bash(git diff *)"]
}
}

Replace the three commands with the checks your spec.md lists. The headless command is in How does the parent check a subagent’s result?, with the schema it uses.

For a writer, add isolation: worktree. By default the worktree branches from your repository’s default branch on the remote, not from the parent session’s HEAD, so a writer sees neither your local commits nor your uncommitted work. Set "worktree": { "baseRef": "head" } in .claude/settings.json so subagent worktrees branch from your current local HEAD (which carries unpushed commits but not uncommitted changes), and commit the work first. The setting does not accept a branch name. The worktree is removed automatically if the subagent makes no changes.

Keep one role contract in your docs and one adapter file per tool. In Claude Code, claude import codex --dry-run (or /import codex in a session) previews what would be brought across; the same works for cursor. Treat the imported subagents as drafts: the boundaries live in different fields in each tool, so re-check every imported role.

Prove the boundaries hold before you rely on them

Section titled “Prove the boundaries hold before you rely on them”

A definition is a claim. Test it the way you would test a permission rule.

  1. Ask the read-only role to write. Delegate: “Add a comment to the first line of README.md.” The explorer or verifier must refuse, or its tool call must be denied. If the file changes, the boundary is prompt-only. Tighten tools, the deny rules or the session profile, then repeat.

  2. Ask the verifier to run an unlisted command. Delegate: “Run curl https://example.com.” In a headless Claude Code run with dontAsk, the call must be denied because no allow rule matches it. In Codex, run the same request inside the read-only codex exec and confirm the command is refused or fails.

  3. Feed the verifier a known-bad diff. Break one acceptance criterion deliberately on a scratch branch. A verifier that returns all PASS on it is not verifying. This is the single most useful test on the page. Rerun it when you change the model or the role prompt, alongside your continuous evals.

  4. Commit the definitions and the test. Keep the role files and the verdict schema under version control, with a short agents/README.md that records the date and tool version each boundary was last tested.

How does the parent check a subagent’s result?

Section titled “How does the parent check a subagent’s result?”

Nobody should re-read the whole diff to trust a verifier. Make the output machine-checkable and let a gate read it.

Keep one schema for both tools at agents/verdict.schema.json. Codex reads it through codex exec --output-schema, and Claude Code through claude -p --json-schema:

{
"type": "object",
"additionalProperties": false,
"required": ["criteria"],
"properties": {
"criteria": {
"type": "array",
"items": {
"type": "object",
"additionalProperties": false,
"required": ["id", "verdict", "evidence"],
"properties": {
"id": { "type": "string" },
"verdict": { "type": "string", "enum": ["PASS", "BLOCKED"] },
"evidence": { "type": "string" }
}
}
}
}
}

Write the verdict to verdict.json. Codex writes the schema object directly with -o verdict.json. Claude Code wraps it in structured_output, so extract that field:

Terminal window
claude -p --agent spec-verifier --permission-mode dontAsk \
--output-format json --json-schema "$(cat agents/verdict.schema.json)" \
"Verify the diff on this branch against spec.md." \
| jq '.structured_output' > verdict.json

Then gate on it in the terminal or in CI:

Terminal window
jq -e '(.criteria | length > 0) and all(.criteria[]; .verdict == "PASS")' verdict.json

The length check matters: jq’s all returns true for an empty array, so without it a verifier that reports no criteria at all would pass the gate. With it, an empty criteria array, a missing field, an empty file or a missing file all exit non-zero, so the gate fails closed. In CI on a pull request, though, the role file, the allow rules in .claude/settings.json, spec.md, the schema and the workflow itself all come from the pull request, so an author or an agent can loosen the gate in the same change. It is a gate only when .claude/, agents/, spec.md and .github/ sit under CODEOWNERS with a required code-owner review; never run it on pull_request_target, and give the job no secret except the model key.

Three checks keep the parent accountable:

  • Criteria coverage. The verdict lists every acceptance criterion ID in spec.md, with none missing and none invented.
  • Evidence spot-check. The parent opens two cited file:line references at random and confirms they say what the verifier claims. One false citation fails the whole report.
  • Deterministic gates still run. Tests, type checks and lint run in CI regardless of the verdict. The verifier covers what those gates cannot express, such as “the error message names the missing field”.

A human signs off at the repository’s existing risk gates, such as production approval or security-sensitive paths. The verdict is evidence for that decision, not a replacement for it. For writing acceptance criteria a verifier can check, see the design stage of the lifecycle.

Does delegation actually pay off? Measure it

Section titled “Does delegation actually pay off? Measure it”

Subagents trade context for coordination. Protecting the main context can save a long session, but every subagent re-reads files, and parallel runs multiply spend. In one practitioner report of an orchestrated multi-agent setup, an hour cost “about 10X the cost of a normal Claude Code session per unit time” (DoltHub, Tim Sehn, 2026-01-15; one user, one day, one setup). Your number will differ. Measure it.

For three to five representative tasks, run each once in a single session and once with delegation, and record:

MeasureSingle agentDelegated
Elapsed time to an accepted result
Tokens or cost, from the tool’s usage view
Findings the parent accepted after spot-checks
Findings rejected as wrong or unsupported
Integration defects (conflicts, reverted edits)

Keep a role when it wins on accepted findings or elapsed time without adding integration defects. On model choice, start subagents on the session’s model and move a role to a cheaper model only when the same comparison shows no drop in accepted findings. See model routing by evidence and the models hub for current options.

When scoped subagents go wrong, and how to recover

Section titled “When scoped subagents go wrong, and how to recover”
SymptomLikely causeRecovery
The “read-only” reviewer changed filesBoundary set only in the prompt; Codex role in a writable session; Claude Code permissionMode overridden by auto modeRestrict tools (Claude Code) or run the verifier as a separate read-only codex exec; rerun boundary test 1
Verifier returns PASS on everythingIt sees only the author’s summary, or it may “fix” failuresPass spec.md and the raw diff, remove edit tools, rerun boundary test 3 with a known-bad diff
Two subagents overwrote each otherWriters shared one checkoutGive each writer a worktree (isolation: worktree, Codex --worktree, Cursor worktrees or cloud subagents) and name the files each one owns
Writer in a worktree ignores your latest changesClaude Code worktrees branch from the remote default branch, not HEADSet "worktree": { "baseRef": "head" } in .claude/settings.json and commit the work first
Concurrent subagent limit reached in Claude Code20 running subagents by defaultBatch the work; raise CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS only after measuring cost
Delegation chains three layers deep and nobody owns the resultNesting allowed by default (three layers in Claude Code)Remove Agent from verifier and explorer tools; set CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH to 1, or agents.max_depth in Codex
New Codex role never appearsMissing name, description or developer_instructions; the file was skippedRead the startup warnings and fix the named field
The parent reports success without checkingThe verdict was treated as the resultGate on the verdict JSON, then spot-check citations before you merge

Stop delegating a task type when its coordination cost (writing the handoff, checking the output, integrating) exceeds the context or time it saves. Put that finding in the role’s README so the next person does not repeat the experiment.

  • Each role has an objective, inputs, allowed actions, an output schema and a stop condition, in a committed file.
  • Read-only roles refuse the write in boundary test 1, and the verifier catches the known-bad diff in boundary test 3.
  • Concurrent writers run in separate worktrees or cloud environments, with named file ownership.
  • The parent gates on the verifier’s structured verdict and spot-checks its citations before integration.
  • At least one single-agent versus delegated comparison is recorded, and the roles you kept won it.
  • Each failure changed a role definition, a deny rule or a test, not only a chat message.

Back to the Developer Scorecard answer key.