Multi-agent orchestration patterns: planner, workers, judge
Multi-agent orchestration patterns are four ways to split software work across coding agents: fan-out/fan-in, pipelines, best-of-N with a judge, and writer/reviewer separation. Each pays off only when the work splits into independently verifiable pieces, or when a wrong answer costs more than the extra runs; otherwise it multiplies token spend.
You start eight agents on the billing service. By afternoon three branches have rewritten the same routes/index.ts, two agents have invented different shapes for one event, the “reviewer” has approved code from its own context, and a day’s budget is gone. The tool did what you asked. The pattern did not fit the work.
This page is for developers who run agents and tech leads who decide how many run at once and what “done” means.
What a working orchestration setup gives you
Section titled “What a working orchestration setup gives you”- A decision table: which pattern a task fits, and which tasks stay with one agent.
- Planner, worker and judge as committable files: a Claude Code worktree subagent, a read-only judge, a Codex
[agents]block. - Four copy-paste prompts: decomposition, fan-out audit, best-of-N judging, adversarial review.
- A verification chain that proves output without a line-by-line read.
- Each tool’s cost controls and the common failure modes.
Commands and settings on this page were checked against Claude Code 2.1.283 and codex 0.157.1 (--help and the Codex source) on 2026-09-26. Cursor specifics were checked against cursor.com on 2026-08-28.
What are the planner, worker and judge roles?
Section titled “What are the planner, worker and judge roles?”Every pattern wires three roles differently. Name them in prompts and agent definitions: most failures come from one agent silently playing two.
| Role | Does | Must not | Output |
|---|---|---|---|
| Planner | Splits the goal into units, names the files each unit owns, writes the shared contracts, and states each unit’s acceptance check | Write implementation code | A plan file: units, owned paths, dependencies, the command that proves each unit |
| Worker | Implements or investigates exactly one unit, in its own worktree or VM | Touch files outside its unit, or edit the tests that judge it | A branch or a report, plus the output of its acceptance command |
| Judge | Runs the deterministic checks first, then compares or grades what the workers produced | Edit the code it judges, or know which agent or model wrote which candidate | A verdict with evidence: commands run, exit codes, file and line |
Teams skip the judge. A judge that only reads code is a second opinion; one that runs the reproduction test, type check and lane tests, and grades only what those miss, is a gate. See model-graded checks.
Which orchestration pattern fits your task?
Section titled “Which orchestration pattern fits your task?”Choose from the shape of the work, not the number of agents you can start.
| Pattern | Shape | Pays off when | Only multiplies cost when | The judge checks |
|---|---|---|---|---|
| Fan-out/fan-in | The same step over many independent items, then one merge | Items own separate files, share no in-flight contracts, and each has its own check | Items share files or a contract that is still changing; the merge step reworks everything | Each item’s own check, then one integration run after the merge |
| Pipeline | Stages in sequence (spec → plan → implement → verify → review), each with its own tools and context | Stages need different permissions or context, and each hands over a file you can validate | Stages hand over prose; one agent could hold the whole task in context | The artifact at each stage boundary, against a schema or a checklist |
| Best-of-N with a judge | N independent attempts at one problem, one judge picks | The solution space is wide, a wrong first answer is expensive, and the judge can decide on evidence | The change is mechanical, so the attempts converge; or the judge can only give an opinion | A reproduction test, a benchmark or a spec check on every candidate |
| Writer/reviewer | One agent writes, a separate agent with fresh context reviews, the writer fixes | Almost always at pull-request level: it is the cheapest pattern with the clearest payoff | The loop has no stop rule and the pair argues for rounds | Findings that carry file, line and a failing command, not style preferences |
Two rules cover the rest. A problem one agent can hold in context stays with one agent: orchestration buys breadth, isolation or reliability, never a cheaper hard problem. A task that ends on one pass/fail condition (“the build is green”) is a goal loop; see the /goal command.
When does fan-out/fan-in pay off?
Section titled “When does fan-out/fan-in pay off?”Fan-out starts one worker per item; fan-in collects and verifies the results, as in audits, file-by-file migrations and test backfills. It pays off when every item owns its files, depends on nothing another worker is still writing, and can prove itself alone with its own tests, type check and lint. An item that fails any of the three belongs in a serial foundation step before the fan-out: migrations, shared types, barrel-file registrations.
Fan-in is where fan-out loses quality. Require three things:
- Coverage accounting. The report lists unfinished items; “no findings” from 60% of the files is not “no findings”.
- Deduplication and refutation. A second pass tries to disprove each finding before it reaches you.
- One integration run. The full suite runs once on the merged result, because each lane was only proven alone.
In Claude Code, prefix it with ultracode or add “use a workflow” to run a dynamic workflow; in Codex and Cursor the agent delegates to subagents.
When does a pipeline beat one long session?
Section titled “When does a pipeline beat one long session?”A pipeline runs stages in order, each with only the context and tools it needs: a read-only planner, an implementer in a worktree, a verifier that runs tests but cannot edit. It pays off when those permissions matter, or when one session’s context would fill up.
The handoff is the whole design. Each stage writes a file with a shape you can check: a plan.json against a schema, a branch against its acceptance command. Stages that hand over prose accumulate each other’s mistakes.
Humans sign off on the plan and the merge (see the verification chain below). For the issue-to-pull-request version, see issue to PR with no hands on the keyboard.
When is best-of-N with a judge worth the extra runs?
Section titled “When is best-of-N with a judge worth the extra runs?”Best-of-N runs N independent attempts at one problem and lets a judge pick. It pays off on a bug with an unclear cause, a performance fix with several plausible approaches, or a design choice, provided a reproduction test or benchmark can tell which attempt worked.
It only multiplies cost on mechanical work (a rename, a codemod, a dependency bump), where attempts converge, or with a judge that has no evidence.
Three rules make the judge worth having:
- Evidence before taste. The judge runs the reproduction test, the affected suite and any benchmark on every candidate, and grades readability only among those that pass.
- Blind comparison. Strip the agent name, model and attempt number from the candidates before the judge sees them, and vary their order across runs.
- No edit rights. The judge may run candidates in throwaway worktrees, but never “fixes up” the winner. If none passes, the verdict is “none”.
Give each attempt a different starting angle (the failing test, recent commits, the data flow at the boundary); identical prompts on one model often produce near-identical attempts. If the diffs still converge, the task is mechanical: use one attempt.
Why separate the writer from the reviewer?
Section titled “Why separate the writer from the reviewer?”An agent reviewing its own work brings the blind spots that produced the bug. A reviewer with fresh context, read-only tools and a fixed report format catches different errors; a different model or vendor diverges further.
Two constraints stop an endless loop. The reviewer reports only findings backed by a file, a line and a command that shows the problem, so style never blocks. The writer gets at most two fix rounds, then open findings go to a human. Approval never replaces CI.
How does each tool run these patterns?
Section titled “How does each tool run these patterns?”All three tools run all four patterns with different building blocks: scriptable workflows in Claude Code, subagents plus cloud attempts in Codex, subagents and per-VM cloud agents in Cursor.
| Pattern | Building block in Cursor |
|---|---|
| Fan-out/fan-in | Subagents; since the 2026-08-19 release they “can now run on their own virtual machines”. Worktrees for lanes you run yourself. |
| Pipeline | Automations start cloud agents on a schedule or on events (GitHub, Slack, Linear, webhooks and more), so each stage can react to the previous stage’s pull request or label. |
| Best-of-N | The same prompt on several Cloud Agents, each from a different angle, judged with the judge prompt above; the Cursor SDK scripts this. |
| Writer/reviewer | Bugbot reviews pull requests as a separate agent. |
Triggers and the API: cloud agents and automations.
| Pattern | Building block in Claude Code 2.1.283 |
|---|---|
| Fan-out/fan-in | Subagents in .claude/agents/. /batch <instruction> splits a change into 5 to 30 units, one worktree subagent each. Dynamic workflows (ultracode or “use a workflow”) fan out dozens to hundreds of agents from a script. |
| Pipeline | Workflow phase() blocks with a JSON schema on each agent() call; or a chain of claude -p --json-schema calls in CI. |
| Best-of-N | A workflow that runs N candidates through pipeline() and a judge agent() afterwards; or several subagents with isolation: worktree. |
| Writer/reviewer | A read-only reviewer subagent; /code-review, or /code-review ultra in the cloud. OpenAI’s Codex plugin adds a second vendor with /codex:adversarial-review. |
Off your machine, claude --cloud "TASK" starts a cloud session per lane. Agent teams are experimental and off by default (CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1). Details: custom subagents and dynamic workflows.
| Pattern | Building block in codex 0.157.1 |
|---|---|
| Fan-out/fan-in | Subagents (on by default), capped with [agents] max_threads. For code lanes, codex --worktree per lane, or git worktree add plus codex exec -C <dir> from a script. |
| Pipeline | A chain of codex exec --output-schema FILE -o out.json calls, each stage reading the previous stage’s file. |
| Best-of-N | codex cloud exec --env ENV_ID --attempts 3 (cloud is experimental), then codex cloud diff and codex cloud apply TASK_ID --attempt N for the winner. |
| Writer/reviewer | codex exec review --base main or /review in a fresh session, or a custom read-only reviewer role in [agents]. |
[agents]max_threads = 4default_subagent_reasoning_effort = "medium"
[agents.reviewer]description = "Read-only reviewer. Reports findings with file and line; never edits files."config_file = "./agents/reviewer.toml"The description does not make the role read-only; the config layer does. config_file resolves relative to config.toml:
sandbox_mode = "read-only"developer_instructions = "Review only. Report each finding with file, line and a command that shows it; never edit files."Lane scripts: Codex multi-agent workflows. Best-of-N attempts: Codex cloud environments.
Run planner, workers and judge in Claude Code
Section titled “Run planner, workers and judge in Claude Code”This example wires all three roles for a fan-out of implementation units. Commit the files so the team shares the same workers and judge.
-
Plan and approve. Run the planner prompt above. Read
plan.json, not code: no path in two units, an acceptance command per unit. Land the foundation serially first. -
Branch workers from your work, not from
main. A subagent withisolation: worktreebranches from the default branch and would miss the foundation. Add this to.claude/settings.json:.claude/settings.json {"worktree": {"baseRef": "head"}} -
Define the worker. One unit per invocation, its own worktree, and the acceptance command as its stop condition:
.claude/agents/unit-implementer.md ---name: unit-implementerdescription: Implements exactly one unit from plan.json in an isolated worktree. Use for each parallel unit after the foundation has landed.tools: Read, Grep, Glob, Edit, Write, Bashisolation: worktree---You implement one unit from plan.json. The unit id is in your task.Edit only files matching that unit's owned_paths. Never edit files under tests/ that existed before you started.When done, run the unit's acceptance commands and commit.Report: unit id, branch name, files changed, each acceptance command with its exit code.If an acceptance command fails after three attempts, stop and report the failure instead of widening your scope. -
Define the judge. No
EditorWrite, so it cannot make its own verdict true:.claude/agents/unit-judge.md ---name: unit-judgedescription: Verifies one finished unit against plan.json. Use after a unit-implementer reports done.tools: Read, Grep, Glob, Bash---You verify one unit branch. You never edit files.1. Run git diff --name-only against the base and fail the unit if any path is outside its owned_paths.2. Fail the unit if any pre-existing test file changed.3. Run the unit's acceptance commands yourself and record exit codes.4. Only then read the diff for correctness issues the tests cannot catch.Output: PASS or FAIL, then each check with its evidence (command, exit code, file:line). -
Fan out. Ask Claude: “Start one unit-implementer per unit in plan.json whose depends_on is satisfied, then run unit-judge on each finished branch.” Watch progress with
/tasks. -
Fan in. Merge passing branches in dependency order, run the
integrationcommands once, and open the pull request. CI reruns the gates; the developer who launched the run signs off.
To make the orchestration repeatable, save it as a dynamic workflow. This best-of-three runs three read-only proposals from different angles, then a judge that tests each in a throwaway worktree.
export const meta = { name: 'best-of-three-fix', description: 'Three independent fix proposals for one bug, judged against its reproduction test',}
const proposal = { type: 'object', required: ['root_cause', 'diff', 'evidence'], properties: { root_cause: { type: 'string' }, diff: { type: 'string' }, evidence: { type: 'string' }, },}
phase('Propose')const angles = [ 'start from the failing test and trace inward', 'start from the recent commits that touched this module', 'start from the data flow at the API boundary',]const candidates = await pipeline(angles, (angle) => agent( `Bug: ${args.bug}. Reproduction: ${args.repro}. Investigate; ${angle}. ` + 'Do not edit files. Return the root cause, a unified diff that fixes it, and command output that supports it.', { label: angle, schema: proposal }, ),)
phase('Judge')return await agent( 'You are the judge; do not edit the main checkout. For each candidate below, create a throwaway worktree, ' + `apply its diff, run ${args.repro} and the module's test suite, record exit codes, then remove the worktree. ` + 'Disqualify any candidate that fails or changes a test file. Pick the smallest passing root-cause fix, or none.\n' + JSON.stringify(candidates.filter(Boolean)),)Run it with a prompt such as Run /best-of-three-fix with bug "proration is off by one day at month end" and repro "npx vitest run tests/billing/proration.test.ts". The candidates carry no agent names, so the judge compares them blind. The order is fixed by the angles array; to shuffle, pass a seed or a permutation in through args (workflow scripts cannot call Math.random()).
Do you need a third-party orchestrator?
Section titled “Do you need a third-party orchestrator?”The built-ins cover single-vendor work. Add an orchestrator for a mixed fleet, a supervision UI, or handoff across tools.
| Tool | Pattern it implements | Install | GitHub stars (2026-09-26) |
|---|---|---|---|
| oh-my-claudecode | Fan-out teams inside Claude Code (/team 3:executor "fix all TypeScript errors"); omc team starts Claude, Codex or Gemini workers in tmux | /plugin marketplace add https://github.com/Yeachan-Heo/oh-my-claudecode, then /plugin install oh-my-claudecode | 39,357 |
| Codex plugin for Claude Code (OpenAI) | Cross-vendor writer/reviewer: /codex:review, /codex:adversarial-review | /plugin marketplace add openai/codex-plugin-cc, then /plugin install codex@openai-codex | 33,594 |
| Gas Town | Planner and workers: a “Mayor” Claude Code instance coordinates worker agents over the Beads ledger | brew install gastown (macOS) | 18,195 |
| Claude Squad | Fan-out you supervise: one tmux session and worktree per agent, with a diff view | brew install claude-squad | 8,533 |
| CLI Agent Orchestrator (AWS Labs) | Supervisor-to-worker handoff over tmux, with a small web UI | uv tool install git+https://github.com/awslabs/cli-agent-orchestrator.git@main --upgrade | 1,350 |
Full catalogue: running many agents at once and the agent tools overview.
How do you verify orchestrated output without reading every diff?
Section titled “How do you verify orchestrated output without reading every diff?”Build the chain from gates that run without you, so every stage produces evidence:
-
Lane gate. Each unit’s own acceptance commands pass on its branch, and a path check proves it stayed in its lane. A plain script does the path check without a model, then runs the acceptance command:
scripts/lane-gate.sh #!/usr/bin/env bash# Usage: scripts/lane-gate.sh BASE_BRANCH 'src/notifications/|src/api/preferences/' npx vitest run tests/notificationsset -euo pipefailoutside=$(git diff --name-only "$1"...HEAD | grep -Ev "^($2)" || true)if [ -n "$outside" ]; thenecho "Lane violation, files outside the lane:"; echo "$outside"; exit 1fi"${@:3}" # the unit's acceptance command -
Judge. A separate, read-only agent reruns the decisive checks; every verdict line carries a command and an exit code.
-
Integration gate. After fan-in, the full suite, type check and lint run once on the combined branch, catching interactions no lane could see.
-
CI. The pull request runs the same gates on a clean machine. An agent’s
APPROVEnever replaces it. -
Human sign-off. The developer who launched the run approves the plan before the fan-out and the merge after CI. The tech lead owns
.claude/agents/,.claude/workflows/and the Codex[agents]roles through CODEOWNERS: changing a judge changes what “verified” means.
Attach the plan, verdicts and CI results to the pull request as an evidence bundle. To stop workers weakening the tests that judge them, see protect the oracle.
What does orchestration cost, and how do you cap it?
Section titled “What does orchestration cost, and how do you cap it?”Cost grows with the number of agents, not the size of the result. One engineer at DoltHub reported that an hour of running Gas Town “cost me about $100 in Claude tokens. That’s about 10X the cost of a normal Claude Code session per unit time” (Tim Sehn, DoltHub blog, 2026-01-15; one user and one setup, not a benchmark). Anthropic’s workflow documentation says a single workflow run “can use meaningfully more tokens than working through the same task in conversation”.
The caps each tool gives you:
| Control | Claude Code 2.1.283 | Codex 0.157.1 |
|---|---|---|
| Concurrency | 20 running subagents per session by default (CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS); 16 concurrent workflow agents by default (CLAUDE_CODE_WORKFLOW_MAX_CONCURRENT_AGENTS, 1 to 256) | [agents] max_threads |
| Size of one run | 1,000 agents per workflow run; a Large workflow warning above 25 agents or a projected 1.5 million tokens; workflowSizeGuideline (/config) sets the target size — default medium (fewer than 10 agents), small on Pro | --attempts on codex cloud exec accepts 1 to 4 |
| Spend | claude -p --max-budget-usd for headless runs | [goals] max_goal_token_budget; tokens spent by nested subagents count toward it |
| Model per role | model and effort in each subagent file | default_subagent_reasoning_effort, default_subagent_model |
Cursor bills cloud agents at API pricing for the selected model (checked 2026-08-28); cancel each scripted @cursor/sdk run on a timer (driving agents from code).
Start every role on the tool’s default model, tune effort before switching models, and move only mechanical workers to a cheaper model once accepted work shows no drop (models hub). Pilot on one directory and measure cost per accepted change.
What breaks when you orchestrate several agents?
Section titled “What breaks when you orchestrate several agents?”Workers collide on the same file
Section titled “Workers collide on the same file”Conflicts in barrel files, route registration or shared types mean the planner put a shared file in two units. Move it into the foundation, land it, and rebase the remaining lanes; the lane gate prevents a repeat.
Workers reimplement code that already exists on your branch
Section titled “Workers reimplement code that already exists on your branch”A Claude Code worker with isolation: worktree branched from the default branch, because worktree.baseRef defaults to "fresh". Set "baseRef": "head", or create the worktrees yourself from the right branch.
The judge approves everything
Section titled “The judge approves everything”Candidates pass the judge, then fail CI: the judge reads instead of running checks, shares the writer’s context, or can edit. Remove Edit and Write, require a command and exit code per verdict line, and compare the last five verdicts with CI; when they disagree, the judge prompt is the bug.
A clean report hides missing coverage
Section titled “A clean report hides missing coverage”Workers were stopped or hit an unrecoverable API error (or a usage limit in a claude -p or background run), and fan-in dropped their empty results (such an agent() call in a Claude Code workflow resolves to null). Require a “not finished” list and rerun only those items.
The run stalls, or writer and reviewer never converge
Section titled “The run stalls, or writer and reviewer never converge”A background worker waiting on an unapproved tool idles for hours: pre-approve test, lint and read commands; keep destructive ones denied. A writer/reviewer pair on round five needs the stop rule: two fix rounds, then a human.
Where to go next with multi-agent orchestration
Section titled “Where to go next with multi-agent orchestration”See also multi-agent harnesses and background and cloud agents compared.