Skip to content

Multi-agent orchestration patterns: planner, workers, judge

Multi-agent orchestration patterns are four ways to split software work across coding agents: fan-out/fan-in, pipelines, best-of-N with a judge, and writer/reviewer separation. Each pays off only when the work splits into independently verifiable pieces, or when a wrong answer costs more than the extra runs; otherwise it multiplies token spend.

You start eight agents on the billing service. By afternoon three branches have rewritten the same routes/index.ts, two agents have invented different shapes for one event, the “reviewer” has approved code from its own context, and a day’s budget is gone. The tool did what you asked. The pattern did not fit the work.

This page is for developers who run agents and tech leads who decide how many run at once and what “done” means.

What a working orchestration setup gives you

Section titled “What a working orchestration setup gives you”
  • A decision table: which pattern a task fits, and which tasks stay with one agent.
  • Planner, worker and judge as committable files: a Claude Code worktree subagent, a read-only judge, a Codex [agents] block.
  • Four copy-paste prompts: decomposition, fan-out audit, best-of-N judging, adversarial review.
  • A verification chain that proves output without a line-by-line read.
  • Each tool’s cost controls and the common failure modes.

Commands and settings on this page were checked against Claude Code 2.1.283 and codex 0.157.1 (--help and the Codex source) on 2026-09-26. Cursor specifics were checked against cursor.com on 2026-08-28.

What are the planner, worker and judge roles?

Section titled “What are the planner, worker and judge roles?”

Every pattern wires three roles differently. Name them in prompts and agent definitions: most failures come from one agent silently playing two.

RoleDoesMust notOutput
PlannerSplits the goal into units, names the files each unit owns, writes the shared contracts, and states each unit’s acceptance checkWrite implementation codeA plan file: units, owned paths, dependencies, the command that proves each unit
WorkerImplements or investigates exactly one unit, in its own worktree or VMTouch files outside its unit, or edit the tests that judge itA branch or a report, plus the output of its acceptance command
JudgeRuns the deterministic checks first, then compares or grades what the workers producedEdit the code it judges, or know which agent or model wrote which candidateA verdict with evidence: commands run, exit codes, file and line

Teams skip the judge. A judge that only reads code is a second opinion; one that runs the reproduction test, type check and lane tests, and grades only what those miss, is a gate. See model-graded checks.

Which orchestration pattern fits your task?

Section titled “Which orchestration pattern fits your task?”

Choose from the shape of the work, not the number of agents you can start.

PatternShapePays off whenOnly multiplies cost whenThe judge checks
Fan-out/fan-inThe same step over many independent items, then one mergeItems own separate files, share no in-flight contracts, and each has its own checkItems share files or a contract that is still changing; the merge step reworks everythingEach item’s own check, then one integration run after the merge
PipelineStages in sequence (spec → plan → implement → verify → review), each with its own tools and contextStages need different permissions or context, and each hands over a file you can validateStages hand over prose; one agent could hold the whole task in contextThe artifact at each stage boundary, against a schema or a checklist
Best-of-N with a judgeN independent attempts at one problem, one judge picksThe solution space is wide, a wrong first answer is expensive, and the judge can decide on evidenceThe change is mechanical, so the attempts converge; or the judge can only give an opinionA reproduction test, a benchmark or a spec check on every candidate
Writer/reviewerOne agent writes, a separate agent with fresh context reviews, the writer fixesAlmost always at pull-request level: it is the cheapest pattern with the clearest payoffThe loop has no stop rule and the pair argues for roundsFindings that carry file, line and a failing command, not style preferences

Two rules cover the rest. A problem one agent can hold in context stays with one agent: orchestration buys breadth, isolation or reliability, never a cheaper hard problem. A task that ends on one pass/fail condition (“the build is green”) is a goal loop; see the /goal command.

Fan-out starts one worker per item; fan-in collects and verifies the results, as in audits, file-by-file migrations and test backfills. It pays off when every item owns its files, depends on nothing another worker is still writing, and can prove itself alone with its own tests, type check and lint. An item that fails any of the three belongs in a serial foundation step before the fan-out: migrations, shared types, barrel-file registrations.

Fan-in is where fan-out loses quality. Require three things:

  • Coverage accounting. The report lists unfinished items; “no findings” from 60% of the files is not “no findings”.
  • Deduplication and refutation. A second pass tries to disprove each finding before it reaches you.
  • One integration run. The full suite runs once on the merged result, because each lane was only proven alone.

In Claude Code, prefix it with ultracode or add “use a workflow” to run a dynamic workflow; in Codex and Cursor the agent delegates to subagents.

When does a pipeline beat one long session?

Section titled “When does a pipeline beat one long session?”

A pipeline runs stages in order, each with only the context and tools it needs: a read-only planner, an implementer in a worktree, a verifier that runs tests but cannot edit. It pays off when those permissions matter, or when one session’s context would fill up.

The handoff is the whole design. Each stage writes a file with a shape you can check: a plan.json against a schema, a branch against its acceptance command. Stages that hand over prose accumulate each other’s mistakes.

Humans sign off on the plan and the merge (see the verification chain below). For the issue-to-pull-request version, see issue to PR with no hands on the keyboard.

When is best-of-N with a judge worth the extra runs?

Section titled “When is best-of-N with a judge worth the extra runs?”

Best-of-N runs N independent attempts at one problem and lets a judge pick. It pays off on a bug with an unclear cause, a performance fix with several plausible approaches, or a design choice, provided a reproduction test or benchmark can tell which attempt worked.

It only multiplies cost on mechanical work (a rename, a codemod, a dependency bump), where attempts converge, or with a judge that has no evidence.

Three rules make the judge worth having:

  • Evidence before taste. The judge runs the reproduction test, the affected suite and any benchmark on every candidate, and grades readability only among those that pass.
  • Blind comparison. Strip the agent name, model and attempt number from the candidates before the judge sees them, and vary their order across runs.
  • No edit rights. The judge may run candidates in throwaway worktrees, but never “fixes up” the winner. If none passes, the verdict is “none”.

Give each attempt a different starting angle (the failing test, recent commits, the data flow at the boundary); identical prompts on one model often produce near-identical attempts. If the diffs still converge, the task is mechanical: use one attempt.

Why separate the writer from the reviewer?

Section titled “Why separate the writer from the reviewer?”

An agent reviewing its own work brings the blind spots that produced the bug. A reviewer with fresh context, read-only tools and a fixed report format catches different errors; a different model or vendor diverges further.

Two constraints stop an endless loop. The reviewer reports only findings backed by a file, a line and a command that shows the problem, so style never blocks. The writer gets at most two fix rounds, then open findings go to a human. Approval never replaces CI.

All three tools run all four patterns with different building blocks: scriptable workflows in Claude Code, subagents plus cloud attempts in Codex, subagents and per-VM cloud agents in Cursor.

PatternBuilding block in Cursor
Fan-out/fan-inSubagents; since the 2026-08-19 release they “can now run on their own virtual machines”. Worktrees for lanes you run yourself.
PipelineAutomations start cloud agents on a schedule or on events (GitHub, Slack, Linear, webhooks and more), so each stage can react to the previous stage’s pull request or label.
Best-of-NThe same prompt on several Cloud Agents, each from a different angle, judged with the judge prompt above; the Cursor SDK scripts this.
Writer/reviewerBugbot reviews pull requests as a separate agent.

Triggers and the API: cloud agents and automations.

Run planner, workers and judge in Claude Code

Section titled “Run planner, workers and judge in Claude Code”

This example wires all three roles for a fan-out of implementation units. Commit the files so the team shares the same workers and judge.

  1. Plan and approve. Run the planner prompt above. Read plan.json, not code: no path in two units, an acceptance command per unit. Land the foundation serially first.

  2. Branch workers from your work, not from main. A subagent with isolation: worktree branches from the default branch and would miss the foundation. Add this to .claude/settings.json:

    .claude/settings.json
    {
    "worktree": {
    "baseRef": "head"
    }
    }
  3. Define the worker. One unit per invocation, its own worktree, and the acceptance command as its stop condition:

    .claude/agents/unit-implementer.md
    ---
    name: unit-implementer
    description: Implements exactly one unit from plan.json in an isolated worktree. Use for each parallel unit after the foundation has landed.
    tools: Read, Grep, Glob, Edit, Write, Bash
    isolation: worktree
    ---
    You implement one unit from plan.json. The unit id is in your task.
    Edit only files matching that unit's owned_paths. Never edit files under tests/ that existed before you started.
    When done, run the unit's acceptance commands and commit.
    Report: unit id, branch name, files changed, each acceptance command with its exit code.
    If an acceptance command fails after three attempts, stop and report the failure instead of widening your scope.
  4. Define the judge. No Edit or Write, so it cannot make its own verdict true:

    .claude/agents/unit-judge.md
    ---
    name: unit-judge
    description: Verifies one finished unit against plan.json. Use after a unit-implementer reports done.
    tools: Read, Grep, Glob, Bash
    ---
    You verify one unit branch. You never edit files.
    1. Run git diff --name-only against the base and fail the unit if any path is outside its owned_paths.
    2. Fail the unit if any pre-existing test file changed.
    3. Run the unit's acceptance commands yourself and record exit codes.
    4. Only then read the diff for correctness issues the tests cannot catch.
    Output: PASS or FAIL, then each check with its evidence (command, exit code, file:line).
  5. Fan out. Ask Claude: “Start one unit-implementer per unit in plan.json whose depends_on is satisfied, then run unit-judge on each finished branch.” Watch progress with /tasks.

  6. Fan in. Merge passing branches in dependency order, run the integration commands once, and open the pull request. CI reruns the gates; the developer who launched the run signs off.

To make the orchestration repeatable, save it as a dynamic workflow. This best-of-three runs three read-only proposals from different angles, then a judge that tests each in a throwaway worktree.

.claude/workflows/best-of-three-fix.js
export const meta = {
name: 'best-of-three-fix',
description: 'Three independent fix proposals for one bug, judged against its reproduction test',
}
const proposal = {
type: 'object',
required: ['root_cause', 'diff', 'evidence'],
properties: {
root_cause: { type: 'string' },
diff: { type: 'string' },
evidence: { type: 'string' },
},
}
phase('Propose')
const angles = [
'start from the failing test and trace inward',
'start from the recent commits that touched this module',
'start from the data flow at the API boundary',
]
const candidates = await pipeline(angles, (angle) =>
agent(
`Bug: ${args.bug}. Reproduction: ${args.repro}. Investigate; ${angle}. ` +
'Do not edit files. Return the root cause, a unified diff that fixes it, and command output that supports it.',
{ label: angle, schema: proposal },
),
)
phase('Judge')
return await agent(
'You are the judge; do not edit the main checkout. For each candidate below, create a throwaway worktree, ' +
`apply its diff, run ${args.repro} and the module's test suite, record exit codes, then remove the worktree. ` +
'Disqualify any candidate that fails or changes a test file. Pick the smallest passing root-cause fix, or none.\n' +
JSON.stringify(candidates.filter(Boolean)),
)

Run it with a prompt such as Run /best-of-three-fix with bug "proration is off by one day at month end" and repro "npx vitest run tests/billing/proration.test.ts". The candidates carry no agent names, so the judge compares them blind. The order is fixed by the angles array; to shuffle, pass a seed or a permutation in through args (workflow scripts cannot call Math.random()).

The built-ins cover single-vendor work. Add an orchestrator for a mixed fleet, a supervision UI, or handoff across tools.

ToolPattern it implementsInstallGitHub stars (2026-09-26)
oh-my-claudecodeFan-out teams inside Claude Code (/team 3:executor "fix all TypeScript errors"); omc team starts Claude, Codex or Gemini workers in tmux/plugin marketplace add https://github.com/Yeachan-Heo/oh-my-claudecode, then /plugin install oh-my-claudecode39,357
Codex plugin for Claude Code (OpenAI)Cross-vendor writer/reviewer: /codex:review, /codex:adversarial-review/plugin marketplace add openai/codex-plugin-cc, then /plugin install codex@openai-codex33,594
Gas TownPlanner and workers: a “Mayor” Claude Code instance coordinates worker agents over the Beads ledgerbrew install gastown (macOS)18,195
Claude SquadFan-out you supervise: one tmux session and worktree per agent, with a diff viewbrew install claude-squad8,533
CLI Agent Orchestrator (AWS Labs)Supervisor-to-worker handoff over tmux, with a small web UIuv tool install git+https://github.com/awslabs/cli-agent-orchestrator.git@main --upgrade1,350

Full catalogue: running many agents at once and the agent tools overview.

How do you verify orchestrated output without reading every diff?

Section titled “How do you verify orchestrated output without reading every diff?”

Build the chain from gates that run without you, so every stage produces evidence:

  1. Lane gate. Each unit’s own acceptance commands pass on its branch, and a path check proves it stayed in its lane. A plain script does the path check without a model, then runs the acceptance command:

    scripts/lane-gate.sh
    #!/usr/bin/env bash
    # Usage: scripts/lane-gate.sh BASE_BRANCH 'src/notifications/|src/api/preferences/' npx vitest run tests/notifications
    set -euo pipefail
    outside=$(git diff --name-only "$1"...HEAD | grep -Ev "^($2)" || true)
    if [ -n "$outside" ]; then
    echo "Lane violation, files outside the lane:"; echo "$outside"; exit 1
    fi
    "${@:3}" # the unit's acceptance command
  2. Judge. A separate, read-only agent reruns the decisive checks; every verdict line carries a command and an exit code.

  3. Integration gate. After fan-in, the full suite, type check and lint run once on the combined branch, catching interactions no lane could see.

  4. CI. The pull request runs the same gates on a clean machine. An agent’s APPROVE never replaces it.

  5. Human sign-off. The developer who launched the run approves the plan before the fan-out and the merge after CI. The tech lead owns .claude/agents/, .claude/workflows/ and the Codex [agents] roles through CODEOWNERS: changing a judge changes what “verified” means.

Attach the plan, verdicts and CI results to the pull request as an evidence bundle. To stop workers weakening the tests that judge them, see protect the oracle.

What does orchestration cost, and how do you cap it?

Section titled “What does orchestration cost, and how do you cap it?”

Cost grows with the number of agents, not the size of the result. One engineer at DoltHub reported that an hour of running Gas Town “cost me about $100 in Claude tokens. That’s about 10X the cost of a normal Claude Code session per unit time” (Tim Sehn, DoltHub blog, 2026-01-15; one user and one setup, not a benchmark). Anthropic’s workflow documentation says a single workflow run “can use meaningfully more tokens than working through the same task in conversation”.

The caps each tool gives you:

ControlClaude Code 2.1.283Codex 0.157.1
Concurrency20 running subagents per session by default (CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS); 16 concurrent workflow agents by default (CLAUDE_CODE_WORKFLOW_MAX_CONCURRENT_AGENTS, 1 to 256)[agents] max_threads
Size of one run1,000 agents per workflow run; a Large workflow warning above 25 agents or a projected 1.5 million tokens; workflowSizeGuideline (/config) sets the target size — default medium (fewer than 10 agents), small on Pro--attempts on codex cloud exec accepts 1 to 4
Spendclaude -p --max-budget-usd for headless runs[goals] max_goal_token_budget; tokens spent by nested subagents count toward it
Model per rolemodel and effort in each subagent filedefault_subagent_reasoning_effort, default_subagent_model

Cursor bills cloud agents at API pricing for the selected model (checked 2026-08-28); cancel each scripted @cursor/sdk run on a timer (driving agents from code).

Start every role on the tool’s default model, tune effort before switching models, and move only mechanical workers to a cheaper model once accepted work shows no drop (models hub). Pilot on one directory and measure cost per accepted change.

What breaks when you orchestrate several agents?

Section titled “What breaks when you orchestrate several agents?”

Conflicts in barrel files, route registration or shared types mean the planner put a shared file in two units. Move it into the foundation, land it, and rebase the remaining lanes; the lane gate prevents a repeat.

Workers reimplement code that already exists on your branch

Section titled “Workers reimplement code that already exists on your branch”

A Claude Code worker with isolation: worktree branched from the default branch, because worktree.baseRef defaults to "fresh". Set "baseRef": "head", or create the worktrees yourself from the right branch.

Candidates pass the judge, then fail CI: the judge reads instead of running checks, shares the writer’s context, or can edit. Remove Edit and Write, require a command and exit code per verdict line, and compare the last five verdicts with CI; when they disagree, the judge prompt is the bug.

Workers were stopped or hit an unrecoverable API error (or a usage limit in a claude -p or background run), and fan-in dropped their empty results (such an agent() call in a Claude Code workflow resolves to null). Require a “not finished” list and rerun only those items.

The run stalls, or writer and reviewer never converge

Section titled “The run stalls, or writer and reviewer never converge”

A background worker waiting on an unapproved tool idles for hours: pre-approve test, lint and read commands; keep destructive ones denied. A writer/reviewer pair on round five needs the stop rule: two fix rounds, then a human.

Where to go next with multi-agent orchestration

Section titled “Where to go next with multi-agent orchestration”

See also multi-agent harnesses and background and cloud agents compared.