Skip to content

Team parallelism at Level 3 — isolate work and cap the queue

Team parallelism at Level 3 of the autonomy ladder means several agent tasks run at once without sharing files, ports or data, under a concurrency cap set per task class from measured review and CI capacity, with a named owner for each merge. Parallelism pays only when accepted throughput beats a serial baseline and rework does not rise.

Your engineers already run two or three agents each. Pull requests arrive faster than ever, yet the release is no closer: two branches rewrote the same pricing module, a test run failed because another agent’s dev server held the port, and nobody knows who should resolve the conflict. That is fan-out without a team workflow, and this page is for the tech lead or CTO who has to turn it into one.

The mechanics of one developer running worktrees live in Run parallel agents in isolated worktrees. The review-side arithmetic lives in Running the review queue. This page covers the team standard between them: which work may run in parallel, how state is isolated, who integrates, and how you prove it pays.

What a Level 3 parallel-work standard gives your team

Section titled “What a Level 3 parallel-work standard gives your team”
  • A decision table for which task classes run in parallel and which stay serial.
  • A task-card template every parallel task carries before an agent starts.
  • An isolation script that refuses work when the task’s class is full and otherwise gives each task its own worktree, branch, ports and database name.
  • A per-class concurrency cap and the rule that changes it.
  • An overlap check that runs before merge, and an integration owner who acts on it.
  • Five metric definitions and a two-week experiment that compares parallel work with a serial baseline.
  • A one-page team policy you can commit to the repository today.

How the CTO Scorecard scores team parallelism

Section titled “How the CTO Scorecard scores team parallelism”

Question 14 of the CTO Scorecard asks whether the team runs parallel agent tasks in isolated worktrees. The four answers map to the evidence below. Only the last one earns full credit, and every item in it must be something you can show, not something you intend.

AnswerPointsWhat you can observeNext move
Nobody0One agent session per engineer, one checkoutPilot one task class with two isolated tasks
Individual developers experimenting1Private worktree habits, no shared conventionsStandardise the isolation script and branch prefix
Recommended or paid for seniors2Tooling is funded, but no cap, owner or measurementSet per-class caps and name integration owners
Team workflow with isolated state, measured concurrency caps, integration ownership and reviewer-queue limits3The policy file, the dispatch gate, the overlap check and the weekly metricsHold the cap where accepted throughput is highest, not where it peaks

Why more agents can lower accepted throughput

Section titled “Why more agents can lower accepted throughput”

Generation scales with agents; review and integration do not. Faros AI’s AI Engineering Report 2026: The Acceleration Whiplash (April 2026; vendor telemetry from 22,000 developers on its own platform) measured both sides in the same period: task throughput per developer rose 33.7%, while median time in review rose 441.5% and incidents per pull request rose 242.7%.

That is the Level 3 ceiling: the queue, not the agent, sets the pace. So the unit you manage is accepted work, meaning merged, green and not reverted, rather than pull requests opened or agents running.

Parallelism is a property of the work, not of the tool. Two tasks can run at once only when neither changes something the other reads. Use this table to classify each task before dispatch, and add your own classes as you learn.

Task classDefaultWhy
Vertical slice in its own module or packageParallelOwned paths do not overlap; conflicts are rare and textual
Tests or characterisation tests for existing codeParallelAdds files; the oracle is reviewed once, not per task
Isolated bug fix with a failing testParallelSmall diff, clear acceptance check
Change to a shared abstraction, base class or API contractSerialEvery parallel task reads it, so every task is invalidated by it
Database schema migrationSerialMigrations are ordered; two in flight produce a numbering conflict or a broken chain
Dependency or framework upgradeSerialTouches lockfiles and build config that every branch shares
Cross-cutting rename or formatting passSerial, aloneConflicts with everything; run it when nothing else is in flight

A serial task is not forbidden. It runs alone in its lane, and parallel tasks that depend on it wait until it merges.

An agent that does not know its boundaries drifts into a neighbour’s files. The task card is the contract: the agent-ready backlog shapes the issue, and the card adds what parallel work needs. Commit the card to tasks/<slug>.yaml on the default branch before dispatch; the scripts read it from there.

tasks/billing-retry.yaml
slug: billing-retry
class: vertical-slice # a row from the task-class table
accepted_plan: docs/plans/billing-retry.md
owned_paths:
- src/billing/retry/**
- tests/billing/retry/**
read_only_paths:
- src/billing/client.ts # shared contract: may read, must not edit
depends_on: [] # slugs that must merge first
checks:
- npm run typecheck
- npm test -- tests/billing/retry
stop_when:
- a change outside owned_paths is needed
- a check fails three times in a row
integration_owner: "@ana"

The stop_when list matters as much as the scope. An agent that hits a shared contract should stop and report, not edit the contract and invalidate three other branches.

Replace PLAN_FILE with the plan’s file name. You review the graph and the overlap table, not the code: if the overlap table is not empty, the batch is not ready.

A Git worktree gives each task its own files and branch while sharing one object database. It does not isolate ports, databases, caches, containers or a running dev server. Two agents with separate worktrees and one shared database still produce two green runs that disagree.

Give every task the full set in one command, so nobody assembles it by hand:

scripts/agent-task.sh
#!/usr/bin/env bash
# Usage: scripts/agent-task.sh SLUG INDEX
# Refuses when the task's class is at its cap. Otherwise creates one worktree,
# one branch, one port block and one labelled draft pull request per task.
set -euo pipefail
slug="$1"
index="$2"
[ "$index" -ge 1 ] || { echo "INDEX must be >= 1; 0 is the main checkout's block"; exit 1; }
# Dispatch gate: class from the task card, cap from tasks/caps.txt.
class=$(yq -r '.class' "tasks/$slug.yaml")
cap=$(awk -v c="$class" '$1 == c { print $2 }' tasks/caps.txt)
[ -n "$cap" ] || { echo "no cap for class $class in tasks/caps.txt"; exit 1; }
open=$(gh pr list --state open --label "class:$class" --limit 200 --json number --jq length)
[ "$open" -lt "$cap" ] || { echo "class $class at cap $cap"; exit 1; }
dir="../$(basename "$PWD")-$slug"
git fetch --quiet origin
git worktree add --quiet -b "agent/$slug" "$dir" origin/main
cat > "$dir/.agent-env" <<ENV
APP_PORT=$((3000 + index * 10))
DEBUG_PORT=$((9229 + index * 10))
DATABASE_NAME=app_${slug//-/_}
ENV
# Take the slot: from now on this draft pull request counts against the cap.
git -C "$dir" commit --quiet --allow-empty -m "Start agent/$slug"
git -C "$dir" push --quiet -u origin "agent/$slug"
gh pr create --draft --head "agent/$slug" --label "class:$class" \
--title "agent: $slug" --body "Task card: tasks/$slug.yaml"
echo "Worktree $dir on agent/$slug"
cat "$dir/.agent-env"

The caps live in tasks/caps.txt, one class and its cap per line, for example vertical-slice 3 and migration 1, and the policy below points to that file rather than repeating the numbers. Create one class:<name> label per class once, with gh label create class:vertical-slice, because gh pr create --label fails on a label that does not exist. The script needs gh, git and yq v4 (mikefarah/yq; -r also works with the Python yq).

Run it from the main checkout, for example scripts/agent-task.sh billing-retry 2. Start INDEX at 1: index 0 would give the task the main checkout’s default ports, which is the collision the script exists to prevent. Four details are deliberate. The draft pull request opens at dispatch, not when the agent finishes, so work in progress counts against the cap and never hides as a branch nobody sees. The agent/ branch prefix lets every later script find parallel work. The port stride of 10 keeps a tool that falls back to “the next free port” inside its own block. Add .agent-env to .gitignore, and have your dev scripts read it, so the ports apply without anyone remembering an export. Replace origin/main if your default branch has another name.

The gate is advisory under concurrent dispatch: two people who dispatch in the same minute both see a free slot, and the class goes over its cap. The overlap check and the cap review catch that. For a hard limit, dispatch from one workflow that runs in a single GitHub Actions concurrency: group.

Each tool can also create the worktree itself. The isolation of ports and data is still yours in all three.

Checked against Claude Code 2.1.283 (claude --help):

  • claude --worktree billing-retry (short form -w) starts a session in a new Git worktree. Add --tmux to open that worktree in its own tmux session; --tmux requires --worktree.
  • claude --bg starts a session in the background and prints its ID; claude agents opens agent view (research preview), which lists background sessions and shows which ones need input.
  • The bundled /batch skill splits one large change into 5–30 units, each in its own worktree. Treat each unit as a task: give it owned paths as its card would, and count it against its class cap.
  • claude rm <id> deletes a background session and its worktree when that is safe, which covers most cleanup.

Clean up on a schedule, not when the disk fills. git worktree list shows every checkout; git worktree remove <dir> refuses when the worktree has uncommitted or untracked files, and git branch -d agent/<slug> refuses when the branch is not merged. Both refusals are the safety you want: never add --force to a cleanup script. Run git worktree prune afterwards to drop records for directories that were deleted by hand. For supervising many sessions at once, see running many agents at once.

A cap per engineer (“three agents each”) ignores the only constraint that matters: how fast the team can verify and integrate. Derive the cap from the queue instead.

  1. Compute the review WIP cap. Use the method in Running the review queue: measured review throughput multiplied by the time in review you accept. That number is the ceiling for all open agent pull requests together.
  2. Split the ceiling across task classes. Give each parallel class a share, and give every serial class a cap of one. Start the parallel classes low; the experiment below tells you when to raise them.
  3. Gate dispatch, not pull request creation. scripts/agent-task.sh above is the gate: it counts open pull requests labelled class:<name>, drafts included, and refuses when the class is at its cap in tasks/caps.txt. Gating at pull request creation instead would leave the agent’s finished work on a branch nobody sees. If agents also start from a workflow or a scheduler, run the same gate lines there first.
  4. Check CI capacity separately. If the median CI run for a batch waits in the queue longer than it runs, CI is the constraint, and the cap comes down until it is not.
  5. Review the caps every two weeks. They move when throughput moves. Lower a class’s cap the week its rework rate rises, without waiting for the review.

Assign an integration owner and a merge order

Section titled “Assign an integration owner and a merge order”

Textual conflicts are the easy case; Git reports them. Semantic conflicts are the dangerous ones: two branches that each pass their own checks and break when combined, such as one renaming a field while another adds a caller of the old name. The integration owner exists for those.

The integration owner for a batch decides the merge order from the task cards’ depends_on, rebases or asks the agent to rebase each branch onto the latest default branch before merge, and runs the full test suite on the combined result, not only the task’s own checks. If your forge offers a merge queue, use it: it tests each pull request against the ones merged ahead of it, which is exactly the semantic-conflict check. The owner signs off on the combined result, not on each diff.

Run an overlap check before merge, so conflicts surface while they are still a decision rather than a surprise:

scripts/agent-overlap.sh
#!/usr/bin/env bash
# Print every file that more than one open agent/* branch changes.
set -euo pipefail
git fetch --quiet origin
for head in $(gh pr list --state open --limit 200 --json headRefName \
--jq '.[].headRefName' | grep '^agent/'); do
git diff --name-only "origin/main...origin/$head" | sed "s|\$| origin/$head|"
done | sort | awk '{ n[$1]++; b[$1] = b[$1] " " $2 }
END { for (f in n) if (n[f] > 1) print f ":" b[f] }'

The loop reads open pull requests, not remote branches, so a merged branch that nobody deleted does not show up as an overlap. Empty output means no two open branches touch the same file. A line such as src/billing/client.ts: origin/agent/billing-retry origin/agent/invoice-pdf means a task broke its card; the owner decides which branch keeps the change and which one rebases.

Replace BRANCH_A and BRANCH_B with the two slugs. The output you check is the list in step 2 and a green full suite, not the merged diff line by line.

Prove that parallelism pays against a serial baseline

Section titled “Prove that parallelism pays against a serial baseline”

A team that counts agents running or pull requests opened will always find that more is better. Measure what reaches production instead, and compare it with the same team working one task at a time.

MetricDefinitionSource
Accepted throughputAgent pull requests merged per engineer-week that pass CI and are not reverted or reopened within 14 daysgh pr list --state merged plus revert search
Rework rateShare of merged agent pull requests that need a follow-up fix within 14 daysa rework label the integration owner applies
Conflict rateShare of agent pull requests that needed a conflict resolution before mergea conflict label, or the overlap check log
Time in reviewReady for review to merged, median and 90th percentile; agent pull requests start as drafts at dispatchthe ready_for_review event in each pull request’s timeline; see Running the review queue
Abandoned workagent/ branches or worktrees with no commit in seven daysgit for-each-ref --sort=-committerdate refs/remotes/origin/agent/

Run the comparison as a two-week experiment on one task class. In week one, each engineer runs one agent task at a time. In week two, the class runs at its proposed cap. Keep the task mix and the reviewers the same. Raise the cap only if accepted throughput per engineer rose and rework rate, conflict rate and time in review stayed within your targets. If throughput rose and rework rose with it, the extra output was not accepted work; hold or lower the cap.

Replace START_DATE with the first day of the window in YYYY-MM-DD form and BASELINE with the week-one table. The recommendation is a proposal; the tech lead sets the cap.

Commit the policy next to the scripts so agents and people read the same rules. Adjust the numbers to your measurements; the structure is what earns the Q14 evidence.

docs/parallel-work.md
# Parallel agent work policy
Owner: TECH_LEAD · Reviewed every two weeks · Last change: DATE
## Eligibility
- Only task classes marked Parallel in the task-class table run concurrently.
- Serial classes (shared contract, migration, dependency upgrade, cross-cutting)
run one at a time, and dependent tasks wait for them to merge.
- No dispatch without a task card: owned paths, checks, stop conditions,
integration owner.
## Isolation
- Every task starts from scripts/agent-task.sh: own worktree, agent/<slug>
branch, port block, database name. No shared dev server or database.
- Secrets come from the approved local config, never copied from a teammate.
## Caps
- Team review WIP cap: N open agent PRs (derived in the review-queue sheet).
- Per-class caps live in tasks/caps.txt, one "class cap" pair per line;
every serial class has a cap of 1.
- scripts/agent-task.sh is the dispatch gate: it refuses new work when a class
is at its cap, and every agent pull request carries a class:<name> label.
## Integration
- Each batch names an integration owner. The owner sets merge order, runs the
full suite on the combined result, and signs off on it.
- scripts/agent-overlap.sh runs before every merge; any overlap blocks it.
## Measurement and cleanup
- Weekly: accepted throughput, rework, conflicts, time in review, abandoned work.
- Worktrees and agent/* branches idle for 7 days are reviewed and removed
without --force.

Replace TECH_LEAD, DATE and N with your own, and put the per-class numbers in tasks/caps.txt. The five sections map one to one to the Q14 full-credit answer, which makes the scorecard evidence a link to this file.

What breaks when a team runs agents in parallel

Section titled “What breaks when a team runs agents in parallel”

Pull request count rises and releases do not. Unfinished work piles up in review while dashboards look healthy. Recovery: stop dispatch for the class, drain the queue to below its cap, then restart at the last cap where accepted throughput was highest.

Two green branches break the build together. A semantic conflict passed each task’s own checks. Recovery: revert the later merge, add the test that exposes the conflict, and require the full suite on the combined result through the integration owner or a merge queue.

Tests pass or fail depending on who else is running. Worktrees isolated the files but not the port, database or cache. Recovery: move every task onto the isolation script, and make dev and test commands read .agent-env rather than defaults.

An agent edits a shared contract to finish its task. The card’s read_only_paths were advice, not a gate. Recovery: revert the contract change, turn it into a serial task, and run a scope check in CI on every pull request from an agent/ branch:

scripts/agent-scope.sh
#!/usr/bin/env bash
# Fail when an agent/<slug> branch changes a file outside its card's owned_paths.
# Fails closed: a missing card, empty owned_paths or a failed diff is an error.
set -euo pipefail
slug="${1:-${GITHUB_HEAD_REF:?pass SLUG or run on a pull request}}"
slug="${slug#agent/}"
head="${2:-HEAD}"
card="$(git show "origin/main:tasks/$slug.yaml")" || { echo "no card tasks/$slug.yaml on origin/main"; exit 1; }
mapfile -t globs < <(yq -r '.owned_paths[]' <<<"$card")
(( ${#globs[@]} > 0 )) || { echo "card has no owned_paths"; exit 1; }
changed="$(git diff --name-only "origin/main...$head")" || { echo "cannot diff origin/main...$head"; exit 1; }
status=0
while read -r file; do
[[ -z $file ]] && continue
[[ $file == tasks/* ]] && { echo "task card edited: $file"; status=1; continue; }
for glob in "${globs[@]}"; do [[ $file == $glob ]] && continue 2; done
echo "outside owned_paths: $file"; status=1
done <<<"$changed"
exit "$status"

The script reads the card from origin/main, not from the branch, and fails on any change under tasks/, so an agent cannot widen its own scope by editing its card. It also fails closed: a missing card, an empty owned_paths or a diff that cannot be computed is an error, not a pass. The script alone is not tamper-proof, though. On the pull_request trigger GitHub runs the workflow file from the pull request itself, so a branch in the same repository can edit the job to skip the step or run its own copy of the script, and the check still goes green. The real control is review: a CODEOWNERS entry that assigns .github/, scripts/agent-scope.sh and tasks/ to the integration owners, plus a branch protection rule (or ruleset) on main that requires code-owner review and makes this check required. Then a branch that touches the gate cannot merge without a person who owns it. On pull_request, give the job permissions: contents: read, no secrets and persist-credentials: false on checkout. The alternative is pull_request_target, which always runs the default branch’s workflow. That job runs with the base repository’s token, so keep it at contents: read, give it no secrets, check out only main with persist-credentials: false, and fetch the pull request head as data without checking it out or running anything from it:

.github/workflows/agent-scope.yml
on:
pull_request_target:
branches: [main]
permissions:
contents: read
jobs:
scope:
if: startsWith(github.head_ref, 'agent/')
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
with:
ref: main
fetch-depth: 0
persist-credentials: false
- name: Fetch the pull request head as data
env:
PR_NUMBER: ${{ github.event.pull_request.number }}
GH_TOKEN: ${{ github.token }}
run: |
auth="$(printf 'x-access-token:%s' "$GH_TOKEN" | base64 -w0)"
echo "::add-mask::$auth"
git -c http.extraheader="AUTHORIZATION: basic $auth" fetch origin "refs/pull/$PR_NUMBER/head"
- run: bash scripts/agent-scope.sh "$GITHUB_HEAD_REF" FETCH_HEAD

The fetch uses refs/pull/<number>/head, which exists for fork pull requests too, where the branch name is not on origin. The read-only token reaches only that one git fetch through -c http.extraheader and is never written to .git/config; a public repository can drop the header. The script needs bash 4 or later for mapfile (CI runners have it; on macOS install bash with Homebrew) and yq v4. Bash pattern matching lets * cross a /, so src/billing/retry/** covers the whole directory tree. Fetch enough history for the three-dot diff to find a merge base (in GitHub Actions, fetch-depth: 0).

Worktrees and branches accumulate. Disks fill, and stale branches clutter every agent/* listing and the idle report. Recovery: run the seven-day idle report weekly, remove what is merged or abandoned with the non-forcing commands, and record who approved each deletion.

Engineers supervise more sessions than they can follow. Review turns into approval. Recovery: lower per-engineer session counts and rotate review duty, as described in the sustainable-pace section of Running the review queue.

Parallelism is stable when the caps hold for a month without a rising rework rate. The next step on the ladder is work that runs without a person watching the session, which only makes sense once this page’s gates are in place.