Skip to content

How teams, roles and headcount change shape

Engineering org design for agentic engineering sizes teams by verification capacity rather than by how much code people can write. Agents raise the rate at which changes arrive, so review bandwidth, eval coverage and accountable owners become the binding constraints. Headcount decisions follow those measured constraints, not a vendor’s productivity multiplier.

This page is for executives and CTOs facing a headcount question they cannot yet answer: the board has read that agents write most of the code at some companies, while your own teams merge more pull requests than ever and wait longer for review.

  • A review capacity worksheet that turns arrival rate and risk mix into reviewer hours.
  • Role charters for harness engineer, eval owner and factory operator, staffed at 8, 40 and 200 engineers.
  • A headcount decision table that maps measured signals to decisions, with no multiplier in it.
  • A freed-capacity policy you can adopt as written.
  • The metric definitions that tell you whether the new shape works, and who signs off on each.

Generation scales faster than verification. Faros AI’s telemetry from 22,000 developers and more than 4,000 teams (April 2026) shows epics completed per developer +66.2% and task throughput per developer +33.7%, alongside median time in review +441.5%, PRs merged without any review +31.3%, and incidents per pull request +242.7%. The work grows, and verifying it falls to reviewers.

The organization’s foundation decides who gains. DORA’s 2025 report (Google Cloud, 2025-09-23) links AI adoption to higher delivery throughput and lower stability, and explains: “Teams working in loosely coupled architectures with fast feedback loops see gains, while those constrained by tightly coupled systems and slow processes see little or no benefit.”

The job shifts toward review, and the reviewer stays. In Anthropic’s internal study of 132 engineers and researchers (2025-12-02, a vendor’s own self-report), some engineers describe their work shifting “70%+ to being a code reviewer/reviser rather than a net-new code writer”. Stripe merges more than 1,300 pull requests a week that are “completely minion-produced, human-reviewed, but containing no human-written code” (Stripe engineering blog, 2026-02-19): removing the writing did not remove the reviewer.

When generation is cheap, the scarce resources are the people who can decide whether a change is correct, the oracles that let them decide without reading every line, and the owners who answer for the result.

Size teams by review capacity, not by output

Section titled “Size teams by review capacity, not by output”

With agents, a team’s capacity is the number of changes it can verify and own, not the number it can write. Plan it in three steps.

  1. Measure the arrival rate. Pull 8–12 weeks of merged pull requests from your forge. The command below runs in a terminal with the GitHub CLI and writes one row per pull request: number, author, opened, first human review, merged, lines changed, distinct human reviewers and their logins.

    Terminal window
    gh pr list --state merged --limit 1000 \
    --search "merged:>=2026-07-01" \
    --json number,author,createdAt,mergedAt,additions,deletions,reviews \
    --jq '.[] | . as $pr
    | [.reviews[] | select((.author.login // "ghost") != $pr.author.login
    and ((.author.login // "ghost") | test("\\[bot\\]$|^app/|copilot|codex|claude|bugbot"; "i") | not))] as $h
    | [.number, .author.login, .createdAt,
    ([$h[].submittedAt] | min), .mergedAt,
    (.additions + .deletions),
    ([$h[] | .author.login // "ghost"] | unique | length),
    ([$h[] | .author.login // "ghost"] | unique | join(","))] | @tsv' \
    > review-load.tsv

    The filter drops the author’s own reviews and review bots, because a bot review or a self-review is not a human review. Check the bot logins in your own export and add any the pattern misses; the pattern also drops any human login containing these words, so check the list. An empty first-review column with zero reviewers is then a pull request that merged without human review. Count those first.

  2. Price each change by risk class. Use the four risk classes (low, medium, high, critical) from governance and autonomy and agree on the minutes a competent reviewer needs per class, including reading the evidence bundle. Time five real reviews per class rather than estimating.

  3. Compare demand with protected supply. Supply is the hours per week each reviewer actually protects for review, restricted to the people allowed to approve each class. High and critical changes usually have far fewer eligible approvers than low-class ones.

The worksheet below is an illustrative example with round numbers, not a benchmark. Replace every value with your own measurements. Read-only work that needs no merge review is left out. The three high-and-critical approvers are among the eight reviewers, so the 48 hours of supply are shared across all three columns.

LineLow (reversible)Medium (shared non-production)High and critical (production or regulated)Total
Merged PRs per week3618660
Review minutes per PR103090—
Review demand, hours per week69924
Eligible reviewers883—
Protected hours per reviewer666—
Review supply, hours per week48 (shared)48 (shared)1848
Utilization——50%50%

At 60 pull requests a week this team has headroom. If agents double the arrival rate and nothing else changes, demand reaches 48 hours: 100% utilization of a queue. Queues do not degrade gently as utilization approaches 100%; waiting time grows steeply, and reviewers start approving without reading. That is a likely mechanism behind the Faros review-time and unreviewed-merge numbers.

You now have three levers, and none of them is “hire more authors”:

  • Lower minutes per PR with stronger oracles and an evidence bundle, so a reviewer judges proof rather than lines.
  • Move a risk class to evidence-only review through the staged trust transfer protocol, once its oracle has earned it.
  • Cap arrivals with WIP limits and PR size budgets, as the review queue page describes.

The builder-to-verifier ratio is the share of a team’s engineering hours spent causing change compared with the hours spent deciding that change is correct. It is a single number that shows whether an org chart matches agentic work.

FieldDefinition
Builder hoursWriting intent and specs, operating agents, fixing what agents could not, hand-written code
Verifier hoursReviewing pull requests and evidence, writing and maintaining tests, evals and acceptance criteria, investigating escaped defects
RatioBuilder hours ÷ verifier hours, per team, per month
SourceA two-week time sample each quarter (calendar blocks plus review timestamps), not continuous tracking
OwnerThe engineering manager collects it; the CTO reviews the trend across teams
Read it withMedian time in review and change failure rate for the same period

No published benchmark gives a “correct” ratio, so read the trend. If agents take over more building and the ratio does not move toward verification, verification is either not being done or hidden inside rubber-stamp approvals. If it moves toward verification and time in review still rises, your oracles are too weak and every review is a full read.

Three responsibilities appear once agents do most of the writing. They start as part-time duties and become roles as the organization grows, each defined by what the person decides.

RoleOwnsDecidesMeasured by
Harness engineerThe shared agent setup: instruction files, skills, hooks, MCP server configuration, sandbox and permission profiles, the CI loops agents run inWhich capabilities agents get, and which checks run before a human is askedShare of agent runs that pass the gates without human repair; time to roll out a harness change to every repository
Eval ownerThe oracles for one domain: acceptance tests, property tests, model-graded checks, the regression suite for the agent setup itselfWhether a check is strong enough to replace a human read for a risk classEscaped defects in the domain; mutation score or seeded-bug catch rate of the suite
Factory operatorThe unattended loops: queue, budgets, stop conditions, the autonomy register, pausing and rolling back a loopWhen a loop runs, stops or loses its autonomyLoop success rate, spend per accepted change, time to stop a misbehaving loop

These roles carry no merge authority of their own. The accountable human stays the named assignee and the code owner, the principle Linear wrote into its product: “issues can only be assigned to humans, and only delegated to agents” (Linear, 2025-08-01).

Staffing depends on size. These are starting points to adjust, not ratios from a study.

Engineering sizeHarness engineerEval ownerFactory operator
About 8 (one team)One engineer, one day a week, rotating each quarterThe tech lead for the team’s critical domainNot needed until the first unattended loop ships
About 40 (five teams)One to two people inside a platform functionOne per critical domain, part-time, named in the autonomy registerOne person owns the loop register and on-call for loops
About 200An agent platform team running the harness as a productA named owner per domain with protected time; evals reviewed like codeA small rota, with the same incident process as production services

The operating model puts these roles into a RACI, and career ladders describes how to reward them.

Manager span used to be sized by people. With agents, the load is the review decisions, loops and incidents the team owns; one engineer supervising several parallel agents produces the review demand of several engineers. Three rules keep span sane:

  1. Count loops, not only people. Each unattended loop adds review demand and an incident path. A team of five with eight loops is not a small team.
  2. Keep one owner per critical domain on every team. Domain knowledge concentrates in fewer heads as agents move fast. Name the owner and a backup before the team shrinks.
  3. Protect the talent pipeline deliberately. Anthropic’s internal study (2025-12-02) describes a “paradox of supervision”: overseeing agents needs the skills that over-delegation erodes. A team with no juniors has no future reviewers. See growing junior developers.

Where does the product–engineering boundary move?

Section titled “Where does the product–engineering boundary move?”

When build time collapses, the bottleneck moves upstream to deciding what to build and how to know it is right. Boris Cherny of Anthropic predicted that the software engineer title will become “builder” or “product manager” (reported second-hand by The San Francisco Standard, 2026-02-19). Treat that as a prediction; today’s decision is narrower: who writes which artifact.

ArtifactWritten byApproved byWhy this split
Intent: problem, user, why nowProductEngineering lead, for feasibility and risk classProduct owns the why; engineering flags what it will cost to verify
Acceptance criteria as examplesProduct and engineering togetherEval ownerThey become the oracle, so they must be testable
Design and risk classEngineeringCode owner, or architecture review for high and criticalArchitecture decides how much acceleration the team keeps (DORA 2025)
ImplementationAgents, operated by engineersChecks, then the named reviewerThe human decision is acceptance, not authorship
Production releasePlatformNamed owner for high and criticalSeparation of duties survives the change

The largest practical change is that product managers write acceptance criteria directly. Product management when build time collapses covers the roadmap and prioritization side.

A productivity multiplier converts a hope into a hiring freeze. Plan from measured constraints instead: collect at least 8 weeks of metrics on a team before you use this table for it, and read the rows in order. The first matching row wins.

Signal, measured over 8+ weeksWhat it meansHeadcount decision
PRs merge without review outside an approved evidence-only risk classThe control is brokenChange nothing in the org until the gate is fixed
Change failure rate or reverted changes risingQuality debt is accumulatingStop extending autonomy; move verifier hours up; no reductions
Time in review rising, reviewer utilization above your targetVerification bindsDo not add builders. Shift people to eval owner and harness work; hire senior reviewers for high and critical changes if eligible approvers are the limit
Review queue healthy, no shaped work waitingIntent bindsMove capacity to product discovery and specification; do not add engineers
Queue healthy, backlog full, stability flat or improvingCapacity is genuinely freedApply your freed-capacity policy; adjust through hiring plans and attrition, reviewed quarterly

The economics page turns the same measurements into cost per accepted change, and the business case page shows how to present the decision with ranges and stop gates.

Freed capacity that nobody allocates turns into more pull requests, which the review station then has to absorb. Faros reports code churn up 861% in the same April 2026 report: code rewritten or deleted soon after it lands, which is where unallocated output tends to show up. Anthropic’s internal study (2025-12-02, self-reported) found 27% of Claude-assisted work consisted of tasks that would not have been done otherwise. Decide in advance which of those tasks you want.

A reorganization is a change to a production system, so give it acceptance criteria before it starts and a rollback if it misses them. Judge it on these signals from the forge, CI and incident tracker.

MetricDefinitionDirection you wantSigned off by
Median time to first reviewOpened → first human review, per risk classFlat or falling as arrivals riseEngineering manager
Unreviewed merge shareMerged PRs with no human review ÷ merged PRs, outside evidence-only classesZeroCTO
Change failure rateDeployments causing a failure in production ÷ deploymentsFlat or fallingPlatform owner
Escaped defects per domainDefects found after release, by domainFalling where an eval owner is namedEval owner
Builder-to-verifier ratioAs defined aboveMoving toward verification as agents build moreCTO
Reviewer concentrationShare of reviewed PRs the top reviewer per team took part in (the reviewers column)No one person above a third of reviewsEngineering manager

Put these on one panel with the frameworks for measuring AI-era delivery, agree the targets before the change, and name the date you will reverse it if two consecutive monthly readings miss.

The org design is the same for every tool. What differs is which review work each tool takes off a human reviewer, and so how many reviewer hours you plan for. Approval stays a human role in all three.

Engineers run /code-review on a local diff before opening a pull request, and claude ultrareview (in a session, /code-review ultra) runs a cloud multi-agent review that reproduces each finding. The managed Code Review service posts inline comments on GitHub pull requests; it is a research preview for Team and Enterprise plans. Neither service is available to Zero Data Retention organizations, and Ultrareview is not available on Bedrock, Google Cloud or Foundry. Plan both as a first pass that lowers minutes per PR, tuned by the harness engineer through CLAUDE.md and REVIEW.md (Code Review docs, checked 2026-09-26).

What goes wrong when teams reshape around agents

Section titled “What goes wrong when teams reshape around agents”

Headcount is cut on a projected multiplier. Review demand stays, reviewers leave, and unreviewed merges appear within weeks. Recovery: freeze reductions, publish the unreviewed-merge share, and restore review supply for medium, high and critical changes before anything else.

The verifier role becomes a rubber stamp. The signature is utilization near 100% and review times falling while change failure rises. Recovery: cap arrivals with WIP limits, spot-check approvals against the evidence bundle, and prioritize eval work until minutes per PR fall honestly.

The harness engineer becomes a gate. Every team waits on one person to add a skill or a hook. Recovery: publish the harness as a versioned product with contribution rules, as the platform team page describes.

Juniors stop being hired. The team looks efficient for a year and then has no one to promote into review. Recovery: reinstate early-career hiring and give juniors review-to-learn work, not only agent operation.

Product ships prototypes to production. Faster building tempts product teams to skip the risk class. Recovery: make the risk class a required field on intent, and keep production release behind the named owner.

  • Which constraint binds on each team today: review, intent or genuine capacity? Show me the eight-week data.
  • What share of merged pull requests had no review last month, and in which risk classes?
  • Who is the named eval owner for each critical domain, and how much protected time do they have?
  • Who can stop an unattended loop, and how long did it take last time?
  • What happens to freed capacity next quarter, and where is that written down?
  • What would make us reverse the reorganization, and when do we check?

Before this page, read the economics of agent-built software and the business case. For hiring against the new role profiles, see hiring for agentic engineering; for why review bandwidth caps the autonomy ladder at Level 3, see Level 3: review diffs.