How teams, roles and headcount change shape
Engineering org design for agentic engineering sizes teams by verification capacity rather than by how much code people can write. Agents raise the rate at which changes arrive, so review bandwidth, eval coverage and accountable owners become the binding constraints. Headcount decisions follow those measured constraints, not a vendor’s productivity multiplier.
This page is for executives and CTOs facing a headcount question they cannot yet answer: the board has read that agents write most of the code at some companies, while your own teams merge more pull requests than ever and wait longer for review.
What you can adopt for your org design
Section titled “What you can adopt for your org design”- A review capacity worksheet that turns arrival rate and risk mix into reviewer hours.
- Role charters for harness engineer, eval owner and factory operator, staffed at 8, 40 and 200 engineers.
- A headcount decision table that maps measured signals to decisions, with no multiplier in it.
- A freed-capacity policy you can adopt as written.
- The metric definitions that tell you whether the new shape works, and who signs off on each.
What the evidence says about team shape
Section titled “What the evidence says about team shape”Generation scales faster than verification. Faros AI’s telemetry from 22,000 developers and more than 4,000 teams (April 2026) shows epics completed per developer +66.2% and task throughput per developer +33.7%, alongside median time in review +441.5%, PRs merged without any review +31.3%, and incidents per pull request +242.7%. The work grows, and verifying it falls to reviewers.
The organization’s foundation decides who gains. DORA’s 2025 report (Google Cloud, 2025-09-23) links AI adoption to higher delivery throughput and lower stability, and explains: “Teams working in loosely coupled architectures with fast feedback loops see gains, while those constrained by tightly coupled systems and slow processes see little or no benefit.”
The job shifts toward review, and the reviewer stays. In Anthropic’s internal study of 132 engineers and researchers (2025-12-02, a vendor’s own self-report), some engineers describe their work shifting “70%+ to being a code reviewer/reviser rather than a net-new code writer”. Stripe merges more than 1,300 pull requests a week that are “completely minion-produced, human-reviewed, but containing no human-written code” (Stripe engineering blog, 2026-02-19): removing the writing did not remove the reviewer.
When generation is cheap, the scarce resources are the people who can decide whether a change is correct, the oracles that let them decide without reading every line, and the owners who answer for the result.
Size teams by review capacity, not by output
Section titled “Size teams by review capacity, not by output”With agents, a team’s capacity is the number of changes it can verify and own, not the number it can write. Plan it in three steps.
-
Measure the arrival rate. Pull 8–12 weeks of merged pull requests from your forge. The command below runs in a terminal with the GitHub CLI and writes one row per pull request: number, author, opened, first human review, merged, lines changed, distinct human reviewers and their logins.
Terminal window gh pr list --state merged --limit 1000 \--search "merged:>=2026-07-01" \--json number,author,createdAt,mergedAt,additions,deletions,reviews \--jq '.[] | . as $pr| [.reviews[] | select((.author.login // "ghost") != $pr.author.loginand ((.author.login // "ghost") | test("\\[bot\\]$|^app/|copilot|codex|claude|bugbot"; "i") | not))] as $h| [.number, .author.login, .createdAt,([$h[].submittedAt] | min), .mergedAt,(.additions + .deletions),([$h[] | .author.login // "ghost"] | unique | length),([$h[] | .author.login // "ghost"] | unique | join(","))] | @tsv' \> review-load.tsvThe filter drops the author’s own reviews and review bots, because a bot review or a self-review is not a human review. Check the bot logins in your own export and add any the pattern misses; the pattern also drops any human login containing these words, so check the list. An empty first-review column with zero reviewers is then a pull request that merged without human review. Count those first.
-
Price each change by risk class. Use the four risk classes (low, medium, high, critical) from governance and autonomy and agree on the minutes a competent reviewer needs per class, including reading the evidence bundle. Time five real reviews per class rather than estimating.
-
Compare demand with protected supply. Supply is the hours per week each reviewer actually protects for review, restricted to the people allowed to approve each class. High and critical changes usually have far fewer eligible approvers than low-class ones.
The worksheet below is an illustrative example with round numbers, not a benchmark. Replace every value with your own measurements. Read-only work that needs no merge review is left out. The three high-and-critical approvers are among the eight reviewers, so the 48 hours of supply are shared across all three columns.
| Line | Low (reversible) | Medium (shared non-production) | High and critical (production or regulated) | Total |
|---|---|---|---|---|
| Merged PRs per week | 36 | 18 | 6 | 60 |
| Review minutes per PR | 10 | 30 | 90 | — |
| Review demand, hours per week | 6 | 9 | 9 | 24 |
| Eligible reviewers | 8 | 8 | 3 | — |
| Protected hours per reviewer | 6 | 6 | 6 | — |
| Review supply, hours per week | 48 (shared) | 48 (shared) | 18 | 48 |
| Utilization | — | — | 50% | 50% |
At 60 pull requests a week this team has headroom. If agents double the arrival rate and nothing else changes, demand reaches 48 hours: 100% utilization of a queue. Queues do not degrade gently as utilization approaches 100%; waiting time grows steeply, and reviewers start approving without reading. That is a likely mechanism behind the Faros review-time and unreviewed-merge numbers.
You now have three levers, and none of them is “hire more authors”:
- Lower minutes per PR with stronger oracles and an evidence bundle, so a reviewer judges proof rather than lines.
- Move a risk class to evidence-only review through the staged trust transfer protocol, once its oracle has earned it.
- Cap arrivals with WIP limits and PR size budgets, as the review queue page describes.
What is the builder-to-verifier ratio?
Section titled “What is the builder-to-verifier ratio?”The builder-to-verifier ratio is the share of a team’s engineering hours spent causing change compared with the hours spent deciding that change is correct. It is a single number that shows whether an org chart matches agentic work.
| Field | Definition |
|---|---|
| Builder hours | Writing intent and specs, operating agents, fixing what agents could not, hand-written code |
| Verifier hours | Reviewing pull requests and evidence, writing and maintaining tests, evals and acceptance criteria, investigating escaped defects |
| Ratio | Builder hours ÷ verifier hours, per team, per month |
| Source | A two-week time sample each quarter (calendar blocks plus review timestamps), not continuous tracking |
| Owner | The engineering manager collects it; the CTO reviews the trend across teams |
| Read it with | Median time in review and change failure rate for the same period |
No published benchmark gives a “correct” ratio, so read the trend. If agents take over more building and the ratio does not move toward verification, verification is either not being done or hidden inside rubber-stamp approvals. If it moves toward verification and time in review still rises, your oracles are too weak and every review is a full read.
Which new roles do agentic teams need?
Section titled “Which new roles do agentic teams need?”Three responsibilities appear once agents do most of the writing. They start as part-time duties and become roles as the organization grows, each defined by what the person decides.
| Role | Owns | Decides | Measured by |
|---|---|---|---|
| Harness engineer | The shared agent setup: instruction files, skills, hooks, MCP server configuration, sandbox and permission profiles, the CI loops agents run in | Which capabilities agents get, and which checks run before a human is asked | Share of agent runs that pass the gates without human repair; time to roll out a harness change to every repository |
| Eval owner | The oracles for one domain: acceptance tests, property tests, model-graded checks, the regression suite for the agent setup itself | Whether a check is strong enough to replace a human read for a risk class | Escaped defects in the domain; mutation score or seeded-bug catch rate of the suite |
| Factory operator | The unattended loops: queue, budgets, stop conditions, the autonomy register, pausing and rolling back a loop | When a loop runs, stops or loses its autonomy | Loop success rate, spend per accepted change, time to stop a misbehaving loop |
These roles carry no merge authority of their own. The accountable human stays the named assignee and the code owner, the principle Linear wrote into its product: “issues can only be assigned to humans, and only delegated to agents” (Linear, 2025-08-01).
Staffing depends on size. These are starting points to adjust, not ratios from a study.
| Engineering size | Harness engineer | Eval owner | Factory operator |
|---|---|---|---|
| About 8 (one team) | One engineer, one day a week, rotating each quarter | The tech lead for the team’s critical domain | Not needed until the first unattended loop ships |
| About 40 (five teams) | One to two people inside a platform function | One per critical domain, part-time, named in the autonomy register | One person owns the loop register and on-call for loops |
| About 200 | An agent platform team running the harness as a product | A named owner per domain with protected time; evals reviewed like code | A small rota, with the same incident process as production services |
The operating model puts these roles into a RACI, and career ladders describes how to reward them.
How does span of control change?
Section titled “How does span of control change?”Manager span used to be sized by people. With agents, the load is the review decisions, loops and incidents the team owns; one engineer supervising several parallel agents produces the review demand of several engineers. Three rules keep span sane:
- Count loops, not only people. Each unattended loop adds review demand and an incident path. A team of five with eight loops is not a small team.
- Keep one owner per critical domain on every team. Domain knowledge concentrates in fewer heads as agents move fast. Name the owner and a backup before the team shrinks.
- Protect the talent pipeline deliberately. Anthropic’s internal study (2025-12-02) describes a “paradox of supervision”: overseeing agents needs the skills that over-delegation erodes. A team with no juniors has no future reviewers. See growing junior developers.
Where does the product–engineering boundary move?
Section titled “Where does the product–engineering boundary move?”When build time collapses, the bottleneck moves upstream to deciding what to build and how to know it is right. Boris Cherny of Anthropic predicted that the software engineer title will become “builder” or “product manager” (reported second-hand by The San Francisco Standard, 2026-02-19). Treat that as a prediction; today’s decision is narrower: who writes which artifact.
| Artifact | Written by | Approved by | Why this split |
|---|---|---|---|
| Intent: problem, user, why now | Product | Engineering lead, for feasibility and risk class | Product owns the why; engineering flags what it will cost to verify |
| Acceptance criteria as examples | Product and engineering together | Eval owner | They become the oracle, so they must be testable |
| Design and risk class | Engineering | Code owner, or architecture review for high and critical | Architecture decides how much acceleration the team keeps (DORA 2025) |
| Implementation | Agents, operated by engineers | Checks, then the named reviewer | The human decision is acceptance, not authorship |
| Production release | Platform | Named owner for high and critical | Separation of duties survives the change |
The largest practical change is that product managers write acceptance criteria directly. Product management when build time collapses covers the roadmap and prioritization side.
Plan headcount without multipliers
Section titled “Plan headcount without multipliers”A productivity multiplier converts a hope into a hiring freeze. Plan from measured constraints instead: collect at least 8 weeks of metrics on a team before you use this table for it, and read the rows in order. The first matching row wins.
| Signal, measured over 8+ weeks | What it means | Headcount decision |
|---|---|---|
| PRs merge without review outside an approved evidence-only risk class | The control is broken | Change nothing in the org until the gate is fixed |
| Change failure rate or reverted changes rising | Quality debt is accumulating | Stop extending autonomy; move verifier hours up; no reductions |
| Time in review rising, reviewer utilization above your target | Verification binds | Do not add builders. Shift people to eval owner and harness work; hire senior reviewers for high and critical changes if eligible approvers are the limit |
| Review queue healthy, no shaped work waiting | Intent binds | Move capacity to product discovery and specification; do not add engineers |
| Queue healthy, backlog full, stability flat or improving | Capacity is genuinely freed | Apply your freed-capacity policy; adjust through hiring plans and attrition, reviewed quarterly |
The economics page turns the same measurements into cost per accepted change, and the business case page shows how to present the decision with ranges and stop gates.
Where should freed capacity go?
Section titled “Where should freed capacity go?”Freed capacity that nobody allocates turns into more pull requests, which the review station then has to absorb. Faros reports code churn up 861% in the same April 2026 report: code rewritten or deleted soon after it lands, which is where unallocated output tends to show up. Anthropic’s internal study (2025-12-02, self-reported) found 27% of Claude-assisted work consisted of tasks that would not have been done otherwise. Decide in advance which of those tasks you want.
How you know the new shape works
Section titled “How you know the new shape works”A reorganization is a change to a production system, so give it acceptance criteria before it starts and a rollback if it misses them. Judge it on these signals from the forge, CI and incident tracker.
| Metric | Definition | Direction you want | Signed off by |
|---|---|---|---|
| Median time to first review | Opened → first human review, per risk class | Flat or falling as arrivals rise | Engineering manager |
| Unreviewed merge share | Merged PRs with no human review ÷ merged PRs, outside evidence-only classes | Zero | CTO |
| Change failure rate | Deployments causing a failure in production ÷ deployments | Flat or falling | Platform owner |
| Escaped defects per domain | Defects found after release, by domain | Falling where an eval owner is named | Eval owner |
| Builder-to-verifier ratio | As defined above | Moving toward verification as agents build more | CTO |
| Reviewer concentration | Share of reviewed PRs the top reviewer per team took part in (the reviewers column) | No one person above a third of reviews | Engineering manager |
Put these on one panel with the frameworks for measuring AI-era delivery, agree the targets before the change, and name the date you will reverse it if two consecutive monthly readings miss.
Where each tool moves the review load
Section titled “Where each tool moves the review load”The org design is the same for every tool. What differs is which review work each tool takes off a human reviewer, and so how many reviewer hours you plan for. Approval stays a human role in all three.
Engineers run /code-review on a local diff before opening a pull request, and claude ultrareview (in a session, /code-review ultra) runs a cloud multi-agent review that reproduces each finding. The managed Code Review service posts inline comments on GitHub pull requests; it is a research preview for Team and Enterprise plans. Neither service is available to Zero Data Retention organizations, and Ultrareview is not available on Bedrock, Google Cloud or Foundry. Plan both as a first pass that lowers minutes per PR, tuned by the harness engineer through CLAUDE.md and REVIEW.md (Code Review docs, checked 2026-09-26).
codex review runs a non-interactive review against --uncommitted changes, a --base branch or a --commit (checked on CLI 0.157.1), so it fits as a CI step the harness engineer owns. On GitHub, @codex review on a pull request requests a review, and the eval owner encodes the domain’s review rules in AGENTS.md (vendor docs, checked 2026-08-28).
Bugbot reviews pull requests for bugs. PR Routing & Approval “assigns reviewers based on code ownership and commit history, and can approve low-risk PRs when your criteria are met” (Cursor docs, checked 2026-08-28). That last clause is an org-design decision, not a setting: only let it approve a risk class that has passed your trust-transfer stages, and name the human accountable.
What goes wrong when teams reshape around agents
Section titled “What goes wrong when teams reshape around agents”Headcount is cut on a projected multiplier. Review demand stays, reviewers leave, and unreviewed merges appear within weeks. Recovery: freeze reductions, publish the unreviewed-merge share, and restore review supply for medium, high and critical changes before anything else.
The verifier role becomes a rubber stamp. The signature is utilization near 100% and review times falling while change failure rises. Recovery: cap arrivals with WIP limits, spot-check approvals against the evidence bundle, and prioritize eval work until minutes per PR fall honestly.
The harness engineer becomes a gate. Every team waits on one person to add a skill or a hook. Recovery: publish the harness as a versioned product with contribution rules, as the platform team page describes.
Juniors stop being hired. The team looks efficient for a year and then has no one to promote into review. Recovery: reinstate early-career hiring and give juniors review-to-learn work, not only agent operation.
Product ships prototypes to production. Faster building tempts product teams to skip the risk class. Recovery: make the risk class a required field on intent, and keep production release behind the named owner.
Questions to ask your CTO
Section titled “Questions to ask your CTO”- Which constraint binds on each team today: review, intent or genuine capacity? Show me the eight-week data.
- What share of merged pull requests had no review last month, and in which risk classes?
- Who is the named eval owner for each critical domain, and how much protected time do they have?
- Who can stop an unattended loop, and how long did it take last time?
- What happens to freed capacity next quarter, and where is that written down?
- What would make us reverse the reorganization, and when do we check?
Where to go next with org design
Section titled “Where to go next with org design”Before this page, read the economics of agent-built software and the business case. For hiring against the new role profiles, see hiring for agentic engineering; for why review bandwidth caps the autonomy ladder at Level 3, see Level 3: review diffs.