Skip to content

Team: leading a team that ships with agents

To lead a team that ships with agents, make the setup a team system, not one enthusiast’s habit: roll out one repository at a time, share one versioned harness, fix review gates before adding parallel agents, and give seniors ownership of those gates. Each part of this hub names the Tech Lead Scorecard questions it answers.

You just got approval to roll agents out to all eight engineers on your team. Each of them already uses a different tool with private settings, CI checks only lint, and nobody has decided who reviews what an agent opens. You have to choose what to standardise first. These pages give you that order and a way to check it worked.

Take the Tech Lead Scorecard first (25 questions, about eight minutes). The answer key maps each question to its page and names the evidence a 3-point answer needs. If you have no score yet, start from the symptom you see:

What you see on the teamStart withScorecard questions
Everyone configures their agent differentlyShared agent rules · Developer onboarding1–4
Agent pull requests wait days for reviewRunning the review queue6 (14, 15 related)
Plausible-but-wrong code reaches mainThe evidence bundle7 (5, 8 related)
Agents get tickets they cannot finish or verifyShaping a backlog for agents17
Seniors re-review everything or opt outSkeptics and senior engineers12
Usage is patchy and nobody knows whyAdoption roadmap9–11
Juniors ship code they cannot explainUpskilling a team · Growing junior developers (related)18

Roll out: from one pilot repository to the whole team

Section titled “Roll out: from one pilot repository to the whole team”

A rollout moves one repository at a time, with a pilot, a support owner and an exit criterion written down before it starts. These pages answer questions 4 and 9–11.

The roll-out hub also links the guides for teams switching from another tool.

Shared harness: one setup instead of eight private ones

Section titled “Shared harness: one setup instead of eight private ones”

The harness is the shared instruction core (AGENTS.md, with thin adapters such as CLAUDE.md and Cursor’s Rules), skills, hooks and MCP servers every agent on the team loads. It lives in the repository and changes only through pull requests. These pages answer questions 1–3, 16 and 20, and they cover Claude Code, Codex and Cursor side by side.

Review and flow: keep the queue moving without rubber-stamping

Section titled “Review and flow: keep the queue moving without rubber-stamping”

Agents raise the arrival rate of pull requests; review capacity stays where it was. Fix the gates before you add parallel agents, or reviewers start approving without looking. These pages answer questions 5–8, 14, 15 and 17; the answer key lists the second page for each.

People: skeptics, seniors, juniors and keeping current

Section titled “People: skeptics, seniors, juniors and keeping current”

The tools change weekly; the people decide whether the system holds. These pages answer questions 12, 18, 19 and 21.

Questions 13 and 22–25 (tooling ownership and budget, cost, secrets, what agents may not touch, which tools to standardise on) sit with the organization. Their pages are listed in the answer key; most sit in Engineering organization, with tooling policy and MCP security under the CTO scorecard guide.

How do you know the team system works without reading every diff?

Section titled “How do you know the team system works without reading every diff?”

Ask for artifacts, not reassurance. Every merged agent pull request carries an evidence bundle: the acceptance criteria, the tests that prove them, and the required checks that passed. Branch protection enforces the same gates on every pull request, whoever or whatever wrote it. Reading evidence instead of code explains why that replaces line-by-line review.

Then watch three numbers each month: time in review, change failure rate (the share of deployments that cause a failure in production), and the share of sampled pull requests where a human found something the gates missed. The tech lead owns the gates; a named engineer owns the evidence for each loop (a repeatable class of change with its own trigger, oracle and stop condition). Treat these as a team contract and commit it next to AGENTS.md:

Replace N with the team-wide WIP cap your review capacity supports; the review queue page shows how to size it.

What goes wrong when a team rolls out agents?

Section titled “What goes wrong when a team rolls out agents?”
  • Agents amplify weak gates. DORA’s 2025 report (Google Cloud, 23 September 2025) puts it plainly: “AI doesn’t fix a team; it amplifies what’s already there.” Recovery: freeze new parallelism, fix questions 5–8, then raise the cap.
  • The harness forked. Each developer tuned a private setup, and agents now disagree on conventions. Recovery: merge the private rules into the shared file through one reviewed pull request, and route later changes the same way.
  • Trust moved faster than evidence. The team skipped review on a loop that had weak tests, and a defect escaped. Recovery: move that loop back one stage, as the trust transfer page describes, and strengthen its tests before moving forward again.
  • Seniors were routed around. Adoption numbers look good while senior engineers re-review everything. Recovery: hand them ownership of the gates and the rules.