Tech lead track: move the whole team up one rung
The tech lead track is the reading path for tech leads who need a whole team, not one power user, to ship with Claude Code, Codex, or Cursor: one engineer merges five agent pull requests a day, two refuse the tools, and review is the slowest part of delivery. It covers three jobs: standardize the harness, own verification, and run review capacity.
What is a tech lead’s job when agents write the code?
Section titled “What is a tech lead’s job when agents write the code?”Your job moves from writing code to designing the system that decides whether code ships:
Standardize the harness. The harness is what the team shares around the model: the instructions file, permissions, MCP servers, skills, and one command that runs tests, lint, and types. Keep it in the repository; review it like code.
Own verification design. Cheap generation moves the bottleneck to checking. DORA 2025 (Google Cloud, 23 September 2025) warns that without controls like strong automated testing, “an increase in change volume leads to instability.” You choose the tests, fitness functions, and acceptance criteria that prove a change correct.
Run review capacity. Review is a queue with fixed capacity: Faros AI’s AI Engineering Report 2026 (April 2026; 22,000 developers, vendor telemetry) measured throughput up 33.7%, median time in review up 441.5%, and incidents per pull request up 242.7%. You decide which changes a person reads and which merge on evidence.
Baseline first, transfer trust last: evidence-based review without strong tests only moves the risk.
The tech lead reading path, step by step
Section titled “The tech lead reading path, step by step”Each step is marked Free or Subscribers; the three ladder steps are free. Read them in order, then skip any step whose exit evidence your team already has.
-
Why now: One frame for maturity, process and capability before you change how eight people work.
Done when: Your team placed on the map, loop by loop.
-
Why now: Most teams stall at Level 3; know its review ceiling before you plan past it.
Done when: The team knows how many diffs a week it can really review.
-
Why now: The standard the team moves to: evidence first, code only for escalation classes.
Done when: Escalation classes agreed with the team.
-
Why now: Set the baseline before the rollout, or you will not be able to show what moved.
Done when: Four weeks of baseline data for the metrics you will report.
-
Why now: Senior refusal quietly caps the whole team; bring the skeptics in before the rollout.
Done when: Each senior has shipped one real change with an agent, paired with you.
-
Why now: Roll out one repository at a time, as transitions on the ladder, not a big bang.
Done when: A roadmap with one repository, one loop and one exit criterion per step.
-
Why now: Agents amplify what the repository already is; make it legible and testable first.
Done when: A one-command bootstrap, a fast test suite and a current AGENTS.md.
-
Why now: One shared, versioned rules core instead of eight private setups.
Done when: Shared rules in the repository, reviewed like code.
-
Why now: Agents do well on well-shaped work; shape the backlog so they get it.
Done when: A ticket template with acceptance criteria and a risk class.
-
Why now: Trusting tests you did not read needs a measure of how strong they are.
Done when: A mutation or oracle-strength score on the repositories agents touch.
-
Why now: Define what every agent pull request must prove, in one manifest.
Done when: The evidence bundle required by CI on agent pull requests.
-
Why now: Maintainability is checked by fitness functions, not by reading the code.
Done when: Two architecture rules enforced in CI.
-
Why now: Review by evidence and risk class, so review capacity stops being the bottleneck.
Done when: Review time per agent pull request down, escaped defects flat.
-
Why now: Run the review queue as a system when agents open most of the pull requests.
Done when: A work-in-progress limit and a queue-age alert on the review queue.
-
Why now: Move the team from reading every diff to trusting the evidence, deliberately.
Done when: One risk class merged on evidence alone for a month without an incident.
-
Why now: Juniors still need to grow judgment when agents write the code.
Done when: A growth plan for each junior with review and design work built in.
Which scorecard question does each step move?
Section titled “Which scorecard question does each step move?”Use this as a four-week plan: pick one Harness, one Verification, and one Review row, then re-take the scorecard. Numbers follow the answer key.
| Step | Question it moves | Exit evidence |
|---|---|---|
| Frame: One map | Your band | Team placed on the ladder |
| Review: Level 3, you review the diffs | Q6, review standards | Weekly review capacity known |
| Verification: Evidence instead of code | Q8, plausible-but-wrong code | Escalation classes agreed |
| Baseline: Metrics frameworks | Q11, measuring impact | Four weeks of baseline data |
| People: Skeptics and seniors | Q12, skeptics and seniors | Each senior shipped one agent change |
| Harness: Adoption roadmap | Q10, how adoption happened | One pilot repository, one exit criterion |
| Harness: Agent-ready codebase | Q2 and Q4, context and onboarding | One-command bootstrap, current instructions |
| Harness: Shared agent rules | Q1, a common agent config | Rules in the repository, reviewed like code |
| Harness: Agent-ready backlog | Q5, consistent quality | Ticket template with a risk class |
| Verification: Oracle strength | Q8, plausible-but-wrong code | Mutation score on agent-touched code |
| Verification: Evidence bundle | Q7, automated gates | CI requires the bundle |
| Verification: Fitness functions | Q5, consistent quality | Two architecture rules enforced in CI |
| Review: Agent PR review | Q6 and Q15, standards and PR loop | Review time down, escaped defects flat |
| Review: Review queue | Q15, the PR and merge loop | Work-in-progress limit and queue-age alert |
| Review: Trust transfer | Q24, what AI cannot touch | One risk class merged on evidence |
| People: Junior developers | Q18, levelling up the team | A growth plan per junior |
Where the shared harness lives in each tool
Section titled “Where the shared harness lives in each tool”In every tool, the team’s setup belongs in the repository, not on laptops.
- Instructions:
CLAUDE.mdin the repository. Claude Code falls back toAGENTS.mdfrom v2.1.277 on thelatestchannel, not yet onstable. - Permissions: the shared
.claude/settings.json, which ignores adefaultModeofautoorbypassPermissions. - MCP servers:
claude mcp add --scope project, checked in. - Review:
/code-reviewlocally,anthropics/claude-code-action@v1in CI. Checked against Claude Code 2.1.283 on 26 September 2026.
- Instructions:
AGENTS.md, loaded once the repository is marked trusted (since Codex CLI 0.150.0). - Guardrails: admin constraints in
requirements.toml. - Review:
/reviewin the CLI,@codex reviewon pull requests (GitHub review verified 28 August 2026), andopenai/codex-action@v1in CI. - Headless checks:
codex -a never exec …in a sandboxed runner (Codex CLI 0.157.1, checked 26 September 2026).
- Instructions: project Rules, versioned with the code.
- Guardrails: Hooks, plus Plugins that bundle rules, skills, MCP servers, and hooks.
- Review: Bugbot on pull requests, and PR Routing & Approval, which assigns reviewers by ownership and can approve low-risk pull requests.
- Cursor details last verified 28 August 2026.
Why does a team rollout stall?
Section titled “Why does a team rollout stall?”- One power user carries the numbers. The team average can hide engineers still at Level 1. Recovery: pair each holdout with the power user on a real ticket.
- Review time rises after the rollout. Recovery: cap agent pull requests in progress and route changes by risk class before adding reviewers.
- The harness drifts. Recovery: change instructions and settings only through pull requests; re-run the audit prompt monthly.
- Trust transfers before the oracle is strong. Recovery: return that risk class to line-by-line review until oracle strength holds.
Where to go next as a tech lead
Section titled “Where to go next as a tech lead”Frequently asked questions
What is the tech lead track?
An ordered reading path for tech leads and heads of team who are moving a whole team, not one power user, up the autonomy ladder. It opens with the free Tech Lead Scorecard as a baseline, then runs through three jobs: standardizing the harness, owning verification design, and running review capacity. Every step names the scorecard question it moves and the evidence that says the team can move on.
Why does the track start with the scorecard?
The scorecard takes about 8 minutes and scores the team on 25 questions. Its answers show which steps the team already meets, and re-taking it after four weeks shows which section scores moved. The answer key maps every question to the page that moves it.
Which steps are free?
The three ladder steps at the start of the path are free. Every step is marked Free or Subscribers.