Skip to content

Tech Lead Scorecard answer key

The Tech Lead Scorecard answer key maps each of the 25 questions on the Tech Lead Scorecard to the page that moves that answer and names the evidence a 3-point answer needs. It is for a tech lead or engineering manager choosing which three gaps to close first and how to prove they closed.

Your result says Level 2 · Paired, 31 points out of 75, with the headline “Team Enabler”. Two developers run parallel agents and ship clean pull requests; the other six paste into chat and hand you diffs to read. The review queue is the bottleneck, and 25 per-question pages is too much reading before Monday. This key picks the first three questions and tells you what counts as done.

What does each Level 1–4 band mean for a team?

Section titled “What does each Level 1–4 band mean for a team?”

The bands borrow the names of the autonomy ladder, but they measure something else: the practices around one team’s agent use, as you report them. The ladder measures how each loop ships merged change. One map shows how to place a team by what actually reaches production, so use it to check your self-reported band. Answer for the team’s median developer, not its strongest one.

BandPointsHeadline on the results pageFirst move
Level 1 · Assisted0–18Lone Wolf: it works for you, now make it work for the teamSection 1: commit one shared, versioned rules file to the most active repository.
Level 2 · Paired19–37Team Enabler: pockets of consistency, connect themSection 2: the same automated gates on every pull request, whoever or whatever wrote it.
Level 3 · Review manager38–56Force Multiplier: the team follows one way of workingSections 4 and 6: parallel agents with a concurrency cap, and cost and secrets under an owner.
Level 4 · Spec manager57–75Org Amplifier: a self-sustaining team systemTeach the system to another team. Moving the team off reading every diff is the work that remains.

How do you pick the first three questions to fix?

Section titled “How do you pick the first three questions to fix?”
  1. Divide each section’s points by its maximum. A 6 out of 15 in rollout (40%) is a bigger gap than 6 out of 12 in enablement (50%).
  2. Take the two sections with the lowest share. In each, pick the question where your answer is furthest from 3 points.
  3. If review and quality gates (questions 5–8) sit below half, make one of the three picks a question from that section, even when another section scores lower. More agent output through weak gates only lengthens the review queue. For the same reason, raise team-scale workflow (questions 14–17) only after the gates hold.
  4. For each pick, write the target answer and the artifact that will prove it into the plan file below, with one owner per question.
  5. Re-take the scorecard after four weeks. Compare section shares, not the band: one question moving from 1 to 3 points rarely changes the band, but it shows in its section.

Commit the plan to the team’s main repository, next to the shared agent rules, so it is reviewed like any other change:

# .agents/scorecard-plan.yaml — one entry per question the team is working on
- question: 7 # What automated gates run on AI-authored code?
current_points: 1 # "Lint only"
target_points: 3 # "Layered: AI review + CI + coverage gates"
evidence: "Branch protection requires lint, types, tests, coverage and the review agent; three merged agent PRs show all five green"
owner: "priya"
check_on: 2026-10-26

Each table gives the question, the page the results card links (Start with), a second page that goes deeper, and the evidence that earns 3 points. Claim a 3 only when you can link that evidence.

Shared standards and context (questions 1–4, 12 points)

Section titled “Shared standards and context (questions 1–4, 12 points)”
#QuestionStart with · Go deeperEvidence for 3 points
1Does your team share a common agent config?Shared agent rules · Concise repository contextA reviewed AGENTS.md or CLAUDE.md in every active repository, changed only through pull requests
2How is project context available to everyone’s agent?Documentation as context · Team collaborationArchitecture and convention docs linked from the rules file and updated in the same pull request as the code they describe
3Are prompt and workflow patterns shared across the team?Shared skills · Building custom skillsA skills or commands directory in the repository, one owner per entry
4How fast is a new developer productive with your setup?Developer onboarding · The developer trackOne command that sets up a fresh clone, and a written first-week path the last joiner followed

Review and quality gates (questions 5–8, 12 points)

Section titled “Review and quality gates (questions 5–8, 12 points)”
#QuestionStart with · Go deeperEvidence for 3 points
5How consistent is AI-generated code quality across the team?Code quality gates · Fitness functionsThe same required checks on every pull request, and architecture rules enforced in CI
6Do you have review standards specific to AI-assisted PRs?Team PR review automation · Running the review queueA written standard for agent pull requests, with its mechanical parts enforced by CI
7What automated gates run on AI-authored code?Layered pull-request review · The evidence bundleBranch protection that requires CI, a coverage gate and an AI review before merge
8How do you catch plausible-but-wrong code before merge?Test: give the session a feedback loop · Catching plausible-but-wrong codeTests the agent must pass before it stops, plus an adversarial review pass with its own prompt

Rollout and adoption (questions 9–13, 15 points)

Section titled “Rollout and adoption (questions 9–13, 15 points)”
#QuestionStart with · Go deeperEvidence for 3 points
9What share of the team actively uses agentic tools?Team adoption rate · Team onboarding and adoptionDaily use by more than 80% of the team, read from a usage report rather than a show of hands
10How did adoption happen on your team?Adoption roadmap · Designing a pilotA written rollout with a pilot repository, a support owner and an exit criterion
11Do you measure the impact of AI?AI metrics panel · Metrics frameworksA dashboard with a pre-rollout baseline and a goal, reviewed on a fixed date
12How do skeptics and senior engineers engage?Team onboarding and adoption · Skeptics and senior engineersSenior engineers authored or approved the shared rules and gates
13Who owns tooling decisions and budget?Tooling policy · Cost governanceA named owner, a budget line and a review date, written down

Team-scale workflow (questions 14–17, 12 points)

Section titled “Team-scale workflow (questions 14–17, 12 points)”
#QuestionStart with · Go deeperEvidence for 3 points
14Does the team use parallel agents or worktrees?Team parallelism · Parallel agents in worktreesA documented worktree convention with a concurrency cap per developer
15How automated is your PR, review and merge loop?A bounded PR review-fix loop · Team PR review automationThe agent opens the pull request and fixes review findings for a capped number of rounds; a human merges
16Do you use shared MCP servers across the team?Internal MCP servers · Essential MCP serversA checked-in MCP configuration with an owner per server
17How do you handle large, cross-cutting changes?Million-line codebase strategies · Shaping a backlog for agentsA reviewed plan before execution, and the change split into tickets an agent can finish and verify

Enablement and mentoring (questions 18–21, 12 points)

Section titled “Enablement and mentoring (questions 18–21, 12 points)”
#QuestionStart with · Go deeperEvidence for 3 points
18How do you level up the team’s AI skills?The developer track · Upskilling a teamA curriculum with practice and review of the output, and a record of who completed it
19Do you run internal knowledge-sharing on AI workflows?Knowledge sharing · Workflow transformationA recurring session on the calendar and the artifacts each one produced
20When someone finds a great workflow, how does it spread?Shared skills · Agent skillsThe workflow lives as a skill in the repository and was taught in a session
21How do you keep the team current as tools change?AI tooling roadmap · Keeping a team currentAn owner who triages releases, a sandbox to try them, and recorded adopt-or-skip decisions

Team ops and governance (questions 22–25, 12 points)

Section titled “Team ops and governance (questions 22–25, 12 points)”
#QuestionStart with · Go deeperEvidence for 3 points
22How do you manage AI cost across the team?Cost governance · Model routing by evidenceSpend per person or per repository reviewed monthly, plus written model-routing guidance
23How are secrets and permissions handled for team agents and MCP?MCP security · Agent identity and secretsSecrets from a secret manager, a scoped token per agent integration, and a written policy
24Is there a policy for what AI cannot touch?Governance and autonomy · Hooks as deterministic guardrailsWritten limits enforced by permission settings or hooks, not only by instructions in a prompt
25How do you decide which tools the team standardizes on?Tooling policy · Tool comparisonA trial on the team’s own tasks, the chosen standard, and a revisit date

How do you prove a higher answer without reading every diff?

Section titled “How do you prove a higher answer without reading every diff?”

Every 3-point answer above names an artifact, not an intention: a branch-protection export, a checked-in config, a dashboard with a baseline, merged pull requests that carry an evidence bundle. The owner in the plan file produces the artifact, and a second engineer checks it before you re-take the scorecard. Reading evidence instead of code explains why that check replaces reading each diff.

An agent can collect most of that evidence from the repository, provided it runs read-only. Do not rely on the prompt alone to keep it read-only; start it in a mode that does not edit files.

Start the session in plan mode, which explores without editing files: claude --permission-mode plan (checked in claude --help, v2.1.283). Without the flag, interactive sessions from v2.1.283 (the latest release channel) usually start in auto mode, which can edit files. Paste a prompt below.

The prompt text is the same in all three tools.

Branch protection, required reviewers and usage reports live in your Git host and your tools’ admin consoles, not in the repository. Export them and attach them to the same plan entry.

When does a Tech Lead Scorecard result mislead?

Section titled “When does a Tech Lead Scorecard result mislead?”
  • You scored intentions. A review checklist nobody enforces earns the same answer as a required check if you let it. Recovery: re-score every answer you cannot back with a linked artifact one point lower; that total is your real starting point.
  • You answered for your two strongest developers. Their parallel agents do not make the team Level 3. Recovery: answer for the median developer, and use question 9’s usage report to check.
  • Workflow rose before the gates. More parallel agents and auto-opened pull requests against the same review capacity grow the queue until reviewers approve without looking. Recovery: cap concurrent agent pull requests per developer, fix questions 6–8 first, then raise the cap.
  • Seniors were routed around, not brought in. Adoption climbs while question 12 stays at 1 and the seniors quietly re-review everything. Recovery: give them ownership of the rules and the gates, as skeptics and senior engineers describes.
  • The band became a target. A jump of a band in four weeks usually means optimistic answers. Recovery: re-check each changed answer against its artifact and report section shares to your manager, not the band.

Where to go next from your Tech Lead Scorecard result

Section titled “Where to go next from your Tech Lead Scorecard result”