Skip to content

CTO Scorecard answer key

The CTO Scorecard answer key explains the four Level 1–4 bands the CTO Scorecard reports and maps each of its 25 questions to the page that moves that answer. It is written for a CTO or VP Engineering who has a result and needs to decide which controls to fix first, and what evidence proves the fix.

Your result says Level 2 · Paired, 34 points out of 75, and the budget review is in six weeks. Every section shows a gap, and 25 per-question pages is too much reading before Monday. This key picks the first three questions, points to where each is answered, and names the artifact that earns the higher answer.

The bands borrow the names of the autonomy ladder but measure something else: the controls around agent use, as you report them. The ladder measures how each loop ships merged change, and one map turns that into an organization-level distribution. Use it to check your self-reported band against what reaches production. Answer for your median team, not your best one.

BandPointsWhat it describesFirst move
Level 1 · Assisted0–18Engineers use AI in spite of the organization: unapproved tools, no audit trail, spend nobody can see. The top risks are a compliance incident and runaway spend.Section 1: inventory who uses what, move to centrally governed accounts, and turn on per-person spend visibility.
Level 2 · Paired19–37Tools are standardized and basic review is in place; each developer works with one agent.Section 2: shared agent rules, shared skills, and governed team accounts.
Level 3 · Review manager38–56Layered review, internal MCP servers, and metrics exist. Agents write, engineers review.Sections 4 and 6: parallel work with concurrency caps, and a funded roadmap with stop gates.
Level 4 · Spec manager57–75Accepted intent and evidence drive delivery, approval authority follows risk, capability bets have stop gates, and vendor continuity is exercised.Improve the loop you have. The scorecard has no Level 5 band; running the software factory is a per-loop decision, not a score target.

How do you pick the first three questions to fix?

Section titled “How do you pick the first three questions to fix?”
  1. Divide each section’s points by its maximum. A 6 out of 15 in Section 3 (40%) is a bigger gap than 6 out of 9 in Section 4 (67%), although the raw points are equal.
  2. Take the two sections with the lowest share. In each, pick the question where your answer is furthest from 3 points.
  3. If Section 3 is below half and not already one of the two, add its weakest question; otherwise take the weakest question from the third-lowest section. Parallel and unattended agent work (Section 4) multiplies the pull requests your gates must absorb. Raising Section 4 before Section 3 moves the bottleneck into the review queue.
  4. For each question, write the target answer and the artifact that will prove it in the plan file below. Name one owner per question.
  5. Re-take the scorecard after 30 days. Compare section shares, not the band: one question moving from 1 to 3 points rarely changes the band, but it shows in its section.

Commit the plan to the repository that holds your agent platform configuration, so changes to it are reviewed like any other change:

# ai-scorecard/plan.yaml — one entry per question you are working on
- question: 9 # What automatically reviews your PRs today?
current_points: 1
target_points: 3
target_answer: "Deterministic checks, a focused risk review, reproducible evidence, and a named human gate by impact"
evidence: "Branch protection export listing required checks; three merged agent PRs with an evidence bundle attached"
owner: "platform-lead"
check_on: 2026-10-26

Each table gives the question, the page to read first and a second one, and the evidence a 3-point answer needs. Questions 4, 10, and 21 link their topic page directly, because their per-question pages are being merged into it.

Adoption and spend (questions 1–4, 12 points)

Section titled “Adoption and spend (questions 1–4, 12 points)”
#QuestionReadEvidence for 3 points
1How broadly and effectively does the team use approved AI workflows?Team adoption · Team onboarding and adoptionCompleted approved workflows across eligible engineers, with accepted-outcome, quality, and control data
2What is the tooling policy?Tooling policy · Managed policyAn approved workflow matrix and exceptions with an expiry date and a cleanup record
3How do you govern AI accounts, access lifecycle, data terms, and spend?Team accounts · Procurement questionnaireA named owner per service, joiner/mover/leaver controls that run, reviewed data terms, auditable spend
4Can you connect total AI workflow cost to accepted outcomes and quality?AI usage cost governance · Cost per accepted changeTotal cost per accepted outcome, with budgets, alerts, and a named owner

Shared infrastructure (questions 5–8, 12 points)

Section titled “Shared infrastructure (questions 5–8, 12 points)”
#QuestionReadEvidence for 3 points
5Does the team share agent rules?Shared agent rules · Team collaborationA versioned shared core with tested per-tool adapters, and hard requirements enforced outside the model
6Does the team share skills?Shared skills · Building your own skillsEach skill has an owner, positive, negative, and boundary fixtures, and usage data
7Do internal MCP services solve measured, recurring access gaps safely?Internal MCP servers · Introduction to MCPEach server traces to a measured gap and has an owner, least privilege, monitoring, and retirement criteria
8What is your MCP security model?MCP security · Agent identity and secretsA reviewed allowlist, per-tool authorization, short-lived credentials, audit logs, and explicit approval for writes

Quality gates (questions 9–13, 15 points)

Section titled “Quality gates (questions 9–13, 15 points)”
#QuestionReadEvidence for 3 points
9What automatically reviews your PRs today?Team PR review automation · Reviewing an agent’s PRRequired deterministic checks, a focused risk review, reproducible evidence, and a named human gate sized to impact
10How do you handle human review bandwidth for agent-authored PRs?Governance and autonomy · The review queueDesign-time boundaries, automated evidence gates, risk-based human review, and queue metrics
11How do you record change provenance and route PRs by risk?Change provenance and risk routing · The evidence bundlePR metadata on intent, affected systems, data, and reversibility routes approvals; authorship is kept for measurement
12What is the E2E requirement for UI changes?E2E policy · Agent-driven browser verificationA browser run against acceptance criteria with retained artifacts; critical paths promoted to deterministic CI tests
13Are agents part of CI/CD itself?AI in CI/CD · From issue to pull requestThe agent prepares, reviews, and tests in the pipeline; deterministic checks gate the PR; a named human keeps merge and production authority

Parallelism at team scale (questions 14–16, 9 points)

Section titled “Parallelism at team scale (questions 14–16, 9 points)”
#QuestionReadEvidence for 3 points
14Does the team run parallel agent tasks in isolated worktrees?Team parallelism · Parallel agents in worktreesIsolated state per task, a measured concurrency cap, an integration owner, and a reviewer-queue limit
15Do you run unattended agent tasks on a schedule or event trigger?Unattended agent runs · Shaping a backlog for agentsA curated low-risk backlog, bounded attempts, retained evidence, and review before merge or any external action
16How does the team classify and graduate rapidly generated prototypes?Prototype policy · Governance and autonomyGraduation criteria per risk class: data, ownership, tests, security, observability, and rollback

Org enablement (questions 17–21, 15 points)

Section titled “Org enablement (questions 17–21, 15 points)”
#QuestionReadEvidence for 3 points
17Do sensitive changes require accepted intent, specification, and plan artifacts before execution?Plan policy · Plan: capture intent.mdAccepted intent, spec, and plan before write access; a hook or CI job verifies them; a named approver controls risky actions
18How do you govern shared agent hooks and their per-tool adapters?Shared hooks governance · Managed policyVersioned, reviewed, least-privilege hooks with fixtures, staged rollout, audit logs, and a tested rollback
19How long until a new engineer has a productive AI workflow?Developer onboarding · The developer trackAn onboarding fixture that passes within one working day, with no shared secrets
20How does the team share AI tooling know-how?Knowledge sharing · Keeping a team currentVersioned patterns with owners and evidence, updated from reviews and incidents
21Do you govern data and compliance for each approved AI service and workflow?Data privacy and enterprise policies · Security standards and complianceA data-flow review per service, with minimization, retention, audit, incident handling, legal mapping, and re-approval

Strategy and ROI (questions 22–25, 12 points)

Section titled “Strategy and ROI (questions 22–25, 12 points)”
#QuestionReadEvidence for 3 points
22What do you measure for AI tooling?AI metrics panel · Metrics frameworksThree or more of spend, throughput, quality, adoption, review-to-merge time, and cost per feature, each reviewed on a schedule
23Do you have numerical ROI from AI tooling?Cost per accepted change · AI tooling ROIMatched baseline and current cohorts on accepted lead time, quality, outcome, and total cost, with uncertainty recorded
24Do you have a 6–12 month AI tooling roadmap?The organization-wide roadmap · Roadmap portfolio itemAn owned roadmap of capability hypotheses with graduation and stop gates, a quarterly review, and a budget
25How do you manage vendor risk?Vendor risk management · Avoiding lock-inA tested fallback and degraded mode, exportable artifacts, contract and data-location review, and a recovery exercise

How do you prove a higher answer without reading every diff?

Section titled “How do you prove a higher answer without reading every diff?”

Every 3-point answer above names an artifact, not an intention: a branch-protection export, an audit log, an evidence bundle on merged PRs, a cost report, a recovery-exercise record. Claim the higher answer only when you can link that artifact. The owner in the plan file produces it, and a second person, such as the platform lead or security, checks it before you re-take the scorecard. For quality gates, ask for the evidence bundle; reading evidence instead of code explains why it replaces line-by-line review.

An agent can collect the evidence for you if it runs read-only. Do not rely on the prompt alone to keep it read-only: start the agent in the mode each tool provides for that, from the repository whose gates you are scoring.

Start the session in plan mode, which explores and plans without editing files: claude --permission-mode plan (checked in claude --help, v2.1.283). Paste the prompt below.

The prompt text itself is the same in all three tools.

Branch protection and required reviewers live in your Git host’s settings, not in the repository, so export them separately and attach them to the same plan entry.

  • You scored intentions. A policy document that nobody enforces earns the same answer as an enforced gate if you let it. Recovery: re-score each answer you cannot back with a linked artifact one point lower. The drop is your real starting point.
  • You answered for your best team. One platform team with evidence-gated pipelines does not make ten product teams Level 3. Recovery: answer for the median team.
  • Question 22 became a checkbox count. Tracking six measures nobody reviews scores 3 points and informs no decision. Recovery: keep only measures tied to a decision they drove last quarter, defined as in metrics frameworks.
  • Section 4 rose before Section 3. More parallel and overnight agent runs, the same review capacity: the queue grows and reviewers start approving without reading the evidence. Recovery: cap concurrency, fix questions 9 and 10 first, then raise the cap.
  • The band became a target. Treat a jump from Level 2 to Level 3 in one quarter as a reason to re-check each changed answer against its artifact. Recovery: report section shares with their artifacts to leadership, not the band. Board reporting shows the format.

Where to go next from your CTO Scorecard result

Section titled “Where to go next from your CTO Scorecard result”