CTO Scorecard answer key
The CTO Scorecard answer key explains the four Level 1–4 bands the CTO Scorecard reports and maps each of its 25 questions to the page that moves that answer. It is written for a CTO or VP Engineering who has a result and needs to decide which controls to fix first, and what evidence proves the fix.
Your result says Level 2 · Paired, 34 points out of 75, and the budget review is in six weeks. Every section shows a gap, and 25 per-question pages is too much reading before Monday. This key picks the first three questions, points to where each is answered, and names the artifact that earns the higher answer.
What does each Level 1–4 band mean?
Section titled “What does each Level 1–4 band mean?”The bands borrow the names of the autonomy ladder but measure something else: the controls around agent use, as you report them. The ladder measures how each loop ships merged change, and one map turns that into an organization-level distribution. Use it to check your self-reported band against what reaches production. Answer for your median team, not your best one.
| Band | Points | What it describes | First move |
|---|---|---|---|
| Level 1 · Assisted | 0–18 | Engineers use AI in spite of the organization: unapproved tools, no audit trail, spend nobody can see. The top risks are a compliance incident and runaway spend. | Section 1: inventory who uses what, move to centrally governed accounts, and turn on per-person spend visibility. |
| Level 2 · Paired | 19–37 | Tools are standardized and basic review is in place; each developer works with one agent. | Section 2: shared agent rules, shared skills, and governed team accounts. |
| Level 3 · Review manager | 38–56 | Layered review, internal MCP servers, and metrics exist. Agents write, engineers review. | Sections 4 and 6: parallel work with concurrency caps, and a funded roadmap with stop gates. |
| Level 4 · Spec manager | 57–75 | Accepted intent and evidence drive delivery, approval authority follows risk, capability bets have stop gates, and vendor continuity is exercised. | Improve the loop you have. The scorecard has no Level 5 band; running the software factory is a per-loop decision, not a score target. |
How do you pick the first three questions to fix?
Section titled “How do you pick the first three questions to fix?”- Divide each section’s points by its maximum. A 6 out of 15 in Section 3 (40%) is a bigger gap than 6 out of 9 in Section 4 (67%), although the raw points are equal.
- Take the two sections with the lowest share. In each, pick the question where your answer is furthest from 3 points.
- If Section 3 is below half and not already one of the two, add its weakest question; otherwise take the weakest question from the third-lowest section. Parallel and unattended agent work (Section 4) multiplies the pull requests your gates must absorb. Raising Section 4 before Section 3 moves the bottleneck into the review queue.
- For each question, write the target answer and the artifact that will prove it in the plan file below. Name one owner per question.
- Re-take the scorecard after 30 days. Compare section shares, not the band: one question moving from 1 to 3 points rarely changes the band, but it shows in its section.
Commit the plan to the repository that holds your agent platform configuration, so changes to it are reviewed like any other change:
# ai-scorecard/plan.yaml — one entry per question you are working on- question: 9 # What automatically reviews your PRs today? current_points: 1 target_points: 3 target_answer: "Deterministic checks, a focused risk review, reproducible evidence, and a named human gate by impact" evidence: "Branch protection export listing required checks; three merged agent PRs with an evidence bundle attached" owner: "platform-lead" check_on: 2026-10-26Which page answers each question?
Section titled “Which page answers each question?”Each table gives the question, the page to read first and a second one, and the evidence a 3-point answer needs. Questions 4, 10, and 21 link their topic page directly, because their per-question pages are being merged into it.
Adoption and spend (questions 1–4, 12 points)
Section titled “Adoption and spend (questions 1–4, 12 points)”| # | Question | Read | Evidence for 3 points |
|---|---|---|---|
| 1 | How broadly and effectively does the team use approved AI workflows? | Team adoption · Team onboarding and adoption | Completed approved workflows across eligible engineers, with accepted-outcome, quality, and control data |
| 2 | What is the tooling policy? | Tooling policy · Managed policy | An approved workflow matrix and exceptions with an expiry date and a cleanup record |
| 3 | How do you govern AI accounts, access lifecycle, data terms, and spend? | Team accounts · Procurement questionnaire | A named owner per service, joiner/mover/leaver controls that run, reviewed data terms, auditable spend |
| 4 | Can you connect total AI workflow cost to accepted outcomes and quality? | AI usage cost governance · Cost per accepted change | Total cost per accepted outcome, with budgets, alerts, and a named owner |
Shared infrastructure (questions 5–8, 12 points)
Section titled “Shared infrastructure (questions 5–8, 12 points)”| # | Question | Read | Evidence for 3 points |
|---|---|---|---|
| 5 | Does the team share agent rules? | Shared agent rules · Team collaboration | A versioned shared core with tested per-tool adapters, and hard requirements enforced outside the model |
| 6 | Does the team share skills? | Shared skills · Building your own skills | Each skill has an owner, positive, negative, and boundary fixtures, and usage data |
| 7 | Do internal MCP services solve measured, recurring access gaps safely? | Internal MCP servers · Introduction to MCP | Each server traces to a measured gap and has an owner, least privilege, monitoring, and retirement criteria |
| 8 | What is your MCP security model? | MCP security · Agent identity and secrets | A reviewed allowlist, per-tool authorization, short-lived credentials, audit logs, and explicit approval for writes |
Quality gates (questions 9–13, 15 points)
Section titled “Quality gates (questions 9–13, 15 points)”| # | Question | Read | Evidence for 3 points |
|---|---|---|---|
| 9 | What automatically reviews your PRs today? | Team PR review automation · Reviewing an agent’s PR | Required deterministic checks, a focused risk review, reproducible evidence, and a named human gate sized to impact |
| 10 | How do you handle human review bandwidth for agent-authored PRs? | Governance and autonomy · The review queue | Design-time boundaries, automated evidence gates, risk-based human review, and queue metrics |
| 11 | How do you record change provenance and route PRs by risk? | Change provenance and risk routing · The evidence bundle | PR metadata on intent, affected systems, data, and reversibility routes approvals; authorship is kept for measurement |
| 12 | What is the E2E requirement for UI changes? | E2E policy · Agent-driven browser verification | A browser run against acceptance criteria with retained artifacts; critical paths promoted to deterministic CI tests |
| 13 | Are agents part of CI/CD itself? | AI in CI/CD · From issue to pull request | The agent prepares, reviews, and tests in the pipeline; deterministic checks gate the PR; a named human keeps merge and production authority |
Parallelism at team scale (questions 14–16, 9 points)
Section titled “Parallelism at team scale (questions 14–16, 9 points)”| # | Question | Read | Evidence for 3 points |
|---|---|---|---|
| 14 | Does the team run parallel agent tasks in isolated worktrees? | Team parallelism · Parallel agents in worktrees | Isolated state per task, a measured concurrency cap, an integration owner, and a reviewer-queue limit |
| 15 | Do you run unattended agent tasks on a schedule or event trigger? | Unattended agent runs · Shaping a backlog for agents | A curated low-risk backlog, bounded attempts, retained evidence, and review before merge or any external action |
| 16 | How does the team classify and graduate rapidly generated prototypes? | Prototype policy · Governance and autonomy | Graduation criteria per risk class: data, ownership, tests, security, observability, and rollback |
Org enablement (questions 17–21, 15 points)
Section titled “Org enablement (questions 17–21, 15 points)”| # | Question | Read | Evidence for 3 points |
|---|---|---|---|
| 17 | Do sensitive changes require accepted intent, specification, and plan artifacts before execution? | Plan policy · Plan: capture intent.md | Accepted intent, spec, and plan before write access; a hook or CI job verifies them; a named approver controls risky actions |
| 18 | How do you govern shared agent hooks and their per-tool adapters? | Shared hooks governance · Managed policy | Versioned, reviewed, least-privilege hooks with fixtures, staged rollout, audit logs, and a tested rollback |
| 19 | How long until a new engineer has a productive AI workflow? | Developer onboarding · The developer track | An onboarding fixture that passes within one working day, with no shared secrets |
| 20 | How does the team share AI tooling know-how? | Knowledge sharing · Keeping a team current | Versioned patterns with owners and evidence, updated from reviews and incidents |
| 21 | Do you govern data and compliance for each approved AI service and workflow? | Data privacy and enterprise policies · Security standards and compliance | A data-flow review per service, with minimization, retention, audit, incident handling, legal mapping, and re-approval |
Strategy and ROI (questions 22–25, 12 points)
Section titled “Strategy and ROI (questions 22–25, 12 points)”| # | Question | Read | Evidence for 3 points |
|---|---|---|---|
| 22 | What do you measure for AI tooling? | AI metrics panel · Metrics frameworks | Three or more of spend, throughput, quality, adoption, review-to-merge time, and cost per feature, each reviewed on a schedule |
| 23 | Do you have numerical ROI from AI tooling? | Cost per accepted change · AI tooling ROI | Matched baseline and current cohorts on accepted lead time, quality, outcome, and total cost, with uncertainty recorded |
| 24 | Do you have a 6–12 month AI tooling roadmap? | The organization-wide roadmap · Roadmap portfolio item | An owned roadmap of capability hypotheses with graduation and stop gates, a quarterly review, and a budget |
| 25 | How do you manage vendor risk? | Vendor risk management · Avoiding lock-in | A tested fallback and degraded mode, exportable artifacts, contract and data-location review, and a recovery exercise |
How do you prove a higher answer without reading every diff?
Section titled “How do you prove a higher answer without reading every diff?”Every 3-point answer above names an artifact, not an intention: a branch-protection export, an audit log, an evidence bundle on merged PRs, a cost report, a recovery-exercise record. Claim the higher answer only when you can link that artifact. The owner in the plan file produces it, and a second person, such as the platform lead or security, checks it before you re-take the scorecard. For quality gates, ask for the evidence bundle; reading evidence instead of code explains why it replaces line-by-line review.
An agent can collect the evidence for you if it runs read-only. Do not rely on the prompt alone to keep it read-only: start the agent in the mode each tool provides for that, from the repository whose gates you are scoring.
Start the session in plan mode, which explores and plans without editing files: claude --permission-mode plan (checked in claude --help, v2.1.283). Paste the prompt below.
Run the prompt non-interactively in the read-only sandbox: codex exec --sandbox read-only "<prompt>" (checked in codex exec --help, v0.157.1).
Switch the agent to Plan Mode, which produces a plan before writing any code, then paste the prompt below.
The prompt text itself is the same in all three tools.
Branch protection and required reviewers live in your Git host’s settings, not in the repository, so export them separately and attach them to the same plan entry.
When does the scorecard result mislead?
Section titled “When does the scorecard result mislead?”- You scored intentions. A policy document that nobody enforces earns the same answer as an enforced gate if you let it. Recovery: re-score each answer you cannot back with a linked artifact one point lower. The drop is your real starting point.
- You answered for your best team. One platform team with evidence-gated pipelines does not make ten product teams Level 3. Recovery: answer for the median team.
- Question 22 became a checkbox count. Tracking six measures nobody reviews scores 3 points and informs no decision. Recovery: keep only measures tied to a decision they drove last quarter, defined as in metrics frameworks.
- Section 4 rose before Section 3. More parallel and overnight agent runs, the same review capacity: the queue grows and reviewers start approving without reading the evidence. Recovery: cap concurrency, fix questions 9 and 10 first, then raise the cap.
- The band became a target. Treat a jump from Level 2 to Level 3 in one quarter as a reason to re-check each changed answer against its artifact. Recovery: report section shares with their artifacts to leadership, not the band. Board reporting shows the format.