Skip to content

AI metrics panel — measure accepted delivery

An AI metrics panel for CTO Scorecard Q22 earns its full three points from any three of the six options, but only a quality measure, a flow measure and a cost measure, reported for the same work, answer the real question. Throughput and adoption belong on the panel as context, never as proof that delivery improved.

This page is for the CTO or VP Engineering who answered Q22 (“What do you measure for AI tooling?”) and for the tech lead who has to produce those numbers. The usual situation: the dashboard shows seat spend, pull requests per developer and active users, all three up and to the right. Then an incident review finds two agent-written changes that nobody read closely, and nobody can say whether delivery improved or only output did. This page shows which ticks you have, what each one lacks, and how to turn them into a panel that survives that question. The canonical metric definitions live in metrics frameworks; this page does not repeat them.

  • A table mapping each of the six Q22 options to the canonical metric it should become, and the metric it must be paired with.
  • The three-tick combination to report first, with the evidence for why throughput alone misleads.
  • A metric card template you can commit to a repository as the panel’s definition file.
  • Where the spend and usage numbers come from in Claude Code, Codex and Cursor.
  • Two copy-paste prompts, a verification routine and the failure modes that make a panel lie.

What each Q22 tick is worth, and what it lacks

Section titled “What each Q22 tick is worth, and what it lacks”

Q22 is a multi-select question in the Strategy and ROI section of the CTO Scorecard: each option scores one point, capped at three. The score cannot tell a balanced panel from three activity counts. The table can. Find the options you ticked and read across.

Q22 optionPointsWhat it says on its ownTurn it intoReport it beside
Spend — $/dev/month1What the seats cost, not what they boughtCost per accepted change, with all cost terms from the economics methodAccepted change rate
Throughput — PR/dev/week1How much agents produced; it rises whether or not value doesAccepted change rate, per repository and cohort, never per engineer14-day follow-up fix rate
Quality — bug regression rate, revert rate1Whether faster work is also holding upAgent cohort change fail rate, 14-day follow-up fix rate, escaped defects per 100 merged changesAny speed metric on the same view
Adoption — active-user %, average session count1Whether people use the tools; session count is pure activityWeekly active agent users; drop session countDeveloper trust score
Review-to-merge time1Where agent output waits for humansReview load: median and p90 time to first review and time in reviewReview share per reviewer
Cost-per-feature — token cost per feature ticket1A partial unit cost that leaves out review, rework and CICost per accepted changeChange fail rate

“Accepted” is the word that matters. A change counts as delivered when it merged and was not reverted or fixed within 14 days, which is the window metrics frameworks fixes for the whole site. A fast draft that waits three days for review, or comes back as a revert, did not make delivery faster.

Which three Q22 ticks make a panel that works?

Section titled “Which three Q22 ticks make a panel that works?”

Report quality, review-to-merge time and cost-per-feature first, upgraded as the table says. Together they answer three questions a budget holder asks: did it hold up, where does it wait, and what did each accepted change cost. Add throughput and adoption later as context rows, each paired with its guardrail.

The reason to lead with quality is that throughput and quality can move in opposite directions in the same organization. Faros AI’s AI Engineering Report 2026 (April 2026), built on two years of telemetry from 22,000 developers, reports task throughput per developer up 33.7%. The same data shows bugs per developer up 54% and median time in review up 441.5%. That is vendor telemetry from a self-selected customer base, not a controlled study, but it shows exactly the pattern a throughput-only panel cannot see. DX reports the same pressure on review from another angle: median pull request size grew from 44 to 72 lines between July 2025 and June 2026 (DX, 2026-06-17).

The 2025 DORA report puts the general rule in one line: “AI doesn’t fix a team; it amplifies what’s already there” (Google Cloud, 2025-09-23). A panel exists to show which of the two happened to your teams.

How do you stand up the panel in four weeks?

Section titled “How do you stand up the panel in four weeks?”

The baseline, the agent-assisted marker and the telemetry setup are described step by step in metrics frameworks. This is the order to do them in for Q22.

  1. Week 1: write the decision first. One sentence with a date: “By 2026-12-31, decide whether to extend agent seats from four teams to all twelve.” Every metric on the panel must be able to change that decision; delete any that cannot.

  2. Week 1: freeze three metric cards. Commit one card per metric, using the template below, to the repository that holds your engineering reporting. A definition change after the period starts invalidates the comparison.

  3. Week 2: mark agent-assisted changes. Use the tool’s own attribution where it exists and a pull request template checkbox that adds an agent-assisted label everywhere else, so every metric splits into two cohorts.

  4. Week 3: reconstruct the baseline. Pull at least one quarter of history from git, the code host and the deploy log for the same repositories. Label a reconstructed baseline as reconstructed.

  5. Week 4: write the decision rule before the first report. For example: “Expand if p90 time in review does not rise by more than 20% and the agent cohort’s 14-day follow-up fix rate stays within two percentage points of the other cohort.” The thresholds are yours to set; the point is that they are dated before anyone sees a result.

  6. Week 4 onward: review monthly and annotate. Record every staffing change, new gate, model switch and release freeze on the panel’s timeline, so a shift can be traced to its cause.

One YAML file per panel metric keeps the definition, owner and decision rule in version control, where a change shows up in review like any other change.

# reporting/ai-panel/review-load.yaml (one file per panel metric)
id: review-load
q22_option: review-time
name: Time to first review and time in review
definition: >
Median and p90 hours from pull request opened to first review, and from
first review to merge, for merged pull requests in the period.
unit: hours
source: code host pull request events, read with a read-only token
cohorts: [agent-assisted, other]
segment_by: [repository, risk_class]
paired_with: follow-up-fix-rate-14d
baseline: 2026-Q2, reconstructed from history
cadence: monthly
definition_owner: VP Engineering
pipeline_owner: platform team
missing_data: shown as a gap, never as zero
never_reported: per engineer
decision_rule: >
Expand agent seats if p90 time in review rises by no more than 20% while the
agent cohort's 14-day follow-up fix rate stays within 2 percentage points of
the other cohort.

Copy it for the other two metrics and change id, q22_option, definition, source and paired_with. The segment_by: risk_class line is what stops a documentation change being averaged with a payment migration.

Where do the spend and usage numbers come from in each tool?

Section titled “Where do the spend and usage numbers come from in each tool?”

The quality and review-time ticks come from your code host, CI and incident tracker, identically for every tool. Only the spend, cost and adoption inputs differ. Full telemetry setup lives in metrics frameworks and agent observability.

  • Per developer: /usage in a session shows session cost, plan limits and activity.
  • Per organization: set CLAUDE_CODE_ENABLE_TELEMETRY=1 and the OTEL_* exporter variables in managed settings, then read claude_code.cost.usage (USD) from your collector for the cost card.
  • Cohorts: on Team and Enterprise plans with the Claude GitHub app and GitHub analytics turned on, merged pull requests with attributed lines carry the claude-code-assisted label. Contribution metrics are in public beta and unavailable with Zero Data Retention, so keep the template checkbox as a fallback.

For a quick cross-tool spend estimate from one machine, npx ccusage@latest monthly --json reads local Claude Code, Codex and other agent logs (ccusage 20.0.26 on npm, 2026-09-26). It prices tokens at API rates, so on a subscription it estimates usage value, not the invoice. Reconcile the cost card against what finance paid.

How do you verify the panel without trusting the dashboard?

Section titled “How do you verify the panel without trusting the dashboard?”

Nobody should have to inspect every row to trust the panel. It holds when it passes these checks each month.

  • Hand-check a sample. Pick 10 merged pull requests at random and verify the cohort label, the review times and any revert or fix link by hand. More than one disagreement in 10 means the extraction logic is wrong.
  • Reconcile cost. The cost card matches what finance paid for seats, usage and CI. A gap usually means an API key billed outside the tools’ own reports.
  • Check the pairing. No speed or throughput number appears on a view without its quality partner. A slide that breaks this rule does not ship.
  • Show missing data. A week with no data is drawn as a gap. A panel that fills gaps with zero shows improvement that never happened.
  • Name the sign-off. The VP Engineering signs the panel monthly; the definition owner approves any change to a metric card through a pull request.

Pull requests per developer becomes a target. Agents make the count cheap to inflate, and review queues grow behind it. Recovery: remove throughput from every goal and performance review, replace it with accepted change rate per repository, and state in the panel’s charter that no metric is used for individual performance.

Only successful agent runs are counted. The cohort leaves out abandoned sessions, rejected pull requests and reverts, so the agent cohort looks better than it is. Recovery: count from pull requests opened, not merged, and include every revert in the denominator.

Cost-per-feature counts tokens only. Token cost per ticket looks small because review, rework and CI are missing. Recovery: move to cost per accepted change with the cost terms from economics, and report token cost as one line inside it.

The early adopters are the comparison group. Teams that adopt agents first are often the strongest teams, so their numbers were better before the tools arrived. Recovery: compare each team with its own baseline, record the confounder, and run a matched-cohort pilot before you claim causation.

The panel stops at the dashboard. Numbers are reviewed and nothing changes. Recovery: attach the dated decision rule to each card, and review the panel only in the meeting that owns the decision it serves.

Q22 supplies the measurements; Q23 turns them into an ROI answer, and the board kit turns that into a quarterly report.