Skip to content

Team adoption — measure effective workflow use

Team adoption, as CTO Scorecard Q1 scores it, is the share of eligible engineers who complete an approved AI workflow through to its evidence boundary, paired with accepted-outcome, quality and control measures. Seats bought, logins and prompt counts measure activity, not adoption. The maximum score needs completion, accepted-outcome, quality and control evidence measured against a baseline.

This page is for the tech lead who runs the adoption review and the CTO who answers Q1. The situation: finance asks why 120 seats cost what they do, the admin dashboard says 87% of engineers opened the tool last week, and nobody can say whether a single workflow got faster or safer because of it. “87% active” earns no points on its own.

  • The Q1 scoring table, with the evidence each answer needs.
  • An adoption card template that defines eligibility, completion and guardrails for one workflow, ready to fill in.
  • A four-layer adoption funnel with a formula and a data source for every layer.
  • Where each layer’s data comes from in Claude Code, Codex and Cursor.
  • A friction taxonomy that turns failed runs into the smallest fix.
  • Three copy-paste prompts, a decision rule, and the failure modes that turn adoption into theater.

How does CTO Scorecard Q1 score team adoption?

Section titled “How does CTO Scorecard Q1 score team adoption?”

The scorecard asks “How broadly and effectively does the engineering team use approved AI workflows?” and scores four answers.

PointsAnswerEvidence that earns it
0I don’t knowNone.
1Unmeasured experiments by a few early adoptersAnecdotes, a Slack channel, individual expense claims.
2Measured use of approved workflows, but no accepted-outcome evidence yetA dashboard of sessions or active users per approved workflow.
3Broad eligible-team use with workflow completion, accepted-outcome, quality and control evidenceAn adoption card per workflow, a completion rate against a stable denominator, accepted-change and follow-up-fix rates against a baseline, and control events reviewed.

The jump from 2 to 3 is the one that matters. Level 2 counts people who started; level 3 shows that the work they finished was accepted and did not cost quality or control.

Why do seat counts and usage percentages mislead?

Section titled “Why do seat counts and usage percentages mislead?”

Usage alone no longer separates teams. DORA’s 2025 report found that “90% of survey respondents report using AI at work” (Google Cloud, 2025-09-23). A high active-user share is the industry norm, not an achievement.

DORA’s central finding from the same report explains why usage says little about value: “AI doesn’t fix a team; it amplifies what’s already there.” A team with weak tests and slow review that adopts agents widely gets more unreviewed change, not more delivery. Adoption becomes meaningful only when it is joined to the outcome the workflow exists to produce.

A usage metric also has no eligibility filter. An engineer who maintains a safety-critical firmware module and avoids an agent workflow the policy has not approved for that code is behaving correctly. Counting that person as a “non-adopter” pressures people toward unsafe use.

Define adoption for one workflow before you count it

Section titled “Define adoption for one workflow before you count it”

Adoption is measured per workflow, never for “AI” in general. Write one adoption card for each approved workflow and publish it before the comparison period starts. The card fixes the denominator, so the number cannot drift.

The approved tools and workflows come from your tooling policy; see Q2 · Tooling policy.

ADOPTION CARD — repair a failing test with an agent
Workflow owner: Tech lead, payments team
Approved tools: Claude Code, Codex, Cursor (per tooling policy v3)
Eligible people: Engineers on payments and ledger repositories with a managed
account and the 90-minute workflow onboarding completed
Eligible tasks: Failing unit or integration tests in services at risk tier 1–2
Excluded tasks: Tier-3 services (card data path), flaky-test quarantine decisions
Evidence boundary: PR merged with the previously failing test green in CI, the
change reviewed against the evidence bundle, not reverted or
fixed within 14 days
Not completion: Opening the tool; a session with no PR; a PR closed unmerged
Guardrails: 14-day follow-up fix rate, change fail rate, time to first
review, denied tool calls and secret-scanner hits
Baseline: Same task type, same teams, previous 12 weeks
Review cadence: Monthly with the team; quarterly in the CTO panel
Decision owner: VP Engineering

Then compute effective adoption against that card:

effective adoption = eligible engineers with ≥ 1 completion at the evidence boundary in the period
÷ eligible engineers with access and onboarding completed

The evidence boundary is the point where the workflow’s output is accepted by the system, not by the person who ran it: a merged PR with green checks, an accepted plan, a resolved incident record. Choose the boundary your existing tools already record, so completion is counted from the code host and CI rather than from self-report.

Measure the four layers of the adoption funnel

Section titled “Measure the four layers of the adoption funnel”

Each layer answers a different question, and each has a failure the next one catches.

LayerQuestionMetricData source
AccessCan eligible people start safely?Eligible engineers with a managed account, policy accepted and a working setup ÷ eligible engineersIdentity provider, admin console
Activation and completionDo they finish the workflow?Effective adoption (formula above); completions ÷ sessions started for this workflowAgent telemetry, PR labels, CI
OutcomeIs the finished work accepted?Accepted change rate for the workflow’s PRs against the baseline cohortCode host, revert and fix links
GuardrailIs harm controlled?14-day follow-up fix rate, change fail rate, review load, control events (denied actions, blocked commits, scanner hits)Code host, deploy log, hook and scanner logs

Use the canonical definitions of accepted change rate, follow-up fix rate and review load from the AI engineering metrics frameworks page, so the adoption review and the CTO panel report the same numbers. Report every layer per team or per repository, never per engineer.

Where the adoption data comes from in each tool

Section titled “Where the adoption data comes from in each tool”

The outcome and guardrail layers come from your code host, CI and deploy log whatever tool wrote the code. Only activation and the agent-assisted marker differ by tool.

Activation. Claude Code exports OpenTelemetry metrics when CLAUDE_CODE_ENABLE_TELEMETRY=1 is set, including claude_code.session.count, claude_code.pull_request.count and claude_code.commit.count. Set the variables for everyone through managed settings; the full block is on the agent telemetry page.

Marker. On Team and Enterprise plans with the Claude GitHub app and GitHub analytics turned on, merged PRs with attributed Claude Code lines get the claude-code-assisted label. The GitHub search is:pr is:merged label:claude-code-assisted lists them. Contribution metrics are in public beta and are unavailable with Zero Data Retention; use a PR template checkbox there.

The analytics dashboard includes a per-user leaderboard. Keep it out of the adoption review: it ranks people, and Q1 is about the workflow.

Run the adoption review, one workflow at a time

Section titled “Run the adoption review, one workflow at a time”
  1. Pick one workflow that already has an evidence boundary. Good first candidates: repairing a failing test, planning a bounded feature against a spec, or reviewing a PR against its acceptance criteria. Skip workflows whose output nobody records.

  2. Write and publish the adoption card. Agree eligibility and exclusions with the people who do the work. Freeze the card for the comparison period; a change restarts the period.

  3. Reconstruct the baseline. Pull the last 12 weeks of the same task type from the code host: count, accepted change rate, follow-up fix rate and review time. The baseline prompt on the metrics frameworks page writes the extraction script.

  4. Pilot with a matched cohort. Run the workflow on two or three teams while comparable teams continue as before. Record team, service risk tier, sample size and confounders such as a reorg or a release freeze. For cohort choice and sample size, use the AI pilot design page.

  5. Classify every failed or abandoned run. Use the friction table below, and interview five to eight engineers. Aggregate telemetry shows where runs stop; interviews show why.

  6. Fix the most frequent friction class, then re-measure. Change one thing per cycle so the effect stays attributable.

  7. Decide against the rule you wrote in advance. Expand, narrow, retrain or stop, and record the decision, its owner and the next review date on the adoption card.

Why do engineers start the workflow and not finish it?

Section titled “Why do engineers start the workflow and not finish it?”

Most adoption problems are friction, not reluctance. Classify each failed run once, then fix the class with the most runs.

Friction classSignalSmallest fix to test
AccessSetup tickets, sessions that end at authentication, missing repository permissionsA managed-config bundle and a pre-provisioned account per eligible engineer
ContextThe agent edits the wrong module or ignores conventionsA tested AGENTS.md or CLAUDE.md for the repository; see shared agent rules
CapabilityRuns stop at the same step for people new to the workflowA 30-minute worked example on the team’s own repository; see developer onboarding
VerificationPRs open but stall because nobody trusts the resultA required evidence bundle and a failing test that proves the fix
Tool reliabilityTimeouts, rate limits, broken MCP serversPin versions, fix limits or quotas, record incidents against the tool
Review queueCompletion is high but time to first review doublesA review-capacity cap per team; see the review queue page
Unsuitable taskThe task is outside the card’s eligible setTighten the card’s eligible tasks; this is not a failure of the person

Copy-paste prompts for the adoption review

Section titled “Copy-paste prompts for the adoption review”

Run these in Claude Code, Codex or Cursor with the adoption card and exported data in the working directory. They produce drafts; the decision owner signs off.

The thresholds in the last prompt are an example. Set your own before the pilot starts, and write them on the adoption card.

How is the adoption evidence verified without reading every run?

Section titled “How is the adoption evidence verified without reading every run?”

No one reads every session transcript or every PR. The evidence holds up because of how it is collected and sampled.

  • Completion comes from systems, not people. The evidence boundary is a merged PR, a green CI run or a closed incident record. Self-reported completion never enters the numerator.
  • The denominator is frozen. The adoption card fixes eligibility for the whole comparison period, and the report prints the denominator next to every rate.
  • A sample is checked by hand. Each month, the workflow owner opens ten completions at random and confirms each one met the evidence boundary. A mismatch in more than one invalidates the month’s number until the query is fixed.
  • Guardrails are read beside adoption. The report shows follow-up fix rate and control events on the same view as effective adoption, so a rise in one cannot hide a fall in the other.
  • A named owner signs off. The decision owner on the card records the decision and the next review date. Q1 evidence for the scorecard is that signed record plus the last two monthly reports.
  • Telemetry passes a privacy review. Prompts and code are not stored unless the review approved it; see privacy and data handling for agents.

What goes wrong when you measure team adoption?

Section titled “What goes wrong when you measure team adoption?”

Adoption theater. Leadership buys seats for everyone and mandates weekly usage. Activity rises and accepted delivery does not. Recovery: retire the usage target, publish adoption cards for two or three workflows, and report effective adoption only with its guardrails.

The denominator drifts. Contractors, new hires or excluded teams are added mid-period, and the rate jumps without anyone changing behaviour. Recovery: restate the period with the frozen card, and restart the comparison when eligibility has to change.

Completion gets gamed. Engineers split one change into several small PRs to raise completions. Recovery: count completions per task, not per PR, and watch accepted change rate and review load, which fall when splitting is artificial.

Adoption turns into surveillance. A per-person leaderboard reaches a performance review. Engineers stop reporting failed runs, and the friction data dries up. Recovery: remove per-person views from every adoption report, say so publicly, and rebuild trust before you trust the next survey. For the people side of this, see bringing skeptics and senior engineers along.

Telemetry has holes. Zero Data Retention turns off Claude Code contribution metrics, and codex exec has had a reported metrics gap (#12913, now closed) that you should check on your version. The adoption rate quietly undercounts CI and regulated teams. Recovery: fall back to the PR template marker for those teams and note the gap on the report.

Safe avoidance is penalized. A team declines an agent workflow for a restricted service and is flagged as lagging. Recovery: move that task class to the card’s exclusions and credit the team for applying the policy.

Prove that a new engineer can complete the workflow, then feed effective adoption into the CTO panel beside outcome and cost.

For every other question on the scorecard, start from the CTO Scorecard answer key.