Team adoption — measure effective workflow use
Team adoption, as CTO Scorecard Q1 scores it, is the share of eligible engineers who complete an approved AI workflow through to its evidence boundary, paired with accepted-outcome, quality and control measures. Seats bought, logins and prompt counts measure activity, not adoption. The maximum score needs completion, accepted-outcome, quality and control evidence measured against a baseline.
This page is for the tech lead who runs the adoption review and the CTO who answers Q1. The situation: finance asks why 120 seats cost what they do, the admin dashboard says 87% of engineers opened the tool last week, and nobody can say whether a single workflow got faster or safer because of it. “87% active” earns no points on its own.
What you get from this adoption page
Section titled “What you get from this adoption page”- The Q1 scoring table, with the evidence each answer needs.
- An adoption card template that defines eligibility, completion and guardrails for one workflow, ready to fill in.
- A four-layer adoption funnel with a formula and a data source for every layer.
- Where each layer’s data comes from in Claude Code, Codex and Cursor.
- A friction taxonomy that turns failed runs into the smallest fix.
- Three copy-paste prompts, a decision rule, and the failure modes that turn adoption into theater.
How does CTO Scorecard Q1 score team adoption?
Section titled “How does CTO Scorecard Q1 score team adoption?”The scorecard asks “How broadly and effectively does the engineering team use approved AI workflows?” and scores four answers.
| Points | Answer | Evidence that earns it |
|---|---|---|
| 0 | I don’t know | None. |
| 1 | Unmeasured experiments by a few early adopters | Anecdotes, a Slack channel, individual expense claims. |
| 2 | Measured use of approved workflows, but no accepted-outcome evidence yet | A dashboard of sessions or active users per approved workflow. |
| 3 | Broad eligible-team use with workflow completion, accepted-outcome, quality and control evidence | An adoption card per workflow, a completion rate against a stable denominator, accepted-change and follow-up-fix rates against a baseline, and control events reviewed. |
The jump from 2 to 3 is the one that matters. Level 2 counts people who started; level 3 shows that the work they finished was accepted and did not cost quality or control.
Why do seat counts and usage percentages mislead?
Section titled “Why do seat counts and usage percentages mislead?”Usage alone no longer separates teams. DORA’s 2025 report found that “90% of survey respondents report using AI at work” (Google Cloud, 2025-09-23). A high active-user share is the industry norm, not an achievement.
DORA’s central finding from the same report explains why usage says little about value: “AI doesn’t fix a team; it amplifies what’s already there.” A team with weak tests and slow review that adopts agents widely gets more unreviewed change, not more delivery. Adoption becomes meaningful only when it is joined to the outcome the workflow exists to produce.
A usage metric also has no eligibility filter. An engineer who maintains a safety-critical firmware module and avoids an agent workflow the policy has not approved for that code is behaving correctly. Counting that person as a “non-adopter” pressures people toward unsafe use.
Define adoption for one workflow before you count it
Section titled “Define adoption for one workflow before you count it”Adoption is measured per workflow, never for “AI” in general. Write one adoption card for each approved workflow and publish it before the comparison period starts. The card fixes the denominator, so the number cannot drift.
The approved tools and workflows come from your tooling policy; see Q2 · Tooling policy.
ADOPTION CARD — repair a failing test with an agentWorkflow owner: Tech lead, payments teamApproved tools: Claude Code, Codex, Cursor (per tooling policy v3)Eligible people: Engineers on payments and ledger repositories with a managed account and the 90-minute workflow onboarding completedEligible tasks: Failing unit or integration tests in services at risk tier 1–2Excluded tasks: Tier-3 services (card data path), flaky-test quarantine decisionsEvidence boundary: PR merged with the previously failing test green in CI, the change reviewed against the evidence bundle, not reverted or fixed within 14 daysNot completion: Opening the tool; a session with no PR; a PR closed unmergedGuardrails: 14-day follow-up fix rate, change fail rate, time to first review, denied tool calls and secret-scanner hitsBaseline: Same task type, same teams, previous 12 weeksReview cadence: Monthly with the team; quarterly in the CTO panelDecision owner: VP EngineeringThen compute effective adoption against that card:
effective adoption = eligible engineers with ≥ 1 completion at the evidence boundary in the period ÷ eligible engineers with access and onboarding completedThe evidence boundary is the point where the workflow’s output is accepted by the system, not by the person who ran it: a merged PR with green checks, an accepted plan, a resolved incident record. Choose the boundary your existing tools already record, so completion is counted from the code host and CI rather than from self-report.
Measure the four layers of the adoption funnel
Section titled “Measure the four layers of the adoption funnel”Each layer answers a different question, and each has a failure the next one catches.
| Layer | Question | Metric | Data source |
|---|---|---|---|
| Access | Can eligible people start safely? | Eligible engineers with a managed account, policy accepted and a working setup ÷ eligible engineers | Identity provider, admin console |
| Activation and completion | Do they finish the workflow? | Effective adoption (formula above); completions ÷ sessions started for this workflow | Agent telemetry, PR labels, CI |
| Outcome | Is the finished work accepted? | Accepted change rate for the workflow’s PRs against the baseline cohort | Code host, revert and fix links |
| Guardrail | Is harm controlled? | 14-day follow-up fix rate, change fail rate, review load, control events (denied actions, blocked commits, scanner hits) | Code host, deploy log, hook and scanner logs |
Use the canonical definitions of accepted change rate, follow-up fix rate and review load from the AI engineering metrics frameworks page, so the adoption review and the CTO panel report the same numbers. Report every layer per team or per repository, never per engineer.
Where the adoption data comes from in each tool
Section titled “Where the adoption data comes from in each tool”The outcome and guardrail layers come from your code host, CI and deploy log whatever tool wrote the code. Only activation and the agent-assisted marker differ by tool.
Activation. Claude Code exports OpenTelemetry metrics when CLAUDE_CODE_ENABLE_TELEMETRY=1 is set, including claude_code.session.count, claude_code.pull_request.count and claude_code.commit.count. Set the variables for everyone through managed settings; the full block is on the agent telemetry page.
Marker. On Team and Enterprise plans with the Claude GitHub app and GitHub analytics turned on, merged PRs with attributed Claude Code lines get the claude-code-assisted label. The GitHub search is:pr is:merged label:claude-code-assisted lists them. Contribution metrics are in public beta and are unavailable with Zero Data Retention; use a PR template checkbox there.
The analytics dashboard includes a per-user leaderboard. Keep it out of the adoption review: it ranks people, and Q1 is about the workflow.
Activation. Codex exports logs, traces and metrics through an [otel] table in ~/.codex/config.toml. Keep log_user_prompt = false unless your privacy review approved storing prompts. An issue in openai/codex (#12913, now closed) reported that codex exec exported traces and logs but no metrics. Before you count CI runs, confirm they reach your collector on your Codex version.
Marker. Codex CLI 0.157.1 adds no PR label of its own (checked 2026-09-26). Use the PR template checkbox and an agent-assisted label, as set out in AI PR labeling, and have the CI job that runs codex exec push to branches named by a rule you can query (for example agent/codex/<ticket>).
Activation. Cursor documents a team analytics dashboard and an Analytics API for admins; check current plan limits and endpoints in Cursor’s admin docs before you build a report on them.
Marker. Use the PR template checkbox and an agent-assisted label. It works on every plan and every tool, so it is the one marker you can compare across teams.
Run the adoption review, one workflow at a time
Section titled “Run the adoption review, one workflow at a time”-
Pick one workflow that already has an evidence boundary. Good first candidates: repairing a failing test, planning a bounded feature against a spec, or reviewing a PR against its acceptance criteria. Skip workflows whose output nobody records.
-
Write and publish the adoption card. Agree eligibility and exclusions with the people who do the work. Freeze the card for the comparison period; a change restarts the period.
-
Reconstruct the baseline. Pull the last 12 weeks of the same task type from the code host: count, accepted change rate, follow-up fix rate and review time. The baseline prompt on the metrics frameworks page writes the extraction script.
-
Pilot with a matched cohort. Run the workflow on two or three teams while comparable teams continue as before. Record team, service risk tier, sample size and confounders such as a reorg or a release freeze. For cohort choice and sample size, use the AI pilot design page.
-
Classify every failed or abandoned run. Use the friction table below, and interview five to eight engineers. Aggregate telemetry shows where runs stop; interviews show why.
-
Fix the most frequent friction class, then re-measure. Change one thing per cycle so the effect stays attributable.
-
Decide against the rule you wrote in advance. Expand, narrow, retrain or stop, and record the decision, its owner and the next review date on the adoption card.
Why do engineers start the workflow and not finish it?
Section titled “Why do engineers start the workflow and not finish it?”Most adoption problems are friction, not reluctance. Classify each failed run once, then fix the class with the most runs.
| Friction class | Signal | Smallest fix to test |
|---|---|---|
| Access | Setup tickets, sessions that end at authentication, missing repository permissions | A managed-config bundle and a pre-provisioned account per eligible engineer |
| Context | The agent edits the wrong module or ignores conventions | A tested AGENTS.md or CLAUDE.md for the repository; see shared agent rules |
| Capability | Runs stop at the same step for people new to the workflow | A 30-minute worked example on the team’s own repository; see developer onboarding |
| Verification | PRs open but stall because nobody trusts the result | A required evidence bundle and a failing test that proves the fix |
| Tool reliability | Timeouts, rate limits, broken MCP servers | Pin versions, fix limits or quotas, record incidents against the tool |
| Review queue | Completion is high but time to first review doubles | A review-capacity cap per team; see the review queue page |
| Unsuitable task | The task is outside the card’s eligible set | Tighten the card’s eligible tasks; this is not a failure of the person |
Copy-paste prompts for the adoption review
Section titled “Copy-paste prompts for the adoption review”Run these in Claude Code, Codex or Cursor with the adoption card and exported data in the working directory. They produce drafts; the decision owner signs off.
The thresholds in the last prompt are an example. Set your own before the pilot starts, and write them on the adoption card.
How is the adoption evidence verified without reading every run?
Section titled “How is the adoption evidence verified without reading every run?”No one reads every session transcript or every PR. The evidence holds up because of how it is collected and sampled.
- Completion comes from systems, not people. The evidence boundary is a merged PR, a green CI run or a closed incident record. Self-reported completion never enters the numerator.
- The denominator is frozen. The adoption card fixes eligibility for the whole comparison period, and the report prints the denominator next to every rate.
- A sample is checked by hand. Each month, the workflow owner opens ten completions at random and confirms each one met the evidence boundary. A mismatch in more than one invalidates the month’s number until the query is fixed.
- Guardrails are read beside adoption. The report shows follow-up fix rate and control events on the same view as effective adoption, so a rise in one cannot hide a fall in the other.
- A named owner signs off. The decision owner on the card records the decision and the next review date. Q1 evidence for the scorecard is that signed record plus the last two monthly reports.
- Telemetry passes a privacy review. Prompts and code are not stored unless the review approved it; see privacy and data handling for agents.
What goes wrong when you measure team adoption?
Section titled “What goes wrong when you measure team adoption?”Adoption theater. Leadership buys seats for everyone and mandates weekly usage. Activity rises and accepted delivery does not. Recovery: retire the usage target, publish adoption cards for two or three workflows, and report effective adoption only with its guardrails.
The denominator drifts. Contractors, new hires or excluded teams are added mid-period, and the rate jumps without anyone changing behaviour. Recovery: restate the period with the frozen card, and restart the comparison when eligibility has to change.
Completion gets gamed. Engineers split one change into several small PRs to raise completions. Recovery: count completions per task, not per PR, and watch accepted change rate and review load, which fall when splitting is artificial.
Adoption turns into surveillance. A per-person leaderboard reaches a performance review. Engineers stop reporting failed runs, and the friction data dries up. Recovery: remove per-person views from every adoption report, say so publicly, and rebuild trust before you trust the next survey. For the people side of this, see bringing skeptics and senior engineers along.
Telemetry has holes. Zero Data Retention turns off Claude Code contribution metrics, and codex exec has had a reported metrics gap (#12913, now closed) that you should check on your version. The adoption rate quietly undercounts CI and regulated teams. Recovery: fall back to the PR template marker for those teams and note the gap on the report.
Safe avoidance is penalized. A team declines an agent workflow for a restricted service and is flagged as lagging. Recovery: move that task class to the card’s exclusions and credit the team for applying the policy.
Where to go next with team adoption
Section titled “Where to go next with team adoption”Prove that a new engineer can complete the workflow, then feed effective adoption into the CTO panel beside outcome and cost.
For every other question on the scorecard, start from the CTO Scorecard answer key.