DORA, SPACE, DX Core 4 and AI measurement when agents write the code
When agents write the code, outcome metrics survive and activity metrics break. DORA’s change lead time, change fail rate, recovery time and rework rate still measure delivery. Pull request counts, lines of code and ”% AI-written” measure agent output, which is now cheap. Measure AI adoption with utilization, impact and cost, and read every speed number beside a stability number.
Your agents went live across twelve teams last quarter. The dashboard shows merged pull requests up, lines changed up, and a vendor chart saying half your code is now AI-written. In the same quarter, review queues doubled and two incidents traced back to changes nobody read closely. The board asks whether the investment is working, and none of the numbers on your screen can answer that question.
This page is for the CTO or VP Engineering who owns the measurement programme, the tech lead who has to report on one team without turning metrics into a leaderboard, and the executive who needs to know which numbers to trust. It is the site’s canonical home for AI metric definitions: other pages link here instead of redefining them.
What you get from this measurement page
Section titled “What you get from this measurement page”- A verdict table: which established metrics survive agent throughput, which break, and what to use instead.
- Current definitions of DORA’s five software delivery metrics, SPACE’s five dimensions and DX Core 4’s four dimensions, with what agents change in each.
- Ten canonical AI metric definitions, each with a formula, a data source and the trap it avoids.
- A minimal panel for a team, an engineering organization and a board, with a review cadence.
- Instrumentation steps for Claude Code, Codex and Cursor, two copy-paste prompts, and a failure-modes section.
Which metrics survive when agents write the code?
Section titled “Which metrics survive when agents write the code?”A metric survives when it measures an outcome the customer or the operator feels. It breaks when it counts something an agent can now produce in bulk at almost no cost. The table is the short answer; the sections after it give the definitions.
| Metric | Verdict | Why | Use instead or alongside |
|---|---|---|---|
| Change lead time (commit to production) | Survives | Measures the whole system, including review and deploy | Split it: time to first review, time in review, time to deploy |
| Change fail rate | Survives, and becomes the key guardrail | Catches the stability cost of higher change volume | Compare agent-assisted and other changes as cohorts |
| Failed deployment recovery time | Survives | Recovery depends on rollback and observability, not on who wrote the code | Track per service |
| Deployment rework rate | Survives | Unplanned deploys caused by incidents expose low-quality volume | Pair with change fail rate |
| Deployment frequency | Survives with a caveat | More, smaller deploys are good; more deploys of large agent diffs are not | Read with PR size |
| Pull requests merged per engineer | Breaks as a success measure | Agents open pull requests in bulk; the count rises whether or not value does | Accepted change rate, cost per accepted change |
| Lines of code added or changed | Breaks | Agents generate verbose code; more lines is often worse | Nothing: stop reporting it as output |
| ”% of code written by AI” | Breaks as a goal | Describes adoption, not value; it rises by itself | Agent-assisted merged change share, read as utilization only |
| Suggestion acceptance rate | Breaks for agents | Built for autocomplete; an agent session has no single suggestion to accept | Accepted change rate per agent loop |
| Story points per sprint | Breaks | Estimation drifts as soon as agents change the effort | Lead time and throughput of accepted changes |
| Developer satisfaction and trust (survey) | Survives, and matters more | Review load, trust and cognitive load move before delivery metrics do | Quarterly survey with fixed questions |
The evidence for the split is consistent across independent sources:
- DORA 2025 (Google Cloud, 2025-09-23): “we observe a positive relationship between AI adoption on both software delivery throughput and product performance.” And in the next sentence: “AI adoption does continue to have a negative relationship with software delivery stability.” A panel that shows only the first half tells half the story.
- DX (Justin Reock, 2026-06-17; last checked 2026-08-28): the median pull request grew from 44 to 72 lines between July 2025 and June 2026, in a sample of more than 400 companies. Pull request counts and sizes change with agent adoption, which makes them unstable units of work.
- Faros AI, AI Engineering Report 2026 (April 2026; last checked 2026-08-28), telemetry from 22,000 developers: task throughput per developer +33.7% and PR merge rate per developer +16.2%, while deployments per week fell 11.7%, incidents per pull request rose 242.7% and median time in review rose 441.5%. Output went up; delivery and stability did not follow.
What are DORA’s five software delivery metrics now?
Section titled “What are DORA’s five software delivery metrics now?”DORA replaced its original four keys with five metrics, and moved from mean time to restore to failed deployment recovery time. The definitions below are from DORA’s own metrics guide (updated 2026-01-05, read from the dora-team/dora.dev source on 2026-09-26). DORA groups them into throughput and instability.
| Group | Metric | DORA definition | What agents change |
|---|---|---|---|
| Throughput | Change lead time | “The amount of time it takes for a change to go from committed to version control to deployed in production.” | Coding time shrinks; review and approval become most of the lead time |
| Throughput | Deployment frequency | “The number of deployments over a given period or the time between deployments.” | Can fall even as pull requests rise, when review is the bottleneck |
| Throughput | Failed deployment recovery time | “The time it takes to recover from a deployment that fails and requires immediate intervention.” | Unchanged if rollback is automated; worse if nobody understands the change |
| Instability | Change fail rate | “The ratio of deployments that require immediate intervention following a deployment.” | The first metric to move when volume outruns verification |
| Instability | Deployment rework rate | “The ratio of deployments that are unplanned but happen as a result of an incident in production.” | Rises when agent changes pass tests that do not check the behaviour |
Two points from DORA’s guide matter more with agents. First, the metrics are “best suited for measuring one application or service at a time”, so do not average them across a portfolio. Second, DORA lists “Setting metrics as a goal” as the first pitfall: a target such as “every team deploys daily” invites gaming, and an agent games a target faster than a person.
For organizational readiness rather than delivery outcomes, DORA publishes a separate AI Capabilities Model with seven capabilities, including working in small batches and quality internal platforms. It has its own DORA AI capabilities self-assessment.
What does the SPACE framework add?
Section titled “What does the SPACE framework add?”SPACE is a model of developer productivity by Nicole Forsgren, Margaret-Anne Storey, Chandra Maddila, Tom Zimmermann, Brian Houck and Jenna Butler (ACM Queue, February 2021). Its central claim, from the abstract: developer productivity “cannot be measured by a single metric or dimension.” It names five dimensions: Satisfaction and well-being, Performance, Activity, Communication and collaboration, and Efficiency and flow.
SPACE is useful now because it explains why one dimension broke. Agents inflate Activity (commits, pull requests, lines) without moving the others, so any panel built mainly on Activity overstates the gain.
| SPACE dimension | Measure it with agents in the loop | Signal to watch |
|---|---|---|
| Satisfaction and well-being | Quarterly survey: trust in agent output, review fatigue, sense of ownership | Trust falling while adoption rises |
| Performance | Change fail rate, escaped defects, accepted change rate | Stability falling behind throughput |
| Activity | Agent-assisted merged change share, as context only | Any target set on it |
| Communication and collaboration | Time to first review, reviewer load per engineer | Review concentrated on a few senior engineers |
| Efficiency and flow | Time in review, interruptions to wait on agents, rework loops | Engineers waiting on review instead of on agents |
What is DX Core 4, and what breaks in it?
Section titled “What is DX Core 4, and what breaks in it?”DX Core 4 is a framework from DX (Abi Noda and Laura Tacho) that consolidates DORA, SPACE and developer-experience research into four dimensions: speed, effectiveness, quality and impact. DX proposes a key metric per dimension, including diffs (pull requests) per engineer for speed and change failure rate for quality. The design intent is that the four pull against each other, so gaming one shows up in another.
With agents, the speed metric is the weak one. Diffs per engineer rises with adoption whatever happens to value, as the DX and Faros data above show. Keep the four dimensions, and for speed report accepted changes rather than raw diffs: pull requests that merged and were not reverted or fixed within 14 days. Quality (change failure rate) becomes the guardrail that tells you whether speed is real.
How does DX’s AI Measurement Framework fit in?
Section titled “How does DX’s AI Measurement Framework fit in?”DX’s AI Measurement Framework (Abi Noda and Laura Tacho) organizes AI measurement into three dimensions: utilization (is AI being used?), impact (is it working?) and cost (does the return justify the spend?). It is the most useful lens for the AI-specific layer of your panel, because it stops utilization numbers from being presented as impact.
You do not need a new framework to use it. DORA’s research team makes the same point in its guidance on measurement frameworks (2025-08-26): “if your overarching goal is the same, you don’t need to change your framework; you can expand your measurements to adapt to changes in technology.” Keep DORA or Core 4 as the outcome layer, and add the definitions below as the AI layer.
Canonical AI metric definitions
Section titled “Canonical AI metric definitions”These ten definitions are the ones every page on this site uses. Each is computable from git, your code host, CI, the incident tracker and the agents’ own telemetry. Where a formula says “agent-assisted”, it means a merged pull request that carries your agent-assisted marker (see the instrumentation steps below).
| # | Metric | Dimension | Definition and formula | Data source | Trap it avoids |
|---|---|---|---|---|---|
| 1 | Weekly active agent users | Utilization | Engineers with at least one agent session in the week ÷ engineers with an agent seat | Agent telemetry or admin analytics | Paying for seats nobody uses |
| 2 | Agent-assisted merged change share | Utilization | Agent-assisted merged pull requests ÷ all merged pull requests, per repository | Pull request labels | Replaces ”% of lines written by AI”, which counts volume |
| 3 | Accepted change rate | Impact | Agent-assisted pull requests merged and not reverted or fixed within 14 days ÷ agent-assisted pull requests opened | Code host, revert and fix links | Counting abandoned or reworked agent output as delivery |
| 4 | Agent cohort change fail rate | Impact | Change fail rate for deployments containing agent-assisted changes, reported beside the rate for the rest | Deploy log joined to pull requests | A blended rate hiding a worse agent cohort |
| 5 | 14-day follow-up fix rate | Impact | Merged pull requests followed within 14 days by a revert or a fix that references them ÷ merged pull requests | Commit messages, pull request links | Quality debt that never reaches production incidents |
| 6 | Review load | Impact | Median and p90 time to first review and time in review, plus reviews per reviewer per week | Code host | Moving the bottleneck to senior reviewers unnoticed |
| 7 | Escaped defects per 100 merged changes | Impact | Production bugs traced to a change ÷ merged changes × 100, per cohort | Incident and bug tracker | Treating green CI as correctness |
| 8 | Cost per accepted change | Cost | Agent spend (seats, usage and tokens) in the period ÷ accepted agent-assisted changes (metric 3’s numerator) | Billing, agent telemetry | Cost per token or per seat, which says nothing about value |
| 9 | Net time gain | Cost | Survey-reported hours saved per engineer per week minus measured extra review and rework hours | Survey plus code host | Self-reported savings presented as measured |
| 10 | Developer trust score | Utilization | Share answering 4 or 5 on “I trust agent-produced changes that passed our gates” (five-point scale), quarterly | Survey | Adoption forced faster than trust |
Three rules keep these definitions honest:
- Report every speed metric beside a stability metric. Metric 2 or 3 never appears without 4 or 5 on the same slide.
- Compare cohorts, not people. Every impact metric is reported per repository or service, split into agent-assisted and other changes. None is ever reported per engineer.
- Fix the window. 14 days is the default follow-up window for metrics 3 and 5. Change it only for the whole organization and only with a note on the dashboard.
What is the minimal panel for each level of the organization?
Section titled “What is the minimal panel for each level of the organization?”A panel that tries to show everything gets read by nobody. Start with the smallest panel that answers the question each level is actually asking, and add metrics only when a decision needs them.
| Level | Question the panel answers | Minimal panel | Cadence | Owner |
|---|---|---|---|---|
| Team (tech lead) | Is agent work reaching production safely, and where is it stuck? | Accepted change rate, time to first review and time in review, 14-day follow-up fix rate, change fail rate for the team’s services | Weekly, in the team’s retro | Tech lead |
| Engineering organization (CTO, VP Engineering) | Is the delivery system improving, and is quality holding? | DORA’s five metrics per critical service, agent cohort change fail rate, escaped defects per 100 changes, review load, weekly active agent users, developer trust score | Monthly | VP Engineering, with a named metrics owner |
| Board and executive | Is the investment paying back, and is risk under control? | Change lead time trend, change fail rate trend, cost per accepted change, net time gain with its method stated, incidents attributable to agent changes | Quarterly | CTO |
The board panel deliberately has no utilization metric except as context. “Half our code is AI-written” is a statement about adoption; DX’s own figure for the industry is an average of 51.9% AI-authored code across more than 400 companies in Q2 2026 (self-reported, DX, 2026-06-17). A number that is already near the industry average tells a board nothing about returns. For board slide structure, see board reporting on agentic engineering.
How do you stand up the baseline?
Section titled “How do you stand up the baseline?”A metric without a baseline cannot show improvement. Collect at least one full quarter of the outcome layer before you change anything, or reconstruct it from history: git, your code host and your deploy log already hold most of it.
-
Write down the decision the measurement serves. For example: “Decide by 2026-12-31 whether to extend agent seats from four teams to all twelve.” DORA’s guidance calls this deciding on the “why”. Without it, you will collect everything and act on nothing.
-
Reconstruct the outcome baseline. Pull the last six months of DORA’s five metrics per service, plus time to first review and time in review. Use the first copy-paste prompt below to have an agent write the extraction script.
-
Mark agent-assisted changes at the source. Every merged pull request needs a marker you can query. Use the tool’s own attribution where it exists and a pull request template checkbox that adds an
agent-assistedlabel everywhere else. See the tabs below. -
Turn on agent telemetry. Route usage and cost to the same collector as your delivery data, so metrics 1, 8 and 9 come from one place.
-
Define decision rules before the data arrives. For example: “Extend if lead time falls, agent cohort change fail rate stays within two percentage points of the other cohort, and cost per accepted change is below the agreed ceiling.” Writing the rule first stops the panel becoming a search for good news.
-
Review on the cadence in the panel table, and record changes. Note every process change (a new gate, a new model, a reorganization) on the dashboard timeline, so a shift in a metric can be traced to its cause. For a structured trial, use a pilot designed to prove something.
Instrument agent-assisted changes in each tool
Section titled “Instrument agent-assisted changes in each tool”The marker and the telemetry differ by tool. The delivery metrics do not: they come from your code host and deploy log whatever wrote the code.
Attribution. On Team and Enterprise plans with the Claude GitHub app installed and GitHub analytics turned on, Claude Code labels merged pull requests that contain Claude Code-assisted lines as claude-code-assisted. Sessions from 21 days before to 2 days after the merge count toward attribution. The GitHub search is:pr is:merged label:claude-code-assisted lists them, which gives you metric 2 without an API. Contribution metrics are in public beta, and lines a human rewrites by more than 20% are not attributed. They are not available to organizations with Zero Data Retention enabled; use the pull request template checkbox instead.
Telemetry. Claude Code exports OpenTelemetry metrics, including claude_code.session.count, claude_code.pull_request.count, claude_code.commit.count, claude_code.cost.usage (USD) and claude_code.token.usage. Set them for everyone through managed settings:
{ "env": { "CLAUDE_CODE_ENABLE_TELEMETRY": "1", "OTEL_METRICS_EXPORTER": "otlp", "OTEL_LOGS_EXPORTER": "otlp", "OTEL_EXPORTER_OTLP_PROTOCOL": "grpc", "OTEL_EXPORTER_OTLP_ENDPOINT": "http://collector.example.com:4317" }}To check the setup, look for claude_code.session.count in your backend after a session starts (Claude Code monitoring docs, checked 2026-09-26 against v2.1.283). Use claude_code.cost.usage for metric 8, not claude_code.lines_of_code.count for anything on the board panel.
Attribution. Codex has no pull request label that this page could verify on 2026-09-26 (checked against Codex CLI 0.157.1). Use the pull request template checkbox and an agent-assisted label for metric 2, and add a label or a naming rule to branches that codex exec creates in CI. Check that your collector receives data from codex exec runs in CI before you trust metric 1 or 8 for CI usage: an open issue in openai/codex reports that non-interactive runs may not emit metrics.
Telemetry. Codex exports logs, traces and metrics through an [otel] table in ~/.codex/config.toml. The exporters accept none, statsig, otlp-http and otlp-grpc:
[otel]environment = "prod"log_user_prompt = falseexporter = { otlp-http = { endpoint = "https://otel.example.com/v1/logs", protocol = "binary" } }Field names come from OtelConfigToml in the openai/codex source (checked 2026-09-26). Keep log_user_prompt = false unless your privacy review has approved storing prompts. ChatGPT Enterprise workspaces also have a Codex analytics dashboard; confirm its API in OpenAI’s admin docs before you build on it (not verified here on 2026-09-26).
Attribution. Cursor documents a team analytics dashboard and API on Teams and Enterprise, and an AI Code Tracking API that maps AI-generated lines to commits on Enterprise. Check the current plan limits and endpoints in Cursor’s admin docs before you build on them (not verified here on 2026-09-26). The pull request template checkbox and agent-assisted label work regardless of plan.
Telemetry. Use the admin dashboard for metric 1 (active users) and your billing export for metric 8. Line-level AI tracking feeds utilization only: keep it off the impact and board panels for the same reason as ”% AI-written”.
The pull request template checkbox is the one marker all three tools share. A minimal version for .github/pull_request_template.md:
## Agent involvement- [ ] An agent (Claude Code, Codex, Cursor or other) wrote or changed code in this PR- [ ] A human reviewed the evidence (tests, acceptance criteria), not only the diffA small GitHub Action or your merge bot turns the first checkbox into the agent-assisted label, so the metric does not depend on memory.
How do you prove the numbers are true?
Section titled “How do you prove the numbers are true?”A metrics programme needs its own verification, or the dashboard becomes the least-checked artifact in the company.
- Reconcile against the source monthly. Pick ten merged pull requests at random and check the label, the lead time and the follow-up link by hand. A disagreement rate above one in ten means the extraction logic is wrong, not the teams.
- Keep a stability metric on every view. If a slide shows metric 2 or 3 without metric 4 or 5, it does not ship.
- Compare cohorts, and state the confounders. Teams that adopt agents first are often the strongest teams. Record team, service and time period with every comparison, and do not claim causation from a before-and-after chart.
- Treat self-report as a hypothesis. In METR’s 2025 randomized trial, experienced open-source developers expected AI to speed them up by 24% and, after tasks took 19% longer, still believed it had sped them up by 20% (METR, 2025-07-10, early-2025 tools). Net time gain (metric 9) therefore always subtracts measured review and rework time.
- Name the sign-off. The VP Engineering signs off the organization panel each month, and a named metrics owner (often in the platform team) owns the definitions. A change to any definition on this page goes through that owner and is dated on the dashboard.
Telemetry pipelines and dashboards are covered in depth in observing agents in operation. The per-stage leading indicators that feed these metrics are in metrics for an AI-native SDLC.
What goes wrong when you measure AI-assisted engineering?
Section titled “What goes wrong when you measure AI-assisted engineering?”The team starts optimizing the pull request count. Once PR throughput appears in a goal or a performance review, agents make it trivially easy to inflate. Recovery: remove PR counts from every goal, report accepted change rate instead, and state in the measurement charter that no metric on this page is used for individual performance.
”% AI-written” becomes the target. A mandate to raise the AI share rewards pushing work through agents whether or not it fits. Recovery: move the metric to the utilization section of the dashboard, drop any target on it, and replace the goal with a lead-time or quality goal.
Stability degrades behind a throughput win. Lead time falls while change fail rate and rework rate creep up, which is the pattern DORA reported for 2025. Recovery: stop expanding agent scope in the affected services, look at the agent cohort’s failures first, and strengthen the gates, starting with the evidence bundle each change must carry.
Review becomes the bottleneck nobody measures. Merged pull requests rise, but time in review grows and a few senior engineers carry the load. Recovery: add reviewer load per engineer to the team panel, set pull request size budgets, and see how to keep the review queue moving.
Attribution disappears. A privacy setting or a new tool removes the agent marker, and metric 2 drops to zero overnight. Recovery: keep the pull request template checkbox as a fallback marker, and flag any week where a marker source changed.
Portfolio averages hide the problem. A blended change fail rate across 40 services looks stable while two critical services get worse. Recovery: report DORA metrics per service, as DORA recommends, and show the worst three services on the organization panel.
Where to go next with metrics
Section titled “Where to go next with metrics”Start from why the product is evidence, not diffs if the idea of measuring accepted changes rather than output is new to your leadership team. Then pick the page for your next decision.