The DORA AI Capabilities Model: a self-assessment
The DORA AI Capabilities Model names seven organizational capabilities that amplify the benefits of AI in software delivery: a clear and communicated AI stance, healthy data ecosystems, AI-accessible internal data, strong version control practices, working in small batches, user-centric focus, and quality internal platforms. DORA publishes no scoring, so this page adds a 0–3 evidence scale for each capability.
Your company bought agent seats for every engineer. Six months later, two teams ship more with the same incident rate, three ship more and break more, and the rest barely use the tools. The difference is rarely the model; it is the organization around the agents.
This page is for the CTO or VP Engineering who owns AI readiness, the tech lead who scores one team, and the executive deciding the next investment.
What you get from the DORA AI capabilities self-assessment
Section titled “What you get from the DORA AI capabilities self-assessment”- The seven capabilities with DORA’s exact names and what each one amplifies.
- A 0–3 scoring scale with the evidence that earns each score in an organization where agents write code.
- A copyable scorecard and a decision table that turns scores into next moves.
- A read-only agent run in Claude Code, Codex and Cursor that collects repository evidence for four capabilities.
What is the DORA AI Capabilities Model?
Section titled “What is the DORA AI Capabilities Model?”DORA, the research program at Google Cloud behind the software delivery metrics, published the model on 2025-09-23 (Kevin M. Storer and Derek DeBellis, “Introducing DORA’s inaugural AI Capabilities Model”). DORA derived candidates from “78 in-depth interviews” and literature, then tested them in a survey “reaching almost 5,000 respondents”. The result is “seven capabilities that substantially either amplify or unlock the benefits of AI”.
The model matters because of what the same year’s research found. Google Cloud’s announcement of the 2025 DORA report (Nathen Harvey and Derek DeBellis, 2025-09-23) reports “a positive relationship between AI adoption on both software delivery throughput and product performance.” In the next sentence: “However, AI adoption does continue to have a negative relationship with software delivery stability.” DORA’s explanation is the reason to assess capabilities at all: “AI accelerates software development, but that acceleration can expose weaknesses downstream. Without robust control systems, like strong automated testing, mature version control practices, and fast feedback loops, an increase in change volume leads to instability.”
The same announcement puts the report’s central theme in one line: “AI doesn’t fix a team; it amplifies what’s already there.”
How do you score each DORA AI capability?
Section titled “How do you score each DORA AI capability?”DORA names the capabilities and what they amplify; it does not publish levels or a score. The scale below is this site’s, and it has one rule: a score counts only if you can link the artifact that proves it. A score someone remembers is a 0.
| Score | Label | Evidence that earns it |
|---|---|---|
| 0 | Absent | Nothing exists, or nobody can point to it. |
| 1 | Ad hoc | Some teams do it, because of specific people. No written standard. |
| 2 | Defined | A written standard applies across the organization, but nothing enforces or measures it. |
| 3 | Enforced and measured | A control enforces it (managed settings, CI check, platform default) and a metric shows it holding, reviewed on a fixed cadence. |
Score each team separately, then take the organization’s score as the median team, and list the lowest team beside it. An average hides the teams where agents are already making things worse.
The seven capabilities, scored
Section titled “The seven capabilities, scored”Each capability gives DORA’s finding, what a 3 looks like, the check, and the page that builds it.
1. Clear and communicated AI stance
Section titled “1. Clear and communicated AI stance”DORA’s finding (2025-09-23): the stance covers “expectations for AI use, support for experimentation, and which tools are permitted”, and it “amplifies AI’s positive impact on individual effectiveness and organizational performance, and can reduce friction for employees.” A later DORA post (Nathen Harvey and Allison Park, 2025-12-10) adds that “Ambiguity creates risk.”
A 3 looks like: a one-page stance states what agents may do and what stays human, its enforceable clauses sit in managed settings, and a quarterly survey question (“Do you know what you may use agents for?”) is tracked.
Check: ask five engineers from different teams where the stance is and which tools are approved; five matching answers with a link is a 2.
Build it: leading people through the change for the stance, and an AI usage policy engineers will follow for the enforceable clauses.
2. Healthy data ecosystems
Section titled “2. Healthy data ecosystems”DORA’s finding (2025-09-23): a healthy data ecosystem, “characterized by high-quality, easily accessible, and unified internal data, substantially amplifies the positive influence of AI adoption on organizational performance.”
A 3 looks like: documentation, runbooks, architecture decisions and API contracts each have a named owner and a scheduled freshness check, and data is classified so it is clear which classes an agent may read.
Check: of ten documents an agent needs for a typical change, count those with an owner, updated in the last six months and matching the code; eight or more is a 2.
Build it: data privacy and enterprise policies for classification, and keeping context files lean and current for freshness.
3. AI-accessible internal data
Section titled “3. AI-accessible internal data”DORA’s finding (2025-09-23): “Connecting AI tools to internal data sources boosts their impact on individual effectiveness and code quality.” The 2025-12-10 post calls this “context engineering”.
A 3 looks like: every active repository has an instruction file (CLAUDE.md, AGENTS.md or Cursor project rules) with build, test and conventions; internal documentation and tickets reach the agent through approved least-privilege MCP servers; and the platform team tracks coverage.
Check: the evidence prompt below.
Build it: documentation as context for instruction files, internal MCP servers for connecting internal systems, and documentation-context MCP servers for the ready-made options.
4. Strong version control practices
Section titled “4. Strong version control practices”DORA’s finding (2025-09-23): “frequent commits amplify AI’s positive influence on individual effectiveness, while the frequent use of rollback features boosts the performance of AI-assisted teams.”
A 3 looks like: agent work reaches the default branch only through a reviewed pull request with required checks, agent-assisted pull requests are labelled as a cohort, and rollback is a rehearsed operation with a measured time.
Check: the evidence prompt, plus the date of the last rehearsed rollback.
Build it: parallel agents with Git worktrees for isolation, and progressive delivery for rollback.
5. Working in small batches
Section titled “5. Working in small batches”DORA’s finding (2025-09-23): working in small batches “amplifies the positive influence of AI on product performance and reduces friction for development teams.” The 2025-12-10 post: AI-generated “massive blocks of code” are “hard to review and test.”
A 3 looks like: a pull request size budget sits in every instruction file and a CI check enforces it, and the median size and share over budget are reported weekly, split by agent-assisted and other changes.
Check: the evidence prompt, against 400 changed lines, the review-queue guide’s starting budget.
Build it: enforce a PR size budget the agents follow, and an agent-ready backlog for slicing work before it reaches an agent.
6. User-centric focus
Section titled “6. User-centric focus”DORA’s finding (2025-09-23): a user-centric focus “amplifies the positive influence of AI on team performance.” And the warning: “in the absence of a user-centric focus, AI adoption can have a negative impact on team performance.”
A 3 looks like: every agent task starts from a user outcome with executable acceptance criteria, and features carry an outcome metric and a kill criterion.
Check: of ten recent agent-assisted changes, count those that trace to a user outcome with acceptance criteria that ran in CI.
Build it: product management when build time collapses for intent and kill criteria, and executable acceptance criteria for the check agents cannot skip.
7. Quality internal platforms
Section titled “7. Quality internal platforms”DORA’s finding (2025-09-23): “In organizations with quality internal platforms, AI’s positive influence on organizational performance is amplified.” Google Cloud’s announcement of the 2025 DORA report (2025-09-23) adds that “90% of organizations have adopted at least one platform and there is a direct correlation between a high quality internal platform and an organization’s ability to unlock the value of AI.”
A 3 looks like: a platform team owns the agent harness as a product (managed settings, shared rules, skills, hooks, CI gates and telemetry), ships it by default and measures adoption and the change fail rate of agent-assisted work.
Check: the evidence prompt, plus one question: can a new repository get the approved agent setup, gates and telemetry without a ticket?
Build it: the agent platform team, managed policy across every coding agent, and shared agent rules.
The DORA AI capabilities scorecard you can copy
Section titled “The DORA AI capabilities scorecard you can copy”Paste this into a document or spreadsheet, one copy per team. The Evidence link column is mandatory; an empty cell scores 0.
# DORA AI capabilities self-assessmentTeam: ______ Assessed on: YYYY-MM-DD Assessors: ______ (lead), ______ (second rater)
| # | Capability (DORA name) | Score 0-3 | Evidence link | Metric and value | Builds it (owner, date) ||---|------------------------------------|-----------|---------------|-----------------------|-------------------------|| 1 | Clear and communicated AI stance | | | survey: % who know it | || 2 | Healthy data ecosystems | | | % docs owned + fresh | || 3 | AI-accessible internal data | | | % repos with context | || 4 | Strong version control practices | | | rollback time | || 5 | Working in small batches | | | median PR lines | || 6 | User-centric focus | | | % changes with AC | || 7 | Quality internal platforms | | | self-serve: yes/no | |
Lowest score: __ (capability #__) Disagreements between raters: __Next move (from the decision table): ______Re-score on: YYYY-MM-DD (90 days)How do you run the assessment?
Section titled “How do you run the assessment?”-
Pick the unit and the raters. Score per team with two raters: the tech lead and someone from outside the team. The CTO scores capabilities 1 and 7 at organization level.
-
Collect the repository evidence first. Run the read-only agent collection below for capabilities 3, 4, 5 and 7, so scoring starts from numbers.
-
Score independently, then reconcile. Each rater scores alone; a two-point difference goes to a 20-minute discussion over the evidence.
-
Map the value stream for the lowest capability. DORA recommends a value stream mapping exercise because “by visualizing your flow from idea to customer, you can identify where work waits and where friction exists” (Nathen Harvey and Allison Park, 2025-12-10). Apply it to the lowest-scoring capability, so the fix targets a system constraint.
-
Assign an owner and a page. Each capability below 2 gets an owner, a due date and the linked page that builds it.
-
Re-score in 90 days with the same raters, against the delivery metrics below.
Collect the repository evidence with an agent
Section titled “Collect the repository evidence with an agent”In all three tools the agent reads Git history, CI configuration and instruction files and writes nothing; only how it is kept read-only differs.
Save the prompt below as dora-evidence.md and run it headlessly from the repository root. --permission-mode dontAsk denies every tool call that is not on the allowlist, so the run cannot use the Edit or Write tools or run any command outside the list (checked against Claude Code v2.1.283). git log still accepts --output=<file>, which writes a file, so run it on a clean checkout and confirm with git status afterwards.
claude -p "$(cat dora-evidence.md)" \ --permission-mode dontAsk \ --allowedTools "Read" "Grep" "Glob" "Bash(git log *)" "Bash(gh pr list *)" "Bash(gh api repos/*/branches/*/protection)"The allowlist narrows gh api to commands that end in the branch-protection path, rather than a bare gh api rule that would also allow write calls such as gh api -X DELETE; the prompt forbids writes, and the clean-checkout + git status check still applies.
List the MCP servers Claude Code can reach in this repository with claude mcp list. It health-checks approved servers, which starts them, and shows unapproved .mcp.json servers as pending; to start no process at all, read .mcp.json directly. That output is the evidence for capability 3.
Run the same prompt non-interactively with the read-only sandbox (checked against Codex CLI 0.157.1):
codex exec --sandbox read-only -o dora-evidence-result.md "$(cat dora-evidence.md)"--sandbox read-only is the legacy flag and still works in 0.157.1; new setups can use the beta permission profile instead, -c default_permissions=":read-only" (Codex CLI 0.157.1), without --sandbox; the two systems do not compose.
The read-only sandbox has no network access (Codex CLI 0.157.1), so the gh calls for capabilities 4 and 5 fail; take those numbers from the Claude Code run or the GitHub UI. The prompt reports the gap instead of guessing.
List the MCP servers Codex is configured to use with codex mcp list. That output is the evidence for capability 3.
Open the repository, switch the agent to Plan Mode, which creates a plan before writing any code (Plan Mode as documented by Cursor, checked 2026-08-28), and paste the prompt. Do not approve a build step afterwards. The read-only guarantee is procedural, not enforced, so run it on a clean branch and check git status afterwards.
For capability 3, record the servers in the project’s .cursor/mcp.json alongside the Claude Code and Codex lists.
How do you turn the scores into a plan?
Section titled “How do you turn the scores into a plan?”Read the scorecard by its minimum, not its total: 17 of 21 with a 0 on small batches is riskier than 12 with no score below 1. The table puts controls before volume, which is this site’s recommendation, not DORA’s.
| If you see | Then | Why |
|---|---|---|
| Capability 4 or 5 at 0–1 | Stop widening agent use in those teams until both reach 2. Start with the PR size budget and branch protection. | These are the control systems DORA names. |
| Capability 1 at 0–1 | Publish the stance within 30 days. | It is the cheapest capability, and only leadership can deliver it. |
| Capability 6 at 0–1 | Remove output targets (pull requests, ”% AI-written”) from team goals; require acceptance criteria on agent tasks. | DORA found AI adoption can hurt team performance without a user-centric focus. |
| Capabilities 2–3 at 0–1, others at 2+ | Fund context work: instruction files first, then one internal MCP server. | Without company context, agents stay generic. |
| Capability 7 at 0–1 and three or more teams using agents | Charter a platform team. | Each team is building its own harness, and the gains stop at the team boundary. |
| All seven at 2+ | Move each 2 to a 3 by adding the control and the metric. | A written standard nobody enforces decays. |
How do you know the self-assessment is telling the truth?
Section titled “How do you know the self-assessment is telling the truth?”An unchecked readiness score drifts upward. Four controls keep it honest.
- Evidence or zero. A score without a working link is a 0, including for the CTO’s own capabilities.
- Two raters, one from outside. A two-point disagreement is reconciled on the evidence, and the readout counts disagreements.
- Pair the scores with delivery outcomes. Capabilities are inputs; the proof is in the delivery metrics. When small batches and version control move from 1 to 3, change fail rate for agent-assisted changes should hold or fall while throughput rises. Use the definitions in DORA, SPACE, DX Core 4 and AI measurement.
- Test an investment before you scale it. When a capability fix is expensive, such as a platform team or an internal MCP server programme, run it first as a pilot with a baseline and a decision rule.
The VP Engineering signs off the readout each quarter, and its first line says the scores never rank teams or individuals.
What goes wrong with a DORA AI capabilities self-assessment?
Section titled “What goes wrong with a DORA AI capabilities self-assessment?”Every team scores itself 2 or 3. Recovery: apply evidence-or-zero retroactively, add the outside rater, and publish how many scores fell.
The total becomes a target. Leadership sets “reach 18 of 21 by Q2”, and teams raise the easy capabilities while small batches stays at 1. Recovery: report the minimum and the lowest team, never a total.
The assessment replaces the metrics. A team reaches 3 everywhere while its change fail rate rises. Recovery: treat the delivery metrics as the arbiter, and re-examine any score that does not show in outcomes after two quarters.
The stance is published and forgotten. Recovery: date it, review it quarterly, and keep the survey question on the team panel.
Capability 3 grows without capability 2. Agents connected to every wiki faithfully repeat stale documentation. Recovery: connect only sources with an owner and a freshness check, and remove MCP servers that point at unowned data.
Where to go next with DORA AI capabilities
Section titled “Where to go next with DORA AI capabilities”If leadership confuses readiness with loop autonomy, start with one map for the ladder, the lifecycle and the factory, then take the page for your lowest capability.
Frequently asked questions
What are the seven capabilities in the DORA AI Capabilities Model?
Clear and communicated AI stance, healthy data ecosystems, AI-accessible internal data, strong version control practices, working in small batches, user-centric focus, and quality internal platforms (DORA, Google Cloud, 2025-09-23).
Does DORA publish maturity levels or a score for the AI Capabilities Model?
No. DORA names the seven capabilities and what each one amplifies. The 0–3 scale on this page is this site's evidence-based scoring, not DORA's.
Which DORA AI capability should an engineering organization fix first?
Fix the lowest-scoring capability among the AI stance, strong version control practices and working in small batches before you widen agent use, because DORA links higher change volume without control systems to instability.
Is the DORA AI Capabilities Model the same as the autonomy ladder?
No. The model is an organizational readiness checklist with no levels. The autonomy ladder measures how far each delivery loop runs without a human, and the two are not mapped onto each other.