Reporting AI engineering to the board
A board report on AI engineering is a one-page quarterly update with four blocks: the autonomy level each critical delivery loop runs at, the cost per accepted change, the quality and incident trend, and a five-row risk register. It never uses lines of code, seat counts or pull-request volume as proof of value.
The board approved the AI tooling budget a year ago. Now the audit committee chair asks two questions: what did the company get for it, and what could go wrong? You have a vendor dashboard with a rising pull-request count, a “62% of code is AI-written” slide from an engineering all-hands, and a case study from someone else’s company. None of those survive the second follow-up question, and the chair knows it.
This page is for the CTO who writes that report and the executive who presents or reads it. It is the last step of the executive and CTO reading tracks: it turns the economics, the business case and the transformation roadmap into something the board can steer by.
What you get from this board reporting kit
Section titled “What you get from this board reporting kit”- A one-page quarterly template you can paste into your board pack today.
- Metric definitions for every number on the page, with the formula, the data source and the gaming risk.
- Where each number comes from in Claude Code, Codex and Cursor, and what to do when a tool cannot supply it.
- A risk register with five standing rows, each with an owner, a leading indicator and dated evidence.
- A “what not to claim” table, with the published evidence behind each line.
- A quarterly cycle with named sign-offs, two copy-paste prompts that draft and attack the pack, and the verification that runs before anything reaches the board.
What goes on the one-page board report?
Section titled “What goes on the one-page board report?”Copy this into your board pack as-is. Every blank is a number you already have or a gap you declare as “not measured”. A declared gap is acceptable in the first two quarters; an invented number never is.
# AI engineering: quarterly board update — Q_ YYYYPrepared by: CTO · Cost block co-signed by: CFO · Risk rows signed by: CISO, General CounselDefinitions version: v_ (link to the metric definitions; unchanged since last quarter: yes/no)
## 1. Headline (three sentences, no adjectives)- What changed: ___- What it cost per accepted change, against last quarter: ___- The one risk the board should watch: ___
## 2. Autonomy level per critical loop| Loop (owner) | Level last Q | Level now | Target next Q | Gate that must hold to move up | Evidence link ||-------------------------|--------------|-----------|---------------|---------------------------------------|---------------|| Dependency updates (__) | L_ | L_ | L_ | e.g. full test suite + canary green | ___ || Bug fixes, service A (__)| L_ | L_ | L_ | ___ | ___ || Feature work, product B (__)| L_ | L_ | L_ | ___ | ___ |
## 3. Economics| Measure | Last Q | This Q | 4-quarter trend | Note ||-----------------------------------------|--------|--------|-----------------|------|| Total AI engineering spend vs budget | | | | || Accepted changes (merged, not reverted) | | | | || Cost per accepted change | | | | || Share of spend that is review and rework| | | | |
## 4. Quality and incidents (split: agent-authored / human-authored)| Measure | Agent | Human | Trend ||-------------------------------------------|-------|-------|-------|| Change failure rate | | | || Incidents per 100 accepted changes | | | || Median time in review | | | || Share of merges with no human or agent review | | | || Median change size (lines) | | | |
## 5. Risk register (five standing rows)| Risk | Rating (L×I) | Trend | Leading indicator | Control in place | Owner | Next action ||-----------------|--------------|-------|-------------------|------------------|-------|-------------|| Security | | | | | CISO | || IP and licensing| | | | | GC | || Vendor | | | | | CTO | || Skills pipeline | | | | | CTO/CHRO | || Regulation | | | | | GC | |
## 6. Decisions requested1. ___ (budget, a loop moving up a level, a policy approval)
## 7. What this report does not claimNo productivity multiplier, no headcount saving, no "% of code written by AI" as value.Which numbers go on the page, and how are they defined?
Section titled “Which numbers go on the page, and how are they defined?”Define every metric once, version the definitions, and change them only between quarters with a note in the header. A metric whose definition moves while the trend is being read is the fastest way to lose a board’s trust. The canonical definitions of the AI metrics live on the metrics frameworks page; the table below is the board-level subset.
| Metric | Definition | Source | Gaming risk and the counter-metric |
|---|---|---|---|
| Level per critical loop | The rung of the autonomy ladder (L0 to L5) at which a recurring unit of work runs, set by the first ladder question the loop fails | Loop owner’s assessment plus the gate evidence | Inflating the level: require a named gate and a link to its last green run for every level above L2 |
| Accepted change | A merged change that met its acceptance criteria and was not reverted within 30 days | Version control plus the revert log | Splitting work into tiny changes: report median change size beside it |
| Cost per accepted change | (licences + metered usage + CI + review hours at loaded cost + rework + incident cost) ÷ accepted changes | Invoices, CI billing, time data; method on the economics page | Leaving out review and rework: finance co-signs the inputs |
| Change failure rate | Share of deployments that cause a failure needing remediation, split by provenance | Deploy and incident systems plus a provenance label | Relabelling incidents: the incident owner, not the author, sets the label |
| Incidents per 100 accepted changes | Production incidents traced to a change, normalised by accepted changes | Incident postmortems | Under-reporting: count every Sev-1 and Sev-2 with a linked change |
| Median time in review | Median time from review request to approval | Code host | Rubber-stamp approvals: report the share of merges with no review beside it |
Two definitions do most of the work. “Accepted” makes the unit of value something the business recognises, and “split by provenance” is what lets the board see whether agent-authored work fails more or less often than human work. Provenance needs a label on every change; the PR labelling guide sets one up that does not depend on any single vendor.
How do you choose the critical loops?
Section titled “How do you choose the critical loops?”Pick three to six loops, not the whole engineering organisation. A loop qualifies when it recurs often enough to trend, when a failure in it reaches customers or regulators, and when an owner can name its gate. Typical first choices: dependency updates, bug fixes in a revenue-bearing service, feature work in the main product, and infrastructure changes.
Report a level per loop, never one level for the company. The ladder itself says a level belongs to a loop: a dependency-update loop with a deterministic check can run at L4 while feature work beside it sits at L2, and both answers are correct at once. A single company-wide level hides exactly the difference the board needs to see.
Where do the numbers come from in Claude Code, Codex and Cursor?
Section titled “Where do the numbers come from in Claude Code, Codex and Cursor?”The spend and adoption inputs differ by tool; the quality inputs do not. Change failure rate, incidents and review time come from your code host and incident system whatever the agent, so build those queries once. For deeper telemetry setup, see agent telemetry.
- Analytics dashboard (Team and Enterprise plans): active users, sessions and, with the GitHub app installed, pull requests with Claude Code and the share of merged pull requests with Claude Code. Merged pull requests with attributed lines carry the GitHub label
claude-code-assisted, so this search on your code host gives a sample without API access:is:pr is:merged label:claude-code-assisted. - Caveat for the board: contribution metrics are a public beta and are unavailable under Zero Data Retention (Anthropic analytics docs, checked 2026-09-26). If you run ZDR, your own provenance label is the only source.
- OpenTelemetry export for spend: set
CLAUDE_CODE_ENABLE_TELEMETRY=1,OTEL_METRICS_EXPORTER=otlpandOTEL_EXPORTER_OTLP_ENDPOINTpointing at your collector (without an endpoint the exporter targets localhost and the board pack gets nothing; agent telemetry has the full variable set); theclaude_code.cost.usagemetric feeds the economics block. - Sanity check for the spend line: Anthropic’s cost documentation gives “around $13 per developer per active day and $150-250 per developer per month” for enterprise deployments (checked 2026-09-26). A figure far outside that range is worth a question before it reaches the board.
- OpenTelemetry export: an
[otel]table in~/.codex/config.tomlexports logs, traces and metrics to your collector (read from the Codex source for CLI 0.157.1, 2026-09-26). Send it to the same collector as your other agents so the economics block has one source. - Local cost reports across agents:
npx ccusage@latest codex monthlyreads local Codex logs and prints cost per month (ccusage 20.0.24, 2026-09-26). It covers one machine, so use it to check a vendor invoice, not to replace it. - Enterprise analytics: OpenAI describes a Codex analytics dashboard and API for ChatGPT Enterprise workspaces; its pages could not be read from the writing environment on 2026-09-26, so confirm with your admin which fields your workspace exposes before you put them in the template.
- Team analytics: Cursor documents a team analytics dashboard and admin APIs for its team plans. cursor.com could not be reached from the writing environment on 2026-09-26, so confirm with your Cursor admin which fields your plan exposes before promising them to the board.
- Provenance: do not depend on a per-commit AI attribution feature until you have confirmed your plan includes it. Label changes in your own pipeline instead; the label survives a tool switch.
- Spend: reconcile the Cursor invoice with the accepted-change count from your code host; the formula is the same as for the other tools.
With more than one agent, the rule holds for each: vendor dashboards describe adoption, your own pipeline describes outcomes, and the board page uses vendor data only for the spend line.
How do you build the risk register?
Section titled “How do you build the risk register?”Keep five standing rows every quarter, even when a row is quiet. A board learns to trust a register that shows green rows turning amber; it stops trusting one where risks appear only after they have happened.
| Risk | What the board needs to know | Leading indicator | Dated evidence you can cite | Canonical page |
|---|---|---|---|---|
| Security | The agent and its toolchain are attack surface: prompt injection through issues and pull requests, stolen credentials, poisoned packages | Agent credentials without an expiry; workflows that let untrusted issue text reach an agent with write access | Cline’s advisory of 2026-02-17 records an unauthorised npm publish of cline@2.3.0; the malicious nx releases (advisory 2025-08-27) collected credentials and posted them to GitHub | Agent threat model |
| IP and licensing | Who owns agent-written code, licence contamination, and what vendor indemnities cover | Share of repositories with licence scanning in CI; client contracts silent on agent use | Checked with counsel each quarter, not from a statistic | Legal and IP |
| Vendor | Models, defaults and plans change monthly, and a change can move quality or cost without any decision on your side | Loops whose evals have not been rerun since the last default-model change | Codex’s bundled default became GPT-6 Astra on 2026-09-04; Claude Code’s default became Claude Opus 5.5 from v2.1.280 (the latest channel) on 2026-09-22; GPT-5.4 was retired in Codex on 2026-08-31 with automatic migration. Versions and prices live on the models hub | Avoiding lock-in |
| Skills pipeline | Oversight needs the skills that heavy delegation erodes; junior engineers learn less by writing code | Juniors per senior reviewer; share of juniors’ merged work they could explain in review | Anthropic’s internal study (2025-12-02, a survey of 132 engineers and researchers): staff worry about “skills atrophying as [they] delegate more”, and one respondent notes that “more junior people don’t come to me with questions as often” | Junior developers |
| Regulation | Which duties reach a company that uses coding agents, and when shipping AI features changes that | AI-literacy training coverage; AI features shipped to EU users | After the Digital Omnibus (Regulation (EU) 2026/1744), EU AI Act Art. 50 transparency duties apply from 2 August 2026 and Annex III high-risk obligations are deferred to 2 December 2027 (secondary: Gibson Dunn and Usercentrics summaries; re-verify on EUR-Lex before the board meeting) | EU AI Act |
Rate each row on likelihood and impact, show the trend arrow, and name the one action due before the next meeting. The vendor row is the one executives underestimate most: in September 2026 both Claude Code and Codex changed their default model within three weeks of each other, and a loop whose quality depends on the model can shift without anyone deciding to shift it. That is why the vendor row’s indicator is “evals not rerun since the last default change”.
When does a risk skip the quarterly cycle?
Section titled “When does a risk skip the quarterly cycle?”Agree escalation triggers with the chair in advance, so nobody has to judge in the moment whether something is “board-worthy”. Escalate between meetings when an incident traced to an agent-authored change reaches customers or a regulator, when a loop drops a level after an incident, when an agent credential is revoked for cause, or when spend runs more than a pre-agreed margin over budget. The agent incidents guide covers containment and the postmortem that feeds the next report.
What should you not claim to the board?
Section titled “What should you not claim to the board?”This section of the template is the one that protects you. Each claim below appears in vendor decks and all-hands slides, and each fails under a prepared question.
| Do not claim | Why it fails | Report instead |
|---|---|---|
| “X% of our code is written by AI” as proof of value | A share of lines is not a share of work or of value. DX’s industry figure, a 51.9% average across more than 400 companies (June 2026), measures authorship, not outcome | Cost per accepted change and the incident trend |
| “Pull requests are up, so productivity is up” | Faros AI’s telemetry of 22,000 developers (April 2026) shows throughput and quality moving apart: PR merge rate per developer +16.2%, but incidents per pull request +242.7% and median time in review +441.5% | Accepted changes and incidents per 100 accepted changes, side by side |
| “AI makes our developers N% faster” | No general measurement supports it. METR’s 2026 follow-up has point estimates in AI’s favour, but the intervals cross zero and METR says selection makes the results hard to interpret | Lead time for accepted changes in named loops, against a baseline from a designed pilot |
| “We can cut headcount by a factor of three to five” | No third-party measurement supports any productivity multiplier | Where freed capacity went, per the business case |
| “Anthropic merges over 80% AI-authored code, so we can too” | Anthropic’s own figure (“more than 80% of the code we merge”, as of May 2026, Anthropic Institute) is vendor-internal: it describes that company’s harness and test suite, not an industry benchmark | Your own level per loop and the gate that holds it |
| “DORA says AI improves delivery” | DORA’s 2025 report pairs a positive relationship with throughput with a negative relationship with stability; quote both halves or neither | Your change failure rate split by provenance |
| “Time saved is money saved” | Saved hours become cash only when capacity is redeployed or a cost is avoided | The redeployment decision, recorded with an owner |
If an executive wants one number, give cost per accepted change with its four-quarter trend. It includes review, rework and incidents, so it cannot improve by generating more code that nobody can verify.
How do you run the quarterly reporting cycle?
Section titled “How do you run the quarterly reporting cycle?”-
Week −3: freeze definitions and export. The engineering operations or platform owner runs the versioned queries for every metric and saves the raw exports, with query IDs, to a folder such as
board/2026-q3/. Nothing is computed by hand. -
Week −3: loop owners set their levels. Each owner answers the ladder questions for their loop and links the last green run of the gate that justifies the level. A level without a gate link is reported one rung lower.
-
Week −2: draft with an agent. Use the drafting prompt below in Claude Code, Codex or Cursor. The agent only arranges numbers that already exist in the exports and footnotes the source of each one.
-
Week −2: attack the draft. Run the sceptical-chair prompt below, then fix or declare every finding. This is the step that catches a flattering trend built on a changed definition.
-
Week −1: verify and sign. Finance reconciles the cost block with invoices and co-signs it. The CISO signs the security row, legal signs the IP and regulation rows, and the CTO signs the page. The signatures go in the header.
-
Week −1: send the pre-read. The chair gets the page and the appendix a week ahead. The meeting discusses the decisions in section 6, not the numbers.
-
After the meeting: record decisions and targets. The target levels per loop become next quarter’s column, and the transformation roadmap gates are updated to match.
The drafting and attacking steps work identically in all three tools: open the agent in the folder that holds the exports and paste the prompt. For an unattended run, use claude -p --permission-mode dontAsk "…" > board/2026-q3/draft.md in Claude Code, or codex exec -s read-only -o board/2026-q3/draft.md "…" in Codex; both only read the exports and print the draft. In Cursor, paste the prompt into the agent with the folder open.
How do you know the report is right before it reaches the board?
Section titled “How do you know the report is right before it reaches the board?”Nobody should have to re-read the page line by line to trust it. Build the checks into the cycle so each one produces evidence someone signs:
- Reproducibility. Every number has a query ID in a versioned repository, and a second person reruns the queries and gets the same values. A mismatch blocks the pack.
- Reconciliation. Finance reconciles total spend with invoices within an agreed tolerance before co-signing.
- Sampled audit of “accepted”. Draw 10 accepted changes at random and confirm each met its acceptance criteria and was not reverted. One failure means the definition or the data is wrong, and the trend line waits.
- Provenance coverage. Report the share of changes carrying a provenance label. Below your threshold, the agent/human split is marked “not comparable”.
- Gate evidence. Each level above L2 links to the last green run of its gate, dated within the quarter.
- Named sign-offs. CTO, CFO, CISO and general counsel sign their blocks, and the signatures sit in the header, not in an email thread.
The agent drafts and attacks; people sign, the same division this site asks engineers to apply to code.
When AI board reporting goes wrong
Section titled “When AI board reporting goes wrong”The metric becomes the target. Teams split work into smaller changes to lower cost per accepted change. Recovery: report median change size beside it; DX measured median pull-request size rising from 44 to 72 lines between July 2025 and June 2026 (June 2026), so a sudden fall in your data is worth a question. Change the unit to accepted outcomes if splitting persists.
The numbers do not reconcile with finance. Vendor dashboards and invoices disagree, usually because of seats bought but unused, metered overage, or a second tool on expenses. Recovery: take spend from invoices only, report the dashboard gap as a note, and move the fix into cost governance.
Attribution breaks. A move to Zero Data Retention switches off Claude Code’s contribution metrics, or a tool switch changes the labels. Recovery: your own provenance label in the pipeline, set by the PR labelling guide, so the series continues across vendors.
A loop drops a level. After an incident, a loop falls from L4 to L3. Report it plainly: the drop is the control working, and hiding it costs more credibility than the incident. Link the postmortem and the new eval that came out of it.
The board asks for “the ROI percentage”. A single ROI figure invites the multipliers the “what not to claim” table rules out. Recovery: give cost per accepted change with its trend and point to the decision recorded in the business case.
The report grows to 12 pages. Recovery: a hard one-page limit, an appendix for everything else, and a rule that a new row replaces an old one.
Where to go next with board reporting
Section titled “Where to go next with board reporting”- Set the cost method behind the economics block: the economics of agent-built software.
- Get the baseline that makes the first trend meaningful: designing a pilot that proves something.
- Settle the metric definitions across the organisation: DORA, SPACE, DX Core 4 and AI measurement.
- Check the gates behind each target level: the organisation-wide transformation roadmap.
- Find where your organisation sits today, with a risk register to start from: take the free C-level scorecard.
Frequently asked questions
What should a board report on AI engineering contain?
One page per quarter with four blocks: the autonomy level each critical delivery loop runs at, the cost per accepted change, the quality and incident trend split by provenance, and a five-row risk register covering security, IP, vendor, skills pipeline and regulation. It ends with the decisions the board is asked to take.
Should the board see the percentage of code written by AI?
Not as a measure of value. A share of lines says nothing about accepted outcomes, and industry figures such as DX's 51.9% average (June 2026) are a share of code, not of work. Report cost per accepted change and the incident trend instead.
Who signs off the AI engineering board report?
The CTO owns the page. Finance co-signs the cost block after reconciling it with invoices, the CISO signs the security row, and legal signs the IP and regulation rows. Every number carries the query that produced it.