Skip to content

Reporting AI engineering to the board

A board report on AI engineering is a one-page quarterly update with four blocks: the autonomy level each critical delivery loop runs at, the cost per accepted change, the quality and incident trend, and a five-row risk register. It never uses lines of code, seat counts or pull-request volume as proof of value.

The board approved the AI tooling budget a year ago. Now the audit committee chair asks two questions: what did the company get for it, and what could go wrong? You have a vendor dashboard with a rising pull-request count, a “62% of code is AI-written” slide from an engineering all-hands, and a case study from someone else’s company. None of those survive the second follow-up question, and the chair knows it.

This page is for the CTO who writes that report and the executive who presents or reads it. It is the last step of the executive and CTO reading tracks: it turns the economics, the business case and the transformation roadmap into something the board can steer by.

What you get from this board reporting kit

Section titled “What you get from this board reporting kit”
  • A one-page quarterly template you can paste into your board pack today.
  • Metric definitions for every number on the page, with the formula, the data source and the gaming risk.
  • Where each number comes from in Claude Code, Codex and Cursor, and what to do when a tool cannot supply it.
  • A risk register with five standing rows, each with an owner, a leading indicator and dated evidence.
  • A “what not to claim” table, with the published evidence behind each line.
  • A quarterly cycle with named sign-offs, two copy-paste prompts that draft and attack the pack, and the verification that runs before anything reaches the board.

Copy this into your board pack as-is. Every blank is a number you already have or a gap you declare as “not measured”. A declared gap is acceptable in the first two quarters; an invented number never is.

# AI engineering: quarterly board update — Q_ YYYY
Prepared by: CTO · Cost block co-signed by: CFO · Risk rows signed by: CISO, General Counsel
Definitions version: v_ (link to the metric definitions; unchanged since last quarter: yes/no)
## 1. Headline (three sentences, no adjectives)
- What changed: ___
- What it cost per accepted change, against last quarter: ___
- The one risk the board should watch: ___
## 2. Autonomy level per critical loop
| Loop (owner) | Level last Q | Level now | Target next Q | Gate that must hold to move up | Evidence link |
|-------------------------|--------------|-----------|---------------|---------------------------------------|---------------|
| Dependency updates (__) | L_ | L_ | L_ | e.g. full test suite + canary green | ___ |
| Bug fixes, service A (__)| L_ | L_ | L_ | ___ | ___ |
| Feature work, product B (__)| L_ | L_ | L_ | ___ | ___ |
## 3. Economics
| Measure | Last Q | This Q | 4-quarter trend | Note |
|-----------------------------------------|--------|--------|-----------------|------|
| Total AI engineering spend vs budget | | | | |
| Accepted changes (merged, not reverted) | | | | |
| Cost per accepted change | | | | |
| Share of spend that is review and rework| | | | |
## 4. Quality and incidents (split: agent-authored / human-authored)
| Measure | Agent | Human | Trend |
|-------------------------------------------|-------|-------|-------|
| Change failure rate | | | |
| Incidents per 100 accepted changes | | | |
| Median time in review | | | |
| Share of merges with no human or agent review | | | |
| Median change size (lines) | | | |
## 5. Risk register (five standing rows)
| Risk | Rating (L×I) | Trend | Leading indicator | Control in place | Owner | Next action |
|-----------------|--------------|-------|-------------------|------------------|-------|-------------|
| Security | | | | | CISO | |
| IP and licensing| | | | | GC | |
| Vendor | | | | | CTO | |
| Skills pipeline | | | | | CTO/CHRO | |
| Regulation | | | | | GC | |
## 6. Decisions requested
1. ___ (budget, a loop moving up a level, a policy approval)
## 7. What this report does not claim
No productivity multiplier, no headcount saving, no "% of code written by AI" as value.

Which numbers go on the page, and how are they defined?

Section titled “Which numbers go on the page, and how are they defined?”

Define every metric once, version the definitions, and change them only between quarters with a note in the header. A metric whose definition moves while the trend is being read is the fastest way to lose a board’s trust. The canonical definitions of the AI metrics live on the metrics frameworks page; the table below is the board-level subset.

MetricDefinitionSourceGaming risk and the counter-metric
Level per critical loopThe rung of the autonomy ladder (L0 to L5) at which a recurring unit of work runs, set by the first ladder question the loop failsLoop owner’s assessment plus the gate evidenceInflating the level: require a named gate and a link to its last green run for every level above L2
Accepted changeA merged change that met its acceptance criteria and was not reverted within 30 daysVersion control plus the revert logSplitting work into tiny changes: report median change size beside it
Cost per accepted change(licences + metered usage + CI + review hours at loaded cost + rework + incident cost) ÷ accepted changesInvoices, CI billing, time data; method on the economics pageLeaving out review and rework: finance co-signs the inputs
Change failure rateShare of deployments that cause a failure needing remediation, split by provenanceDeploy and incident systems plus a provenance labelRelabelling incidents: the incident owner, not the author, sets the label
Incidents per 100 accepted changesProduction incidents traced to a change, normalised by accepted changesIncident postmortemsUnder-reporting: count every Sev-1 and Sev-2 with a linked change
Median time in reviewMedian time from review request to approvalCode hostRubber-stamp approvals: report the share of merges with no review beside it

Two definitions do most of the work. “Accepted” makes the unit of value something the business recognises, and “split by provenance” is what lets the board see whether agent-authored work fails more or less often than human work. Provenance needs a label on every change; the PR labelling guide sets one up that does not depend on any single vendor.

Pick three to six loops, not the whole engineering organisation. A loop qualifies when it recurs often enough to trend, when a failure in it reaches customers or regulators, and when an owner can name its gate. Typical first choices: dependency updates, bug fixes in a revenue-bearing service, feature work in the main product, and infrastructure changes.

Report a level per loop, never one level for the company. The ladder itself says a level belongs to a loop: a dependency-update loop with a deterministic check can run at L4 while feature work beside it sits at L2, and both answers are correct at once. A single company-wide level hides exactly the difference the board needs to see.

Where do the numbers come from in Claude Code, Codex and Cursor?

Section titled “Where do the numbers come from in Claude Code, Codex and Cursor?”

The spend and adoption inputs differ by tool; the quality inputs do not. Change failure rate, incidents and review time come from your code host and incident system whatever the agent, so build those queries once. For deeper telemetry setup, see agent telemetry.

  • Analytics dashboard (Team and Enterprise plans): active users, sessions and, with the GitHub app installed, pull requests with Claude Code and the share of merged pull requests with Claude Code. Merged pull requests with attributed lines carry the GitHub label claude-code-assisted, so this search on your code host gives a sample without API access: is:pr is:merged label:claude-code-assisted.
  • Caveat for the board: contribution metrics are a public beta and are unavailable under Zero Data Retention (Anthropic analytics docs, checked 2026-09-26). If you run ZDR, your own provenance label is the only source.
  • OpenTelemetry export for spend: set CLAUDE_CODE_ENABLE_TELEMETRY=1, OTEL_METRICS_EXPORTER=otlp and OTEL_EXPORTER_OTLP_ENDPOINT pointing at your collector (without an endpoint the exporter targets localhost and the board pack gets nothing; agent telemetry has the full variable set); the claude_code.cost.usage metric feeds the economics block.
  • Sanity check for the spend line: Anthropic’s cost documentation gives “around $13 per developer per active day and $150-250 per developer per month” for enterprise deployments (checked 2026-09-26). A figure far outside that range is worth a question before it reaches the board.

With more than one agent, the rule holds for each: vendor dashboards describe adoption, your own pipeline describes outcomes, and the board page uses vendor data only for the spend line.

Keep five standing rows every quarter, even when a row is quiet. A board learns to trust a register that shows green rows turning amber; it stops trusting one where risks appear only after they have happened.

RiskWhat the board needs to knowLeading indicatorDated evidence you can citeCanonical page
SecurityThe agent and its toolchain are attack surface: prompt injection through issues and pull requests, stolen credentials, poisoned packagesAgent credentials without an expiry; workflows that let untrusted issue text reach an agent with write accessCline’s advisory of 2026-02-17 records an unauthorised npm publish of cline@2.3.0; the malicious nx releases (advisory 2025-08-27) collected credentials and posted them to GitHubAgent threat model
IP and licensingWho owns agent-written code, licence contamination, and what vendor indemnities coverShare of repositories with licence scanning in CI; client contracts silent on agent useChecked with counsel each quarter, not from a statisticLegal and IP
VendorModels, defaults and plans change monthly, and a change can move quality or cost without any decision on your sideLoops whose evals have not been rerun since the last default-model changeCodex’s bundled default became GPT-6 Astra on 2026-09-04; Claude Code’s default became Claude Opus 5.5 from v2.1.280 (the latest channel) on 2026-09-22; GPT-5.4 was retired in Codex on 2026-08-31 with automatic migration. Versions and prices live on the models hubAvoiding lock-in
Skills pipelineOversight needs the skills that heavy delegation erodes; junior engineers learn less by writing codeJuniors per senior reviewer; share of juniors’ merged work they could explain in reviewAnthropic’s internal study (2025-12-02, a survey of 132 engineers and researchers): staff worry about “skills atrophying as [they] delegate more”, and one respondent notes that “more junior people don’t come to me with questions as often”Junior developers
RegulationWhich duties reach a company that uses coding agents, and when shipping AI features changes thatAI-literacy training coverage; AI features shipped to EU usersAfter the Digital Omnibus (Regulation (EU) 2026/1744), EU AI Act Art. 50 transparency duties apply from 2 August 2026 and Annex III high-risk obligations are deferred to 2 December 2027 (secondary: Gibson Dunn and Usercentrics summaries; re-verify on EUR-Lex before the board meeting)EU AI Act

Rate each row on likelihood and impact, show the trend arrow, and name the one action due before the next meeting. The vendor row is the one executives underestimate most: in September 2026 both Claude Code and Codex changed their default model within three weeks of each other, and a loop whose quality depends on the model can shift without anyone deciding to shift it. That is why the vendor row’s indicator is “evals not rerun since the last default change”.

When does a risk skip the quarterly cycle?

Section titled “When does a risk skip the quarterly cycle?”

Agree escalation triggers with the chair in advance, so nobody has to judge in the moment whether something is “board-worthy”. Escalate between meetings when an incident traced to an agent-authored change reaches customers or a regulator, when a loop drops a level after an incident, when an agent credential is revoked for cause, or when spend runs more than a pre-agreed margin over budget. The agent incidents guide covers containment and the postmortem that feeds the next report.

This section of the template is the one that protects you. Each claim below appears in vendor decks and all-hands slides, and each fails under a prepared question.

Do not claimWhy it failsReport instead
“X% of our code is written by AI” as proof of valueA share of lines is not a share of work or of value. DX’s industry figure, a 51.9% average across more than 400 companies (June 2026), measures authorship, not outcomeCost per accepted change and the incident trend
“Pull requests are up, so productivity is up”Faros AI’s telemetry of 22,000 developers (April 2026) shows throughput and quality moving apart: PR merge rate per developer +16.2%, but incidents per pull request +242.7% and median time in review +441.5%Accepted changes and incidents per 100 accepted changes, side by side
“AI makes our developers N% faster”No general measurement supports it. METR’s 2026 follow-up has point estimates in AI’s favour, but the intervals cross zero and METR says selection makes the results hard to interpretLead time for accepted changes in named loops, against a baseline from a designed pilot
“We can cut headcount by a factor of three to five”No third-party measurement supports any productivity multiplierWhere freed capacity went, per the business case
“Anthropic merges over 80% AI-authored code, so we can too”Anthropic’s own figure (“more than 80% of the code we merge”, as of May 2026, Anthropic Institute) is vendor-internal: it describes that company’s harness and test suite, not an industry benchmarkYour own level per loop and the gate that holds it
“DORA says AI improves delivery”DORA’s 2025 report pairs a positive relationship with throughput with a negative relationship with stability; quote both halves or neitherYour change failure rate split by provenance
“Time saved is money saved”Saved hours become cash only when capacity is redeployed or a cost is avoidedThe redeployment decision, recorded with an owner

If an executive wants one number, give cost per accepted change with its four-quarter trend. It includes review, rework and incidents, so it cannot improve by generating more code that nobody can verify.

How do you run the quarterly reporting cycle?

Section titled “How do you run the quarterly reporting cycle?”
  1. Week −3: freeze definitions and export. The engineering operations or platform owner runs the versioned queries for every metric and saves the raw exports, with query IDs, to a folder such as board/2026-q3/. Nothing is computed by hand.

  2. Week −3: loop owners set their levels. Each owner answers the ladder questions for their loop and links the last green run of the gate that justifies the level. A level without a gate link is reported one rung lower.

  3. Week −2: draft with an agent. Use the drafting prompt below in Claude Code, Codex or Cursor. The agent only arranges numbers that already exist in the exports and footnotes the source of each one.

  4. Week −2: attack the draft. Run the sceptical-chair prompt below, then fix or declare every finding. This is the step that catches a flattering trend built on a changed definition.

  5. Week −1: verify and sign. Finance reconciles the cost block with invoices and co-signs it. The CISO signs the security row, legal signs the IP and regulation rows, and the CTO signs the page. The signatures go in the header.

  6. Week −1: send the pre-read. The chair gets the page and the appendix a week ahead. The meeting discusses the decisions in section 6, not the numbers.

  7. After the meeting: record decisions and targets. The target levels per loop become next quarter’s column, and the transformation roadmap gates are updated to match.

The drafting and attacking steps work identically in all three tools: open the agent in the folder that holds the exports and paste the prompt. For an unattended run, use claude -p --permission-mode dontAsk "…" > board/2026-q3/draft.md in Claude Code, or codex exec -s read-only -o board/2026-q3/draft.md "…" in Codex; both only read the exports and print the draft. In Cursor, paste the prompt into the agent with the folder open.

How do you know the report is right before it reaches the board?

Section titled “How do you know the report is right before it reaches the board?”

Nobody should have to re-read the page line by line to trust it. Build the checks into the cycle so each one produces evidence someone signs:

  • Reproducibility. Every number has a query ID in a versioned repository, and a second person reruns the queries and gets the same values. A mismatch blocks the pack.
  • Reconciliation. Finance reconciles total spend with invoices within an agreed tolerance before co-signing.
  • Sampled audit of “accepted”. Draw 10 accepted changes at random and confirm each met its acceptance criteria and was not reverted. One failure means the definition or the data is wrong, and the trend line waits.
  • Provenance coverage. Report the share of changes carrying a provenance label. Below your threshold, the agent/human split is marked “not comparable”.
  • Gate evidence. Each level above L2 links to the last green run of its gate, dated within the quarter.
  • Named sign-offs. CTO, CFO, CISO and general counsel sign their blocks, and the signatures sit in the header, not in an email thread.

The agent drafts and attacks; people sign, the same division this site asks engineers to apply to code.

The metric becomes the target. Teams split work into smaller changes to lower cost per accepted change. Recovery: report median change size beside it; DX measured median pull-request size rising from 44 to 72 lines between July 2025 and June 2026 (June 2026), so a sudden fall in your data is worth a question. Change the unit to accepted outcomes if splitting persists.

The numbers do not reconcile with finance. Vendor dashboards and invoices disagree, usually because of seats bought but unused, metered overage, or a second tool on expenses. Recovery: take spend from invoices only, report the dashboard gap as a note, and move the fix into cost governance.

Attribution breaks. A move to Zero Data Retention switches off Claude Code’s contribution metrics, or a tool switch changes the labels. Recovery: your own provenance label in the pipeline, set by the PR labelling guide, so the series continues across vendors.

A loop drops a level. After an incident, a loop falls from L4 to L3. Report it plainly: the drop is the control working, and hiding it costs more credibility than the incident. Link the postmortem and the new eval that came out of it.

The board asks for “the ROI percentage”. A single ROI figure invites the multipliers the “what not to claim” table rules out. Recovery: give cost per accepted change with its trend and point to the decision recorded in the business case.

The report grows to 12 pages. Recovery: a hard one-page limit, an appendix for everything else, and a rule that a new row replaces an old one.

Frequently asked questions

What should a board report on AI engineering contain?

One page per quarter with four blocks: the autonomy level each critical delivery loop runs at, the cost per accepted change, the quality and incident trend split by provenance, and a five-row risk register covering security, IP, vendor, skills pipeline and regulation. It ends with the decisions the board is asked to take.

Should the board see the percentage of code written by AI?

Not as a measure of value. A share of lines says nothing about accepted outcomes, and industry figures such as DX's 51.9% average (June 2026) are a share of code, not of work. Report cost per accepted change and the incident trend instead.

Who signs off the AI engineering board report?

The CTO owns the page. Finance co-signs the cost block after reconciling it with invoices, the CISO signs the security row, and legal signs the IP and regulation rows. Every number carries the query that produced it.

Edit page

Last updated:

Cite this page — https://developertoolkit.ai/en/strategy/board-reporting/, developertoolkit.ai