Skip to content

CTO and VP Engineering track: an operating model for agent-built software

The CTO and VP Engineering track is the reading path on AI Developer Toolkit for engineering leaders who standardize how an organization builds software with Claude Code, Codex and Cursor. It orders ten decisions, from evidence and a measured pilot to autonomy, security, spend, regulation and board reporting, and each decision ends in a named artifact with an owner.

Every team already runs agents, invoices are rising, and security asks who approved an agent that opens pull requests. The board wants a number; you have seat counts.

What the CTO track produces for your organization

Section titled “What the CTO track produces for your organization”
  • An artifact and an exit criterion per step, so “done” is evidence, not a meeting.
  • A pilot before a platform. You measure one team before you staff a platform team or sign a larger contract.

Agents increase the volume of change faster than your ability to verify it. The 2025 DORA report (Google Cloud, 23 September 2025) links AI adoption to higher delivery throughput and lower delivery stability. Faros AI’s telemetry (April 2026) shows tasks per developer up 33.7% alongside incidents per pull request up 242.7% and median time in review up 441.5%. So measurement comes before the operating model.

Which artifact does each CTO decision produce?

Section titled “Which artifact does each CTO decision produce?”
#DecisionArtifact it producesCanonical page
1What does the evidence say, and where do our teams sit?Evidence brief with three dated sources; org map by ladder levelState of agentic engineering, one map, evidence, not diffs
2What do we measure, and how do we prove it?Metric definitions v1, signed by engineering and finance; pilot design v1 with a control group and a stop ruleMetrics frameworks, pilot design
3Who owns the harness, the gates and the evals?Operating model v1: a RACI with named ownersOperating model
4What does an accepted change cost?Cost-per-accepted-change baseline for one teamEconomics
5Which change classes may merge on evidence alone?Autonomy register v1Governance and autonomy
6What can an agent reach, and as whom?Agent threat model v1; one scoped, attributable identity per agentThreat model, identity and secrets
7How is policy enforced, and who runs the platform?Managed policy per approved tool; platform team charterManaged policy, platform team
8Is spend visible, and can we leave a vendor?Spend per team and per accepted change on one dashboard; exit test for the primary toolCost governance, vendor risk
9Which regulatory duties apply to us?EU AI Act obligations checklist reviewed by legalEU AI Act
10What do we report upward?One-page quarterly board updateBoard reporting

Published steps follow this order:

  1. Why now: The dated third-party evidence, publisher by publisher, before any decision.

    Done when: You can cite three sources with dates for the shift you are planning.

  2. Why now: One map of maturity, process and capability that every team in the org can place itself on.

    Done when: The organization described as a distribution of merged change across levels.

  3. Why now: The operating principle: humans judge evidence and risk, not every line.

    Done when: The principle adopted as policy for at least one risk class.

  4. Why now: Pick the measurement frame before the pilot, so the pilot can prove something.

    Done when: Metric definitions signed off by engineering and finance.

  5. Why now: A pilot with a control group and a stop rule, not a demo.

    Done when: Pilot design v1: teams, duration, metrics and the decision it feeds.

  6. Why now: Decide who owns the harness, the gates and the evals across teams.

    Done when: An operating model v1 with named owners.

  7. Why now: Cost per accepted change, not cost per seat, is the number to manage.

    Done when: A cost-per-accepted-change baseline for one team.

  8. Why now: Autonomy is granted by risk class, with humans at the gates.

    Done when: An autonomy register v1: which change classes may merge on evidence.

  9. Why now: Prompt injection and blast radius are new threats; model them before scaling.

    Done when: A threat model reviewed by security for the agent pipeline.

  10. Why now: Agents need their own identities and scoped credentials, never a developer token.

    Done when: Every agent credential scoped, rotated and attributable.

  11. Why now: Enforce one policy across every coding agent the organization runs.

    Done when: Managed settings deployed for each approved tool.

  12. Why now: Run the harness as a product, with a team that owns it.

    Done when: A platform team charter and backlog.

  13. 13 AI usage cost governance Subscribers

    Why now: Make spend visible and attributable before it becomes a budget surprise.

    Done when: Monthly spend per team and per accepted change on one dashboard.

  14. Why now: Test continuity with each vendor, not its logo.

    Done when: An exit test run for your primary tool.

  15. Why now: Know which obligations apply to building software with agents in the EU.

    Done when: An obligations checklist reviewed by legal.

  16. Why now: Report outcomes, risk and spend to the board in a form it can act on.

    Done when: A one-page board update template filled with your numbers.

Where does each tool enforce the organization’s policy?

Section titled “Where does each tool enforce the organization’s policy?”
ToolWhere the organization’s policy lives
Claude CodeManaged settings: a managed-settings.json file, an MDM policy, or server-managed settings from claude.ai (Team and Enterprise). permissions.disableAutoMode: "disable" turns auto mode off for everyone.
Codexrequirements.toml holds administrator constraints (for example approval_policy, mcp_servers, allow_managed_hooks_only), kept separate from the config.toml defaults.
CursorAdmin settings on Cursor’s business plans; its SDK exposes team and mdm settings layers (@cursor/sdk 1.0.32). Confirm setting names on cursor.com.

Claude Code facts checked on 26 September 2026 against version 2.1.283; Codex facts against version 0.157.1.

How do you prove the operating model works?

Section titled “How do you prove the operating model works?”

You do not read agent diffs; auditable evidence proves quality. Four checks close each quarter:

  • Accepted change, not activity. The metric definitions from decision 2 count merged, non-reverted changes that passed their gates, never lines of code or ”% AI-written”.
  • Stability beside throughput. Change failure rate and incidents per change sit beside throughput.
  • Autonomy tied to evidence. Each class in the autonomy register names the tests, evals and review agents that must pass, and the human who signs off on production.
  • Artifacts with owners and dates. A decision record past its review date is a finding in the board update.
  • The pilot proves nothing. Without a baseline or control group it shows only that people like the tool. Freeze the rollout and rerun it with the design from decision 2.
  • Autonomy before controls. Agents got merge rights before the threat model and scoped identities existed. Revoke agent credentials, then re-grant them per risk class.
  • Percent of code as the headline. DX reports a self-reported average of 51.9% AI-authored code across more than 400 companies (Q2 2026, published 17 June 2026), so the share alone does not distinguish you. Report cost per accepted change and stability instead.
  • The platform team becomes a gate. Give it service levels to product teams and a public backlog.