The operating model: who owns the harness, the gates and the evals
An operating model for agent-built software names one accountable owner for every artifact that steers or checks the agents: shared rules, skills, hooks, the MCP registry, agent identities, eval suites, the autonomy register, unattended loops and the production gate. Product teams own intent and acceptance, a platform function owns the harness, and security and the release owner keep the gates.
Six teams adopted Claude Code, Codex and Cursor in the same quarter, and now every repository has its own CLAUDE.md, three copies of a “secure review” skill have diverged, and an MCP server runs on someone’s personal token. On Monday a hook update blocks every commit in two teams, and nobody can say who is allowed to roll it back. This is an ownership problem, not a tooling one, and each agent you add makes it bigger.
This page is for the CTO or VP Engineering who has to answer “who owns this?” for each piece, and for the executive who wants to know whether anyone does.
What this operating model gives you
Section titled “What this operating model gives you”- A RACI table for the eleven artifacts that control agent work, with exactly one accountable owner per row.
- A map of where each artifact lives and is enforced in Claude Code, Codex and Cursor.
- An autonomy register template: one row per loop, with its level, risk class, oracle, identity and owner.
- Decision criteria for when a named owner is enough and when a platform team is justified.
- What a platform team does to review capacity, and how to stop it becoming the new bottleneck.
- A one-page charter template you can adopt as-is, and five questions an executive can ask the CTO.
What is the harness, and why does each part need an owner?
Section titled “What is the harness, and why does each part need an owner?”The harness is everything around the model that makes its output predictable and checkable: instructions, reusable procedures, runtime controls, tool access, credentials, tests and evals, and the gates that decide what reaches production. The artifact chain describes how work flows through it. The operating model decides who is answerable when each link fails.
Every artifact in the table below is shared state. One change reaches every agent session that loads it, which is why “whoever touched it last” is not an owner.
| Artifact | What it controls | What goes wrong without an owner |
|---|---|---|
Repository rules (CLAUDE.md, AGENTS.md, Cursor Rules) | What the agent knows about this codebase | Contradictory rules; agents follow the wrong one |
| Managed policy | Permissions, allowed models, allowed hooks, MCP servers and plugin marketplaces for everyone | A different effective policy per developer |
| Shared skills and plugins | Reusable procedures (review, migration, release notes) | Drifting forks; one broken skill hits every team |
| Shared hooks | Deterministic checks at tool-call time | All work blocked, or everything silently allowed |
| MCP registry | Which tools and data sources agents can reach | Unvetted servers with broad tokens |
| Agent identities and credentials | What an agent can do outside the sandbox | Loops on a person’s token |
| Oracles (acceptance criteria and tests) | Whether a change does what was asked | Agents pass by weakening tests |
| Harness eval suites | Whether a rule, skill, hook or model change makes agents better or worse | Regressions surface as incidents |
| Autonomy register | Which loop runs at which ladder level and risk class | Autonomy rises by accident |
| Unattended loops | Scheduled and event-triggered agent runs | Orphaned jobs keep spending |
| Production gate | What reaches users, and who authorized it | Authors approve their own changes |
The RACI: who owns each harness artifact?
Section titled “The RACI: who owns each harness artifact?”Read the table by row. A (accountable) is the one person or role who answers for the artifact and can say no; there is exactly one per row. R (responsible) does the work. C (consulted) reviews before a change. I (informed) hears after. Agents are never A for anything, and never R for a gate.
| Artifact | Product team (tech lead) | Platform owner or team | Security | Release owner (service owner) | CTO |
|---|---|---|---|---|---|
| Repository rules | A, R | C | I | — | — |
| Managed policy | C | R | A | — | I |
| Shared skills and plugins | C, contributes | A, R | C | — | — |
| Shared hooks | C | A, R | C | I | — |
| MCP registry (allowlist) | C, requests | R | A | — | I |
| Agent identities and credentials | I | R | A | C | — |
| Oracles (acceptance criteria, tests) | A, R | C | C (security tests) | C | — |
| Harness eval suites | R (task cases) | A, R | C | — | I |
| Autonomy register | R (proposes changes) | C | C | C | A |
| Unattended loops | A, R (per loop) | C | C | I | — |
| Production gate | C | C | C | A, R | I |
Four rules make the table work in practice:
- The team that ships the product owns what “correct” means. Acceptance criteria and tests stay with the product team because only they can say what the behaviour should be. The platform team can supply tooling to protect the oracle, but it cannot decide what the oracle checks.
- Security is accountable for reach, the platform for behaviour. Whatever widens what an agent can touch (MCP servers, credentials, managed permissions) needs security’s yes. Whatever changes how agents work inside that reach (skills, hooks, evals) belongs to the platform owner. Identities themselves are defined in agent identity, credentials and secrets.
- Raising autonomy is a leadership decision. Moving a loop up the ladder changes the organization’s risk, so the CTO (or a delegated head of engineering) is accountable for the register. Teams propose; the register records the evidence.
- The production gate has a human owner who did not author the change. An author agent plus a reviewer agent is not separation of duties: they can share credentials, configuration and blind spots. The release owner approves through a platform control, as governance and autonomy describes.
Where each owned artifact lives in Claude Code, Codex and Cursor
Section titled “Where each owned artifact lives in Claude Code, Codex and Cursor”The RACI is tool-neutral; enforcement is not. The platform owner needs to know which file or console makes each row binding, because an instruction in a rules file is advice and a managed setting is policy.
Policy layer. Managed settings rank first in the settings precedence, above command-line arguments, project and user settings. They are delivered as managed-settings.json, through MDM, or as server-managed settings from the claude.ai console (Team and Enterprise).
Keys that map to RACI rows (all present in the Claude Code changelog through v2.1.283):
- Shared hooks:
allowManagedHooksOnlyignores user and project hooks. Hooks from plugins that managed settings force-enable still run, so the platform team ships hooks inside its own plugin. - Skills and plugins:
strictKnownMarketplacesandblockedMarketplacesrestrict plugin sources;enabledPluginsturns the platform’s plugins on. - MCP registry:
allowManagedMcpServersOnlyanddeniedMcpServers. - Managed permissions:
allowManagedPermissionRulesOnly;permissions.disableAutoModeis the organization off-switch for auto mode. - Models:
availableModelsandenforceAvailableModels.
{ "strictKnownMarketplaces": [ { "source": "github", "repo": "anthropics/claude-plugins-official" }, { "source": "github", "repo": "acme/*" } ], "extraKnownMarketplaces": { "acme-plugins": { "source": { "source": "github", "repo": "acme/acme-plugins" }, "autoUpdate": true } }, "enabledPlugins": { "acme-harness@acme-plugins": true }, "allowManagedHooksOnly": true}acme/acme-plugins is the platform team’s marketplace repository and acme-harness its plugin carrying the shared skills and hooks. Evals: claude plugin eval runs a plugin’s eval cases and reports scored results against a no-plugin baseline, which gives the platform owner a gate for every skill or hook release.
Policy layer. Codex separates defaults from constraints. config.toml holds defaults; requirements.toml holds the constraints an administrator enforces. OpenAI’s own guidance: “Keep config.toml defaults, requirements.toml constraints, and managed or administrator policy separate.”
Keys that map to RACI rows (from config_requirements.rs on openai/codex, checked 2026-09-26): allow_managed_hooks_only and managed_hooks for shared hooks; mcp_servers for the MCP registry; plugins and marketplaces for skills and plugins; permission_profile, approval_policy and approvals_reviewer for managed permissions; enforce_residency for data location.
# requirements.toml: ignore user, project and session hook configsallow_managed_hooks_only = trueSetting allow_managed_hooks_only in config.toml does nothing; OpenAI documents it as a requirements.toml-only key. Repository rules live in AGENTS.md. Since Codex 0.150.0, untrusted projects do not supply project-level AGENTS.md, so the trust decision is itself part of the policy.
Policy layer. Cursor packages the same artifact types: Rules, Agent Skills, Hooks, Plugins (“package rules, skills, agents, commands, MCP servers, and hooks”) and MCP (Cursor docs, checked 2026-08-28). For the gate rows, Bugbot reviews pull requests, and PR Routing & Approval “assigns reviewers based on code ownership and commit history, and can approve low-risk PRs when your criteria are met.”
That last feature moves part of the production-gate row into a tool, so the RACI has to say who owns the approval criteria. Put it with the release owner, and treat an auto-approval rule as a change to the autonomy register.
Cursor’s team admin controls (last checked 2026-08-28) change often; confirm the enforcement options in Cursor’s admin docs before you write them into the charter. The page on enforcing one policy across every coding agent compares the three tools side by side.
The same pattern holds in all three tools: the platform owner publishes skills, hooks and MCP configuration through one versioned channel (a plugin marketplace or an internal repository), and managed policy makes that channel the only one allowed. See running a team plugin marketplace and MCP registries and gateways for the distribution side.
What goes in the autonomy register?
Section titled “What goes in the autonomy register?”The autonomy register is one row per agent loop: a named workflow such as “dependency upgrades in the payments service” or “issue-to-PR for UI copy changes”. It records the loop’s level on the autonomy ladder and the risk class of the changes it makes. Levels measure trust in the loop; risk classes (low, medium, high, critical) measure what a mistake costs. Keeping them apart stops a mature loop from earning its way into critical changes by track record alone.
| Loop | Owner (A) | Level | Risk class | Oracle | Eval suite | Agent identity | Stop condition | Next review |
|---|---|---|---|---|---|---|---|---|
Dependency patch upgrades, web-app | Web tech lead | L4 | low | Full test suite plus lockfile audit | evals/deps (12 cases) | bot-deps-web (repo write, no deploy) | Two CI rounds, then hand back | 2026-12-01 |
| Issue-to-PR, UI copy | Growth tech lead | L3 | low | Visual snapshot plus copy linter | evals/copy (8 cases) | bot-issues-growth | One PR per issue; human merge | 2026-11-15 |
| Payments refactor | Payments tech lead | L2 | critical | Characterization tests plus property tests | none yet | developer session only | Human approves every plan | on request |
The stop condition in the first row borrows Stripe’s published bound for its Minions: “at most two rounds of CI” before the branch goes back to a human (Stripe engineering blog, 2026-02-09). Every row needs a stop condition like it, a named identity, and an eval suite before the level can rise above L3.
A change to any row goes through the CTO or the delegate named in the charter, with the evidence attached: eval results, the loop’s accepted-change rate, and any incidents. Unattended agent runs covers the operating rules for the loops themselves.
When is a platform team justified?
Section titled “When is a platform team justified?”A platform team is justified when harness artifacts are shared across teams and a mistake in one reaches all of them. Before that point, a named owner is enough. DORA’s 2025 report found that “90% of organizations have adopted at least one platform and there is a direct correlation between a high quality internal platform and an organization’s ability to unlock the value of AI” (Google Cloud, 2025-09-23). That is correlation, not causation, but it puts the platform in the critical path.
Stripe shows what the platform owns at scale: its Toolshed “currently contains nearly 500 MCP tools for internal systems and SaaS platforms” (Stripe engineering blog, 2026-02-19), one catalogue shared by every Minion run.
Use this decision table. The thresholds are this guide’s working rule, not a published benchmark; adjust them to your risk profile.
| Your situation | Ownership model | What the owner does |
|---|---|---|
| One team, one or two repositories, no unattended loops | Named harness owner: the tech lead, a few hours a week | Keeps rules, skills and hooks in the repository; runs evals before changing them |
| Two to three teams sharing skills or MCP servers, or the first unattended loop with write access | Virtual platform group: one owner per artifact type from the teams, a shared repository, a budget line | Publishes one marketplace, owns the eval suites, keeps the register |
| Three or more teams duplicating harness work, a second agent vendor, unattended loops in several repositories, or an auditor asking for evidence the teams cannot produce | Dedicated agent platform team | Runs the harness as a product; see the agent platform team |
| Regulated workloads or critical-class changes handled by agents | Dedicated team, with security embedded or on call for MCP and identity changes | Adds control evidence for audit and change management |
Before you staff a team, run a pilot that proves something on one loop. A platform built before a pilot tends to standardize guesses.
What does a platform team do to review capacity?
Section titled “What does a platform team do to review capacity?”Faros AI’s 2025 “AI Productivity Paradox” analysis of more than 10,000 developers reported that developers on high-AI teams merged about 98% more pull requests while PR review time rose about 91% (Faros AI, 2025; secondary source). A platform team fixes that by making most review unnecessary for most changes, not by reviewing more, and by staying out of the product review path.
-
Move checks from reading to gates. The platform team ships the hooks, fitness functions and eval suites that catch what reviewers used to catch by eye, and the evidence bundle that tells a reviewer what was verified.
-
Route by risk class. Low-class changes with a complete evidence bundle get a lighter review; high and critical changes get the release owner. The routing rules sit in the register, not in individual reviewers’ heads.
-
Cap agent work in progress per reviewer. A loop that opens more pull requests than its owner can review is paying to grow a queue, so its stop condition caps open pull requests.
-
Review the harness, not the product. The platform team approves changes to shared skills, hooks, policy and evals. It never sits in a product pull request’s approval path; if it does, it has become the gate this model exists to avoid.
-
Measure the queue. Track review wait time per risk class and the share of agent pull requests merged with a complete evidence bundle. Definitions live in metrics frameworks; the day-to-day practice is in running the review queue.
How do you prove the operating model works?
Section titled “How do you prove the operating model works?”A RACI on a wiki proves nothing. The model works when each accountable owner can show evidence without anyone reading every diff:
- Harness changes pass evals before rollout. A skill, hook, policy or model change runs the harness eval suite and is compared with the current version; the platform owner signs off on the result, not the diff. See evals for coding agents.
- Every agent action traces to an identity and a loop. Commits, pull requests and deploys carry the agent identity and the register row, so the release owner and auditors can see which loop produced what.
- Gates are platform controls. Branch protection, required checks and deployment approvals are enforced by your Git host (GitHub, GitLab) and the deploy system, and each has an A in the RACI.
- The register matches reality. A monthly job lists every scheduled or event-triggered agent run, and anything absent from the register is paused.
Rehearse it once a quarter, the way you rehearse a restore:
- Change a shared skill in a branch and confirm that the eval gate blocks a regression.
- Revoke one agent identity and confirm that its loops stop and no human loses access.
- Try to install an unlisted MCP server on a managed machine and confirm that policy blocks it.
- Propose raising a loop by one level and confirm that the change needs the named approver and the evidence.
- Roll back a shared hook and time how long it takes to reach every developer.
Copy-paste prompts to map your current harness
Section titled “Copy-paste prompts to map your current harness”Run these in Claude Code, Codex or Cursor from the root of a repository. They work the same way in all three tools, because they only read files and write a report.
One-page charter template for the operating model
Section titled “One-page charter template for the operating model”Copy this into your engineering handbook, fill in the names, and have the CTO sign it. Review it every quarter or after any agent incident.
# Agent operating model charter: <organization>, v<N>, <date>
## PurposeAgents write and change code here under one policy. This charter names who ownseach artifact that steers or checks them, and how autonomy is raised.
## Accountable owners<Paste the RACI table from the operating-model page. Put a named person, and adeputy, in every A cell.>
## Rules1. Agents are never accountable and never approve a gate.2. Shared skills, hooks, MCP configuration and policy ship through <channel> only, and pass the harness eval suite before rollout.3. Every unattended loop has a register row, a named identity, an eval suite and a stop condition before it runs. Unregistered loops are paused.4. Raising a loop's level or risk class needs <approver> and attached evidence.5. The platform group reviews harness changes, never product pull requests.
## Service levels- Harness change requests answered within <N> working days.- Broken shared hook or skill rolled back within <N> hours.- Agent identity revoked within <N> minutes of an incident call.
## Evidence we keepEval results per harness release; register history; identity-to-commit trace;gate approvals; quarterly rehearsal results.
## ReviewQuarterly, and after every agent-caused incident. Signed: <CTO>, <security lead>.What breaks in an agent operating model, and how to recover
Section titled “What breaks in an agent operating model, and how to recover”The platform team becomes the review gate. Product pull requests wait on platform approval. Recovery: remove the platform team from product CODEOWNERS, restate rule 5 of the charter, and move the check they were doing by hand into a hook or eval they own.
Everyone owns the rules, so nobody does. Repository rules grow to hundreds of lines and contradict the shared skills. Recovery: one tech lead is A per repository; move procedures out of rules into skills, and run the harness evals before merging a rules change. See shared agent rules.
A loop runs on a person’s token. When that person leaves, the loop fails or keeps running with their full access. Recovery: issue a per-loop identity, record it in the register, revoke the personal token. The procedure is in agent identity, credentials and secrets.
A model or client change ships without evals. Defaults move under you. On 2026-09-26, Claude Code’s stable release channel (2.1.274) still gave Pro and Team Standard seats Sonnet 5 as the default model, while the latest channel defaulted to Opus 5.5 from v2.1.280. Two teams on different channels can run different models under the same policy. Recovery: the platform owner pins the channel with autoUpdatesChannel ("stable" or "latest") and the models with availableModels plus enforceAvailableModels in managed settings, and re-runs the harness evals before moving them; see the models hub for current defaults.
The register drifts from reality. New scheduled jobs appear without rows. Recovery: the monthly inventory job from the second prompt above, and a standing rule that unregistered loops are paused, not grandfathered.
A broken shared hook stops everyone and no one can roll it back. Recovery: every hook release is versioned with a named rollback owner and a tested rollback; the quarterly rehearsal times it. See shared hooks governance. When an agent does cause an incident, the postmortem names the failed row of this RACI; when an agent causes an incident covers the procedure.
Five questions an executive can ask the CTO
Section titled “Five questions an executive can ask the CTO”- For each loop that runs without a human in the session, who is accountable, and where is it written down?
- Which agent identities can write to production or customer data, and how fast can we revoke them?
- When we change a model, skill or policy, what evidence tells us it did not get worse?
- Who can approve a production change, and can that person be the one who asked the agent for it?
- Is our platform group making review faster or adding a queue? Show the review wait time by risk class.
Where to go next with the operating model
Section titled “Where to go next with the operating model”The CTO track reaches this page after the pilot design and continues with the economics of agent-built software, which puts a cost on each accountable row.