Skip to content

The operating model: who owns the harness, the gates and the evals

An operating model for agent-built software names one accountable owner for every artifact that steers or checks the agents: shared rules, skills, hooks, the MCP registry, agent identities, eval suites, the autonomy register, unattended loops and the production gate. Product teams own intent and acceptance, a platform function owns the harness, and security and the release owner keep the gates.

Six teams adopted Claude Code, Codex and Cursor in the same quarter, and now every repository has its own CLAUDE.md, three copies of a “secure review” skill have diverged, and an MCP server runs on someone’s personal token. On Monday a hook update blocks every commit in two teams, and nobody can say who is allowed to roll it back. This is an ownership problem, not a tooling one, and each agent you add makes it bigger.

This page is for the CTO or VP Engineering who has to answer “who owns this?” for each piece, and for the executive who wants to know whether anyone does.

  • A RACI table for the eleven artifacts that control agent work, with exactly one accountable owner per row.
  • A map of where each artifact lives and is enforced in Claude Code, Codex and Cursor.
  • An autonomy register template: one row per loop, with its level, risk class, oracle, identity and owner.
  • Decision criteria for when a named owner is enough and when a platform team is justified.
  • What a platform team does to review capacity, and how to stop it becoming the new bottleneck.
  • A one-page charter template you can adopt as-is, and five questions an executive can ask the CTO.

What is the harness, and why does each part need an owner?

Section titled “What is the harness, and why does each part need an owner?”

The harness is everything around the model that makes its output predictable and checkable: instructions, reusable procedures, runtime controls, tool access, credentials, tests and evals, and the gates that decide what reaches production. The artifact chain describes how work flows through it. The operating model decides who is answerable when each link fails.

Every artifact in the table below is shared state. One change reaches every agent session that loads it, which is why “whoever touched it last” is not an owner.

ArtifactWhat it controlsWhat goes wrong without an owner
Repository rules (CLAUDE.md, AGENTS.md, Cursor Rules)What the agent knows about this codebaseContradictory rules; agents follow the wrong one
Managed policyPermissions, allowed models, allowed hooks, MCP servers and plugin marketplaces for everyoneA different effective policy per developer
Shared skills and pluginsReusable procedures (review, migration, release notes)Drifting forks; one broken skill hits every team
Shared hooksDeterministic checks at tool-call timeAll work blocked, or everything silently allowed
MCP registryWhich tools and data sources agents can reachUnvetted servers with broad tokens
Agent identities and credentialsWhat an agent can do outside the sandboxLoops on a person’s token
Oracles (acceptance criteria and tests)Whether a change does what was askedAgents pass by weakening tests
Harness eval suitesWhether a rule, skill, hook or model change makes agents better or worseRegressions surface as incidents
Autonomy registerWhich loop runs at which ladder level and risk classAutonomy rises by accident
Unattended loopsScheduled and event-triggered agent runsOrphaned jobs keep spending
Production gateWhat reaches users, and who authorized itAuthors approve their own changes

Read the table by row. A (accountable) is the one person or role who answers for the artifact and can say no; there is exactly one per row. R (responsible) does the work. C (consulted) reviews before a change. I (informed) hears after. Agents are never A for anything, and never R for a gate.

ArtifactProduct team (tech lead)Platform owner or teamSecurityRelease owner (service owner)CTO
Repository rulesA, RCI——
Managed policyCRA—I
Shared skills and pluginsC, contributesA, RC——
Shared hooksCA, RCI—
MCP registry (allowlist)C, requestsRA—I
Agent identities and credentialsIRAC—
Oracles (acceptance criteria, tests)A, RCC (security tests)C—
Harness eval suitesR (task cases)A, RC—I
Autonomy registerR (proposes changes)CCCA
Unattended loopsA, R (per loop)CCI—
Production gateCCCA, RI

Four rules make the table work in practice:

  1. The team that ships the product owns what “correct” means. Acceptance criteria and tests stay with the product team because only they can say what the behaviour should be. The platform team can supply tooling to protect the oracle, but it cannot decide what the oracle checks.
  2. Security is accountable for reach, the platform for behaviour. Whatever widens what an agent can touch (MCP servers, credentials, managed permissions) needs security’s yes. Whatever changes how agents work inside that reach (skills, hooks, evals) belongs to the platform owner. Identities themselves are defined in agent identity, credentials and secrets.
  3. Raising autonomy is a leadership decision. Moving a loop up the ladder changes the organization’s risk, so the CTO (or a delegated head of engineering) is accountable for the register. Teams propose; the register records the evidence.
  4. The production gate has a human owner who did not author the change. An author agent plus a reviewer agent is not separation of duties: they can share credentials, configuration and blind spots. The release owner approves through a platform control, as governance and autonomy describes.

Where each owned artifact lives in Claude Code, Codex and Cursor

Section titled “Where each owned artifact lives in Claude Code, Codex and Cursor”

The RACI is tool-neutral; enforcement is not. The platform owner needs to know which file or console makes each row binding, because an instruction in a rules file is advice and a managed setting is policy.

Policy layer. Managed settings rank first in the settings precedence, above command-line arguments, project and user settings. They are delivered as managed-settings.json, through MDM, or as server-managed settings from the claude.ai console (Team and Enterprise).

Keys that map to RACI rows (all present in the Claude Code changelog through v2.1.283):

  • Shared hooks: allowManagedHooksOnly ignores user and project hooks. Hooks from plugins that managed settings force-enable still run, so the platform team ships hooks inside its own plugin.
  • Skills and plugins: strictKnownMarketplaces and blockedMarketplaces restrict plugin sources; enabledPlugins turns the platform’s plugins on.
  • MCP registry: allowManagedMcpServersOnly and deniedMcpServers.
  • Managed permissions: allowManagedPermissionRulesOnly; permissions.disableAutoMode is the organization off-switch for auto mode.
  • Models: availableModels and enforceAvailableModels.
{
"strictKnownMarketplaces": [
{ "source": "github", "repo": "anthropics/claude-plugins-official" },
{ "source": "github", "repo": "acme/*" }
],
"extraKnownMarketplaces": {
"acme-plugins": { "source": { "source": "github", "repo": "acme/acme-plugins" }, "autoUpdate": true }
},
"enabledPlugins": { "acme-harness@acme-plugins": true },
"allowManagedHooksOnly": true
}

acme/acme-plugins is the platform team’s marketplace repository and acme-harness its plugin carrying the shared skills and hooks. Evals: claude plugin eval runs a plugin’s eval cases and reports scored results against a no-plugin baseline, which gives the platform owner a gate for every skill or hook release.

The same pattern holds in all three tools: the platform owner publishes skills, hooks and MCP configuration through one versioned channel (a plugin marketplace or an internal repository), and managed policy makes that channel the only one allowed. See running a team plugin marketplace and MCP registries and gateways for the distribution side.

The autonomy register is one row per agent loop: a named workflow such as “dependency upgrades in the payments service” or “issue-to-PR for UI copy changes”. It records the loop’s level on the autonomy ladder and the risk class of the changes it makes. Levels measure trust in the loop; risk classes (low, medium, high, critical) measure what a mistake costs. Keeping them apart stops a mature loop from earning its way into critical changes by track record alone.

LoopOwner (A)LevelRisk classOracleEval suiteAgent identityStop conditionNext review
Dependency patch upgrades, web-appWeb tech leadL4lowFull test suite plus lockfile auditevals/deps (12 cases)bot-deps-web (repo write, no deploy)Two CI rounds, then hand back2026-12-01
Issue-to-PR, UI copyGrowth tech leadL3lowVisual snapshot plus copy linterevals/copy (8 cases)bot-issues-growthOne PR per issue; human merge2026-11-15
Payments refactorPayments tech leadL2criticalCharacterization tests plus property testsnone yetdeveloper session onlyHuman approves every planon request

The stop condition in the first row borrows Stripe’s published bound for its Minions: “at most two rounds of CI” before the branch goes back to a human (Stripe engineering blog, 2026-02-09). Every row needs a stop condition like it, a named identity, and an eval suite before the level can rise above L3.

A change to any row goes through the CTO or the delegate named in the charter, with the evidence attached: eval results, the loop’s accepted-change rate, and any incidents. Unattended agent runs covers the operating rules for the loops themselves.

A platform team is justified when harness artifacts are shared across teams and a mistake in one reaches all of them. Before that point, a named owner is enough. DORA’s 2025 report found that “90% of organizations have adopted at least one platform and there is a direct correlation between a high quality internal platform and an organization’s ability to unlock the value of AI” (Google Cloud, 2025-09-23). That is correlation, not causation, but it puts the platform in the critical path.

Stripe shows what the platform owns at scale: its Toolshed “currently contains nearly 500 MCP tools for internal systems and SaaS platforms” (Stripe engineering blog, 2026-02-19), one catalogue shared by every Minion run.

Use this decision table. The thresholds are this guide’s working rule, not a published benchmark; adjust them to your risk profile.

Your situationOwnership modelWhat the owner does
One team, one or two repositories, no unattended loopsNamed harness owner: the tech lead, a few hours a weekKeeps rules, skills and hooks in the repository; runs evals before changing them
Two to three teams sharing skills or MCP servers, or the first unattended loop with write accessVirtual platform group: one owner per artifact type from the teams, a shared repository, a budget linePublishes one marketplace, owns the eval suites, keeps the register
Three or more teams duplicating harness work, a second agent vendor, unattended loops in several repositories, or an auditor asking for evidence the teams cannot produceDedicated agent platform teamRuns the harness as a product; see the agent platform team
Regulated workloads or critical-class changes handled by agentsDedicated team, with security embedded or on call for MCP and identity changesAdds control evidence for audit and change management

Before you staff a team, run a pilot that proves something on one loop. A platform built before a pilot tends to standardize guesses.

What does a platform team do to review capacity?

Section titled “What does a platform team do to review capacity?”

Faros AI’s 2025 “AI Productivity Paradox” analysis of more than 10,000 developers reported that developers on high-AI teams merged about 98% more pull requests while PR review time rose about 91% (Faros AI, 2025; secondary source). A platform team fixes that by making most review unnecessary for most changes, not by reviewing more, and by staying out of the product review path.

  1. Move checks from reading to gates. The platform team ships the hooks, fitness functions and eval suites that catch what reviewers used to catch by eye, and the evidence bundle that tells a reviewer what was verified.

  2. Route by risk class. Low-class changes with a complete evidence bundle get a lighter review; high and critical changes get the release owner. The routing rules sit in the register, not in individual reviewers’ heads.

  3. Cap agent work in progress per reviewer. A loop that opens more pull requests than its owner can review is paying to grow a queue, so its stop condition caps open pull requests.

  4. Review the harness, not the product. The platform team approves changes to shared skills, hooks, policy and evals. It never sits in a product pull request’s approval path; if it does, it has become the gate this model exists to avoid.

  5. Measure the queue. Track review wait time per risk class and the share of agent pull requests merged with a complete evidence bundle. Definitions live in metrics frameworks; the day-to-day practice is in running the review queue.

How do you prove the operating model works?

Section titled “How do you prove the operating model works?”

A RACI on a wiki proves nothing. The model works when each accountable owner can show evidence without anyone reading every diff:

  • Harness changes pass evals before rollout. A skill, hook, policy or model change runs the harness eval suite and is compared with the current version; the platform owner signs off on the result, not the diff. See evals for coding agents.
  • Every agent action traces to an identity and a loop. Commits, pull requests and deploys carry the agent identity and the register row, so the release owner and auditors can see which loop produced what.
  • Gates are platform controls. Branch protection, required checks and deployment approvals are enforced by your Git host (GitHub, GitLab) and the deploy system, and each has an A in the RACI.
  • The register matches reality. A monthly job lists every scheduled or event-triggered agent run, and anything absent from the register is paused.

Rehearse it once a quarter, the way you rehearse a restore:

  1. Change a shared skill in a branch and confirm that the eval gate blocks a regression.
  2. Revoke one agent identity and confirm that its loops stop and no human loses access.
  3. Try to install an unlisted MCP server on a managed machine and confirm that policy blocks it.
  4. Propose raising a loop by one level and confirm that the change needs the named approver and the evidence.
  5. Roll back a shared hook and time how long it takes to reach every developer.

Copy-paste prompts to map your current harness

Section titled “Copy-paste prompts to map your current harness”

Run these in Claude Code, Codex or Cursor from the root of a repository. They work the same way in all three tools, because they only read files and write a report.

One-page charter template for the operating model

Section titled “One-page charter template for the operating model”

Copy this into your engineering handbook, fill in the names, and have the CTO sign it. Review it every quarter or after any agent incident.

# Agent operating model charter: <organization>, v<N>, <date>
## Purpose
Agents write and change code here under one policy. This charter names who owns
each artifact that steers or checks them, and how autonomy is raised.
## Accountable owners
<Paste the RACI table from the operating-model page. Put a named person, and a
deputy, in every A cell.>
## Rules
1. Agents are never accountable and never approve a gate.
2. Shared skills, hooks, MCP configuration and policy ship through <channel>
only, and pass the harness eval suite before rollout.
3. Every unattended loop has a register row, a named identity, an eval suite and
a stop condition before it runs. Unregistered loops are paused.
4. Raising a loop's level or risk class needs <approver> and attached evidence.
5. The platform group reviews harness changes, never product pull requests.
## Service levels
- Harness change requests answered within <N> working days.
- Broken shared hook or skill rolled back within <N> hours.
- Agent identity revoked within <N> minutes of an incident call.
## Evidence we keep
Eval results per harness release; register history; identity-to-commit trace;
gate approvals; quarterly rehearsal results.
## Review
Quarterly, and after every agent-caused incident. Signed: <CTO>, <security lead>.

What breaks in an agent operating model, and how to recover

Section titled “What breaks in an agent operating model, and how to recover”

The platform team becomes the review gate. Product pull requests wait on platform approval. Recovery: remove the platform team from product CODEOWNERS, restate rule 5 of the charter, and move the check they were doing by hand into a hook or eval they own.

Everyone owns the rules, so nobody does. Repository rules grow to hundreds of lines and contradict the shared skills. Recovery: one tech lead is A per repository; move procedures out of rules into skills, and run the harness evals before merging a rules change. See shared agent rules.

A loop runs on a person’s token. When that person leaves, the loop fails or keeps running with their full access. Recovery: issue a per-loop identity, record it in the register, revoke the personal token. The procedure is in agent identity, credentials and secrets.

A model or client change ships without evals. Defaults move under you. On 2026-09-26, Claude Code’s stable release channel (2.1.274) still gave Pro and Team Standard seats Sonnet 5 as the default model, while the latest channel defaulted to Opus 5.5 from v2.1.280. Two teams on different channels can run different models under the same policy. Recovery: the platform owner pins the channel with autoUpdatesChannel ("stable" or "latest") and the models with availableModels plus enforceAvailableModels in managed settings, and re-runs the harness evals before moving them; see the models hub for current defaults.

The register drifts from reality. New scheduled jobs appear without rows. Recovery: the monthly inventory job from the second prompt above, and a standing rule that unregistered loops are paused, not grandfathered.

A broken shared hook stops everyone and no one can roll it back. Recovery: every hook release is versioned with a named rollback owner and a tested rollback; the quarterly rehearsal times it. See shared hooks governance. When an agent does cause an incident, the postmortem names the failed row of this RACI; when an agent causes an incident covers the procedure.

Five questions an executive can ask the CTO

Section titled “Five questions an executive can ask the CTO”
  1. For each loop that runs without a human in the session, who is accountable, and where is it written down?
  2. Which agent identities can write to production or customer data, and how fast can we revoke them?
  3. When we change a model, skill or policy, what evidence tells us it did not get worse?
  4. Who can approve a production change, and can that person be the one who asked the agent for it?
  5. Is our platform group making review faster or adding a queue? Show the review wait time by risk class.

The CTO track reaches this page after the pilot design and continues with the economics of agent-built software, which puts a cost on each accountable row.