Skip to content

The agent platform team: running the harness as a product

An agent platform team runs the shared harness for coding agents as an internal product: managed policy, shared skills and hooks, the MCP catalogue, harness evals, agent environments and telemetry. It publishes service levels to product teams, ships every harness change through an eval gate, and stays out of product teams’ pull-request approval path.

You approved a platform group three months ago. It has built a plugin nobody installed, a policy file that blocks two teams’ MCP servers without warning, and a request queue that now takes eleven days to answer. Product teams now copy skills into their own repositories to get around it. The harness exists, but it is run like a ticket desk, not a product.

This page is for the CTO who funds the team and the tech lead who either runs it or depends on it. Who owns what, and when a dedicated team is justified, is in the operating model; this page is about running the team once it exists.

What an agent platform team gives your organization

Section titled “What an agent platform team gives your organization”
  • A one-page charter: mission, scope, what the team refuses to do, service levels.
  • A prioritized backlog of nine harness products, each with a “done when” test.
  • That backlog mapped to Claude Code, Codex and Cursor config.
  • A CI eval gate for harness releases, with a decision rule and a named sign-off.
  • A service-level table and five product metrics.
  • Five rules that keep the platform team from becoming a gate.

Why run the harness as a product, not a project?

Section titled “Why run the harness as a product, not a project?”

Harness artifacts are shared state. A skill, hook, policy key or MCP server changes every agent session that loads it, and keeps changing as models, clients and codebases move. A project that “rolls out the harness” and disbands leaves an unowned artifact that decays.

The evidence for investing here is correlational but consistent. DORA’s 2025 report found that “90% of organizations have adopted at least one platform and there is a direct correlation between a high quality internal platform and an organization’s ability to unlock the value of AI” (Google Cloud, 2025-09-23). “Quality internal platforms” is also one of the seven capabilities in DORA’s AI Capabilities Model (Google Cloud, 2025-09-23). Stripe shows the shape at scale: its Toolshed “currently contains nearly 500 MCP tools for internal systems and SaaS platforms we use at Stripe” (Stripe engineering blog, 2026-02-19), one catalogue that every agent run draws from.

“Product” means three concrete things here: product teams can decline most of what you ship, every release is versioned and can be rolled back, and success is measured by what changes for those teams.

What goes in the agent platform team’s charter?

Section titled “What goes in the agent platform team’s charter?”

Copy this into your engineering handbook and fill in the angle-bracket fields. The “we do not” list matters most: product teams will quote it back when the platform team drifts into their path.

# Agent platform team charter: <organization>, v<N>, <date>
## Mission
Make it cheap and safe for every product team to ship with coding agents,
and prove it with evidence rather than review effort.
## Customers
Product teams using Claude Code, Codex or Cursor. Security and the release
owner are partners, not customers.
## We own (the harness products)
1. Policy floor: managed settings and requirements for every supported tool.
2. Paved-road plugin: shared skills, hooks and agents, one marketplace.
3. MCP catalogue: vetted servers, scoped credentials, a request path.
4. Harness evals: the suite every harness change must pass.
5. Agent environments: dev containers and cloud or self-hosted sandboxes.
6. Telemetry: usage, cost and outcome data in our own collector.
7. Version channels: which client versions and models teams run.
8. Onboarding: the first-week path to a first accepted agent change.
(Agent identities are owned with security; see the operating model.)
## We do not
- Approve, block or review product teams' pull requests.
- Decide what a product's tests or acceptance criteria check.
- Make any paved-road component mandatory unless it is in the policy floor.
- Ship a harness change that has not passed the harness eval suite.
## Policy floor (the only mandatory part)
<list the controls security requires: allowed models, MCP allowlist,
managed hooks for secrets scanning, telemetry export, permission limits>
## Service levels
See the service-level table; published on <URL>, reported monthly.
## Contribution
Any engineer can propose a skill, hook or MCP server by pull request to
<marketplace repo> with eval cases. We review within <N> working days.
## Exceptions
A team may opt out of any paved-road component by recording it in
<exceptions file> with an owner and an expiry date (max 90 days).
## Review
Quarterly with the CTO and two product-team tech leads. Signed: <CTO>.

These are the nine harness products, in the order most organizations need them. The “done when” column is each one’s acceptance test: without that evidence, the item is not done.

#Harness productWhat the team shipsDone when
1Policy floorManaged settings (Claude Code), requirements.toml (Codex), team admin settings (Cursor) with the controls security requiresA test machine with a developer’s own config still gets the floor; an unlisted MCP server is refused
2Telemetry and costOpenTelemetry export from every client to your collector; a dashboard per teamCost and sessions per team are visible within a day, without anyone asking developers
3Version channelsPinned client channel and allowed models, a promotion process between themA new model or client reaches everyone only after the eval suite runs on it
4Harness evalsGolden tasks from your own repositories, graders, a CI jobEvery harness pull request shows a score against the current release
5Paved-road pluginOne marketplace with shared skills, hooks and subagents, versioned releasesTwo or more teams adopt it without being required to, and its context cost is published
6MCP catalogueVetted servers with scoped, per-agent credentials and a request formA new server request is answered inside the service level, with security’s decision recorded
7Agent environmentsDev container images, shared cloud or self-hosted environments with network allowlistsAn agent starts in a clean environment that builds and tests the repository with no manual setup
8OnboardingA first-week path, starter tasks, the prompts that work in your codebaseA new engineer lands a first accepted agent change within the target time
9Contribution pathTemplates, eval-case examples, review rotaMost new skills come from product teams, not from the platform team

Agent identities and credentials are not on this list on purpose. The platform team does the work, but security is accountable for them; the procedure is in agent identity, credentials and secrets.

In what order should a new platform team build it?

Section titled “In what order should a new platform team build it?”
  1. Weeks 1–3: the floor and the telemetry. Ship them together: without telemetry you cannot tell whether anything later is used, and without the floor security will not let you widen anything.

  2. Weeks 3–6: version channels and the eval suite. Collect 20–40 golden tasks from real merged work in two or three repositories, with a grader for each (usually the repository’s own tests). Run them on the current client and model for a baseline. The range is this guide’s working number, not a published benchmark.

  3. Weeks 6–10: the paved-road plugin. Start with the two or three skills teams already copy between repositories; an existing habit is adopted faster than a new idea.

  4. Weeks 8–12: MCP catalogue and environments. Vet first the servers teams already run on personal tokens; each is a credential risk today.

  5. From week 12: onboarding and the contribution path. Publish the first service-level report and invite contributions.

How the backlog maps to Claude Code, Codex and Cursor

Section titled “How the backlog maps to Claude Code, Codex and Cursor”

The backlog is tool-neutral; the config is not. Each tool’s team/ section covers its own rollout in detail: Claude Code for teams, Codex for teams and Cursor for teams. The side-by-side policy comparison is in enforcing one policy across every coding agent.

Policy floor, channels and telemetry live in managed settings, delivered as managed-settings.json, through MDM, or as server-managed settings from the claude.ai console (Team and Enterprise). Managed settings rank above every other level.

{
"autoUpdatesChannel": "stable",
"availableModels": ["opus", "sonnet"],
"enforceAvailableModels": true,
"allowManagedHooksOnly": true,
"allowManagedMcpServersOnly": true,
"extraKnownMarketplaces": {
"acme-plugins": { "source": { "source": "github", "repo": "acme/acme-plugins" } }
},
"enabledPlugins": { "acme-harness@acme-plugins": true },
"env": {
"CLAUDE_CODE_ENABLE_TELEMETRY": "1",
"OTEL_METRICS_EXPORTER": "otlp",
"OTEL_LOGS_EXPORTER": "otlp",
"OTEL_EXPORTER_OTLP_PROTOCOL": "grpc",
"OTEL_EXPORTER_OTLP_ENDPOINT": "http://otel-collector.acme.internal:4317"
}
}

acme/acme-plugins is the platform team’s marketplace repository and acme-harness its plugin. Put the telemetry variables in managed or user settings. From v2.1.282 (the latest channel) Claude Code ignores telemetry export variables in a project’s .claude/settings.json env block; on stable 2.1.274 a project file can still set them, so managed settings is the only level a repository cannot override on both channels. The metrics include claude_code.cost.usage, claude_code.token.usage and claude_code.session.count.

Channels: pin stable (2.1.274 on 2026-09-26; latest was 2.1.283) and promote to latest only for a canary group.

Environments: self-hosted environments (a Team and Enterprise beta, claude --environment ccpool_…) read server-managed settings and, when it is one of the managed sources Claude Code applies, the managed settings file in the runner image. On an Anthropic-managed cloud VM only server-managed settings apply, and repository-enabled plugins are not installed. See also the dev container setup and the LLM gateway guide.

Context cost: claude plugin details acme-harness prints the plugin’s component inventory and projected token cost. Publish that number with every release.

How do you prove a harness release does not make agents worse?

Section titled “How do you prove a harness release does not make agents worse?”

Every harness change (a skill edit, hook, policy key, client version or model) changes every agent’s behaviour, so each goes through the eval suite before release. The platform team signs off on the eval result, not a diff read.

The suite has three layers:

  • Golden tasks from your own merged work, each with a deterministic grader: the repository’s tests, a type check, a lint gate or a fitness function. These decide the release.
  • Model-graded checks for what tests cannot see, such as whether a review skill found a planted defect. These inform the release, and a human samples them.
  • Canary telemetry after release: cost per session and rollback requests from the canary group.

claude plugin eval runs a plugin’s eval cases (evals/**/case.yaml, or prompt.md plus graders/*.md) and, when the plugin resolves, adds a no-plugin baseline arm so you see the delta the plugin makes. In CI:

Terminal window
# CI job on every pull request to the marketplace repository
claude plugin eval ./plugins/acme-harness \
--trust-plugin --runs 3 --threshold 0.8 \
--max-cost-usd 25 --no-publish --json eval-results.json

--threshold exits 1 if any case scores below it; --max-cost-usd aborts with partial results (exit 2) at the ceiling. --trust-plugin runs the plugin’s code as the job’s user (the help text compares it to --dangerously-skip-permissions), so run it only for branches inside the marketplace repository, never forks, on a disposable runner whose only secret is a spend-capped eval key. The LLM grader defaults to Haiku; override it with --judge-model. All flags checked against Claude Code 2.1.283.

Decision rule. Write it into the charter before the first release: a harness release ships when golden-task pass rate is not lower than the current release on any repository in the suite, and cost per passing task has not risen by more than an agreed margin. The platform team lead signs off; security co-signs anything that touches the policy floor. Methods for building the suite are in evals for coding agents and model-graded checks.

What service levels should the platform team promise product teams?

Section titled “What service levels should the platform team promise product teams?”

Service levels turn “the platform team is slow” into a number. Publish them, report monthly, and let product teams escalate a breach to the CTO. The targets are this guide’s starting values, not an industry benchmark; reset them from your first month of data.

ServiceMeasured asStarting target
Harness request (new skill, MCP server, policy change)Time from request to a decision: yes, no, or a date3 working days
Contribution reviewTime from a product team’s pull request to the marketplace to merge or actionable feedback2 working days
Broken shared hook or skillTime from first report to rollback on every machine4 hours
Agent identity revocation (with security)Time from incident call to credential revoked15 minutes
New client or model versionTime from vendor release to “evaluated, promoted or held” decision5 working days
Environment imageShare of agent sessions that start in a clean, building environment95%

Rehearse the rollback and revocation rows: once a quarter, break a hook in a canary branch and time the fix reaching every developer; shared hooks governance covers the release mechanics.

How do you keep the platform team from becoming a gate?

Section titled “How do you keep the platform team from becoming a gate?”

A platform team becomes a gate the moment product teams wait on it to ship product work. These five checkable rules prevent that.

  1. Separate the floor from the paved road. Only the policy floor is mandatory, and security owns its contents. Everything else is an offer; if you have to mandate a skill, it is not good enough yet.

  2. Never sit in the product review path. The platform team is absent from product repositories’ CODEOWNERS and from branch protection. It reviews harness changes only. Review of product changes is covered in running the review queue.

  3. Accept contributions with eval cases, not permission requests. A team that wants a new skill opens a pull request to the marketplace with the skill and two or three eval cases, and the platform team reviews against the eval result within the service level.

  4. Allow exceptions with an expiry. A team may replace a paved-road skill or hook with its own by recording an owner and an expiry date. Every exception is a product gap on the platform backlog.

  5. Publish the queue. Request age, contribution review time and open exceptions sit next to the service levels. A visible queue gets fixed.

How do you measure an agent platform team as a product?

Section titled “How do you measure an agent platform team as a product?”

Measure what changes for product teams, not what the platform team produces: skills shipped is an output, these five are outcomes. Canonical metric definitions for the wider organization are in metrics frameworks.

MetricDefinitionWhy it matters
Voluntary adoptionShare of active agent sessions that load the paved-road plugin where it is not force-enabledThe only honest signal that the paved road is better than the alternatives
Time to first accepted agent changeDays from a new engineer’s first session to their first merged, agent-authored change that passed all gatesMeasures onboarding and environments together
Fork rateNumber of shared skills or hooks copied into product repositories and modifiedEach fork is a missing feature or a service-level miss
Harness-caused incidentsIncidents whose postmortem names a shared skill, hook, policy or environmentQuality of the platform’s own releases
Cost per accepted changeAgent spend divided by merged agent-authored changes that passed all gates, per teamLinks telemetry to outcomes; see cost governance

Report per team, never per person: telemetry that ranks individuals becomes a performance target, then surveillance; career ladders and performance reviews covers how to keep it out of reviews.

Copy-paste prompts for the agent platform team

Section titled “Copy-paste prompts for the agent platform team”

Run these in Claude Code, Codex or Cursor. They behave the same way in all three tools.

What breaks when you run an agent platform team, and how to recover

Section titled “What breaks when you run an agent platform team, and how to recover”

The platform team becomes the review gate. Product pull requests wait on platform approval because “they own the agents”. Recovery: remove the platform team from product CODEOWNERS, restate the “we do not” list, and turn the check they were doing by hand into a hook or eval case.

Nobody adopts the paved road. The team built what it guessed teams needed. Recovery: run the first prompt above, rebuild the plugin around the three most-copied skills, and measure voluntary adoption for a month before building more.

A policy-floor change breaks teams without warning. An allowlist update refuses an MCP server two teams depend on. Recovery: roll back through the same channel, then add a canary group that receives floor changes a week early, and announce floor changes with a date.

The plugin eats the context window. Each team’s favourite skill joined the shared plugin, and every session now starts with thousands of tokens of descriptions. Recovery: publish the token cost per release (claude plugin details on Claude Code), set a budget, and move rarely used skills to an optional plugin.

A model or client update ships before the evals run. A new default arrives through auto-update, and agent behaviour changes across the organization overnight. Recovery: pin the channel and allowed models in managed policy, and make the version-channel service level the only route to promotion. Current defaults are on the models hub.

Telemetry turns into a leaderboard. Someone exports per-person token counts into a performance review. Recovery: aggregate to team level in the collector, restrict raw data to the platform team and security, and put that rule in the charter.

Where to go next with the agent platform team

Section titled “Where to go next with the agent platform team”

In the CTO track this page follows managed policy, which defines the policy floor, and leads to cost governance, which puts budgets on the telemetry you now collect.

Edit page

Last updated:

Cite this page — https://developertoolkit.ai/en/org/platform-team/, developertoolkit.ai