The agent platform team: running the harness as a product
An agent platform team runs the shared harness for coding agents as an internal product: managed policy, shared skills and hooks, the MCP catalogue, harness evals, agent environments and telemetry. It publishes service levels to product teams, ships every harness change through an eval gate, and stays out of product teams’ pull-request approval path.
You approved a platform group three months ago. It has built a plugin nobody installed, a policy file that blocks two teams’ MCP servers without warning, and a request queue that now takes eleven days to answer. Product teams now copy skills into their own repositories to get around it. The harness exists, but it is run like a ticket desk, not a product.
This page is for the CTO who funds the team and the tech lead who either runs it or depends on it. Who owns what, and when a dedicated team is justified, is in the operating model; this page is about running the team once it exists.
What an agent platform team gives your organization
Section titled “What an agent platform team gives your organization”- A one-page charter: mission, scope, what the team refuses to do, service levels.
- A prioritized backlog of nine harness products, each with a “done when” test.
- That backlog mapped to Claude Code, Codex and Cursor config.
- A CI eval gate for harness releases, with a decision rule and a named sign-off.
- A service-level table and five product metrics.
- Five rules that keep the platform team from becoming a gate.
Why run the harness as a product, not a project?
Section titled “Why run the harness as a product, not a project?”Harness artifacts are shared state. A skill, hook, policy key or MCP server changes every agent session that loads it, and keeps changing as models, clients and codebases move. A project that “rolls out the harness” and disbands leaves an unowned artifact that decays.
The evidence for investing here is correlational but consistent. DORA’s 2025 report found that “90% of organizations have adopted at least one platform and there is a direct correlation between a high quality internal platform and an organization’s ability to unlock the value of AI” (Google Cloud, 2025-09-23). “Quality internal platforms” is also one of the seven capabilities in DORA’s AI Capabilities Model (Google Cloud, 2025-09-23). Stripe shows the shape at scale: its Toolshed “currently contains nearly 500 MCP tools for internal systems and SaaS platforms we use at Stripe” (Stripe engineering blog, 2026-02-19), one catalogue that every agent run draws from.
“Product” means three concrete things here: product teams can decline most of what you ship, every release is versioned and can be rolled back, and success is measured by what changes for those teams.
What goes in the agent platform team’s charter?
Section titled “What goes in the agent platform team’s charter?”Copy this into your engineering handbook and fill in the angle-bracket fields. The “we do not” list matters most: product teams will quote it back when the platform team drifts into their path.
# Agent platform team charter: <organization>, v<N>, <date>
## MissionMake it cheap and safe for every product team to ship with coding agents,and prove it with evidence rather than review effort.
## CustomersProduct teams using Claude Code, Codex or Cursor. Security and the releaseowner are partners, not customers.
## We own (the harness products)1. Policy floor: managed settings and requirements for every supported tool.2. Paved-road plugin: shared skills, hooks and agents, one marketplace.3. MCP catalogue: vetted servers, scoped credentials, a request path.4. Harness evals: the suite every harness change must pass.5. Agent environments: dev containers and cloud or self-hosted sandboxes.6. Telemetry: usage, cost and outcome data in our own collector.7. Version channels: which client versions and models teams run.8. Onboarding: the first-week path to a first accepted agent change.(Agent identities are owned with security; see the operating model.)
## We do not- Approve, block or review product teams' pull requests.- Decide what a product's tests or acceptance criteria check.- Make any paved-road component mandatory unless it is in the policy floor.- Ship a harness change that has not passed the harness eval suite.
## Policy floor (the only mandatory part)<list the controls security requires: allowed models, MCP allowlist,managed hooks for secrets scanning, telemetry export, permission limits>
## Service levelsSee the service-level table; published on <URL>, reported monthly.
## ContributionAny engineer can propose a skill, hook or MCP server by pull request to<marketplace repo> with eval cases. We review within <N> working days.
## ExceptionsA team may opt out of any paved-road component by recording it in<exceptions file> with an owner and an expiry date (max 90 days).
## ReviewQuarterly with the CTO and two product-team tech leads. Signed: <CTO>.What belongs on the platform backlog?
Section titled “What belongs on the platform backlog?”These are the nine harness products, in the order most organizations need them. The “done when” column is each one’s acceptance test: without that evidence, the item is not done.
| # | Harness product | What the team ships | Done when |
|---|---|---|---|
| 1 | Policy floor | Managed settings (Claude Code), requirements.toml (Codex), team admin settings (Cursor) with the controls security requires | A test machine with a developer’s own config still gets the floor; an unlisted MCP server is refused |
| 2 | Telemetry and cost | OpenTelemetry export from every client to your collector; a dashboard per team | Cost and sessions per team are visible within a day, without anyone asking developers |
| 3 | Version channels | Pinned client channel and allowed models, a promotion process between them | A new model or client reaches everyone only after the eval suite runs on it |
| 4 | Harness evals | Golden tasks from your own repositories, graders, a CI job | Every harness pull request shows a score against the current release |
| 5 | Paved-road plugin | One marketplace with shared skills, hooks and subagents, versioned releases | Two or more teams adopt it without being required to, and its context cost is published |
| 6 | MCP catalogue | Vetted servers with scoped, per-agent credentials and a request form | A new server request is answered inside the service level, with security’s decision recorded |
| 7 | Agent environments | Dev container images, shared cloud or self-hosted environments with network allowlists | An agent starts in a clean environment that builds and tests the repository with no manual setup |
| 8 | Onboarding | A first-week path, starter tasks, the prompts that work in your codebase | A new engineer lands a first accepted agent change within the target time |
| 9 | Contribution path | Templates, eval-case examples, review rota | Most new skills come from product teams, not from the platform team |
Agent identities and credentials are not on this list on purpose. The platform team does the work, but security is accountable for them; the procedure is in agent identity, credentials and secrets.
In what order should a new platform team build it?
Section titled “In what order should a new platform team build it?”-
Weeks 1–3: the floor and the telemetry. Ship them together: without telemetry you cannot tell whether anything later is used, and without the floor security will not let you widen anything.
-
Weeks 3–6: version channels and the eval suite. Collect 20–40 golden tasks from real merged work in two or three repositories, with a grader for each (usually the repository’s own tests). Run them on the current client and model for a baseline. The range is this guide’s working number, not a published benchmark.
-
Weeks 6–10: the paved-road plugin. Start with the two or three skills teams already copy between repositories; an existing habit is adopted faster than a new idea.
-
Weeks 8–12: MCP catalogue and environments. Vet first the servers teams already run on personal tokens; each is a credential risk today.
-
From week 12: onboarding and the contribution path. Publish the first service-level report and invite contributions.
How the backlog maps to Claude Code, Codex and Cursor
Section titled “How the backlog maps to Claude Code, Codex and Cursor”The backlog is tool-neutral; the config is not. Each tool’s team/ section covers its own rollout in detail: Claude Code for teams, Codex for teams and Cursor for teams. The side-by-side policy comparison is in enforcing one policy across every coding agent.
Policy floor, channels and telemetry live in managed settings, delivered as managed-settings.json, through MDM, or as server-managed settings from the claude.ai console (Team and Enterprise). Managed settings rank above every other level.
{ "autoUpdatesChannel": "stable", "availableModels": ["opus", "sonnet"], "enforceAvailableModels": true, "allowManagedHooksOnly": true, "allowManagedMcpServersOnly": true, "extraKnownMarketplaces": { "acme-plugins": { "source": { "source": "github", "repo": "acme/acme-plugins" } } }, "enabledPlugins": { "acme-harness@acme-plugins": true }, "env": { "CLAUDE_CODE_ENABLE_TELEMETRY": "1", "OTEL_METRICS_EXPORTER": "otlp", "OTEL_LOGS_EXPORTER": "otlp", "OTEL_EXPORTER_OTLP_PROTOCOL": "grpc", "OTEL_EXPORTER_OTLP_ENDPOINT": "http://otel-collector.acme.internal:4317" }}acme/acme-plugins is the platform team’s marketplace repository and acme-harness its plugin. Put the telemetry variables in managed or user settings. From v2.1.282 (the latest channel) Claude Code ignores telemetry export variables in a project’s .claude/settings.json env block; on stable 2.1.274 a project file can still set them, so managed settings is the only level a repository cannot override on both channels. The metrics include claude_code.cost.usage, claude_code.token.usage and claude_code.session.count.
Channels: pin stable (2.1.274 on 2026-09-26; latest was 2.1.283) and promote to latest only for a canary group.
Environments: self-hosted environments (a Team and Enterprise beta, claude --environment ccpool_…) read server-managed settings and, when it is one of the managed sources Claude Code applies, the managed settings file in the runner image. On an Anthropic-managed cloud VM only server-managed settings apply, and repository-enabled plugins are not installed. See also the dev container setup and the LLM gateway guide.
Context cost: claude plugin details acme-harness prints the plugin’s component inventory and projected token cost. Publish that number with every release.
Policy floor lives in requirements.toml, which holds the constraints an administrator enforces; config.toml holds defaults. Keys from ConfigRequirementsToml in config_requirements.rs (openai/codex rust-v0.157.1, checked 2026-09-26): allowed_approval_policies, allowed_sandbox_modes, allowed_permission_profiles, default_permissions, allow_managed_hooks_only, hooks (managed hooks), mcp_servers, plugins, marketplaces and enforce_residency. allow_managed_hooks_only has no effect in config.toml; it is a requirements.toml key.
Paved-road plugin: teams add the platform marketplace pinned to a release tag.
# Terminal, on a developer machine or in an environment imagecodex plugin marketplace add acme/acme-plugins --ref v1.4.0Telemetry: the [otel] table in config.toml exports logs, traces and metrics (field names from OtelConfigToml in the 0.157.1 source).
[otel]environment = "prod"log_user_prompt = falseexporter = { otlp-http = { endpoint = "https://otel.acme.internal/v1/logs", protocol = "binary" } }Keep log_user_prompt = false unless privacy review approves; prompts carry source code and customer data.
Environments: Codex Cloud runs tasks remotely; codex cloud (experimental in 0.157.1) browses those tasks and applies their changes locally. Governance and admin setup is in Codex enterprise governance.
Paved-road plugin: Cursor Plugins “package rules, skills, agents, commands, MCP servers, and hooks” (Cursor docs, checked 2026-08-28), so the same marketplace content can ship as a Cursor plugin next to the Claude Code and Codex ones. Skills follow the Agent Skills standard, which keeps one SKILL.md usable across all three tools.
Environments: Cloud Agents “run in isolated VMs in the cloud with full development environments”, and Builds “prepare your Cloud Agent environment in the background” (Cursor docs, checked 2026-08-28). Version the Cloud Agent environment definition like a container image.
Evals: the Cursor SDK (@cursor/sdk 1.0.32 on npm, 2026-09-22) calls Cursor’s agent from code, so the same golden-task runner can drive Cursor.
Policy floor: check Cursor’s team admin controls against its live admin docs before you rely on them; start from Cursor privacy and security.
How do you prove a harness release does not make agents worse?
Section titled “How do you prove a harness release does not make agents worse?”Every harness change (a skill edit, hook, policy key, client version or model) changes every agent’s behaviour, so each goes through the eval suite before release. The platform team signs off on the eval result, not a diff read.
The suite has three layers:
- Golden tasks from your own merged work, each with a deterministic grader: the repository’s tests, a type check, a lint gate or a fitness function. These decide the release.
- Model-graded checks for what tests cannot see, such as whether a review skill found a planted defect. These inform the release, and a human samples them.
- Canary telemetry after release: cost per session and rollback requests from the canary group.
claude plugin eval runs a plugin’s eval cases (evals/**/case.yaml, or prompt.md plus graders/*.md) and, when the plugin resolves, adds a no-plugin baseline arm so you see the delta the plugin makes. In CI:
# CI job on every pull request to the marketplace repositoryclaude plugin eval ./plugins/acme-harness \ --trust-plugin --runs 3 --threshold 0.8 \ --max-cost-usd 25 --no-publish --json eval-results.json--threshold exits 1 if any case scores below it; --max-cost-usd aborts with partial results (exit 2) at the ceiling. --trust-plugin runs the plugin’s code as the job’s user (the help text compares it to --dangerously-skip-permissions), so run it only for branches inside the marketplace repository, never forks, on a disposable runner whose only secret is a spend-capped eval key. The LLM grader defaults to Haiku; override it with --judge-model. All flags checked against Claude Code 2.1.283.
Codex has no plugin eval command in 0.157.1, so drive each golden task through codex exec and grade the result with the repository’s own checks:
# CI job: one golden task, in a fresh checkout of the task's base commitcodex -a never exec --ephemeral \ -c default_permissions=":workspace" \ --output-schema evals/schema/result.json \ -o out/flaky-retry.json \ "$(cat evals/cases/flaky-retry/prompt.md)"npm test -- --run tests/retry.test.ts # the grader--output-schema makes the final message machine-readable; --ephemeral runs without persisting session files to disk. Compare pass rates for the current and candidate harness.
Drive the same golden-task set through the Cursor SDK (@cursor/sdk) and grade with the repository’s tests. Keep cases and graders in the marketplace repository so one suite scores all three tools.
Decision rule. Write it into the charter before the first release: a harness release ships when golden-task pass rate is not lower than the current release on any repository in the suite, and cost per passing task has not risen by more than an agreed margin. The platform team lead signs off; security co-signs anything that touches the policy floor. Methods for building the suite are in evals for coding agents and model-graded checks.
What service levels should the platform team promise product teams?
Section titled “What service levels should the platform team promise product teams?”Service levels turn “the platform team is slow” into a number. Publish them, report monthly, and let product teams escalate a breach to the CTO. The targets are this guide’s starting values, not an industry benchmark; reset them from your first month of data.
| Service | Measured as | Starting target |
|---|---|---|
| Harness request (new skill, MCP server, policy change) | Time from request to a decision: yes, no, or a date | 3 working days |
| Contribution review | Time from a product team’s pull request to the marketplace to merge or actionable feedback | 2 working days |
| Broken shared hook or skill | Time from first report to rollback on every machine | 4 hours |
| Agent identity revocation (with security) | Time from incident call to credential revoked | 15 minutes |
| New client or model version | Time from vendor release to “evaluated, promoted or held” decision | 5 working days |
| Environment image | Share of agent sessions that start in a clean, building environment | 95% |
Rehearse the rollback and revocation rows: once a quarter, break a hook in a canary branch and time the fix reaching every developer; shared hooks governance covers the release mechanics.
How do you keep the platform team from becoming a gate?
Section titled “How do you keep the platform team from becoming a gate?”A platform team becomes a gate the moment product teams wait on it to ship product work. These five checkable rules prevent that.
-
Separate the floor from the paved road. Only the policy floor is mandatory, and security owns its contents. Everything else is an offer; if you have to mandate a skill, it is not good enough yet.
-
Never sit in the product review path. The platform team is absent from product repositories’
CODEOWNERSand from branch protection. It reviews harness changes only. Review of product changes is covered in running the review queue. -
Accept contributions with eval cases, not permission requests. A team that wants a new skill opens a pull request to the marketplace with the skill and two or three eval cases, and the platform team reviews against the eval result within the service level.
-
Allow exceptions with an expiry. A team may replace a paved-road skill or hook with its own by recording an owner and an expiry date. Every exception is a product gap on the platform backlog.
-
Publish the queue. Request age, contribution review time and open exceptions sit next to the service levels. A visible queue gets fixed.
How do you measure an agent platform team as a product?
Section titled “How do you measure an agent platform team as a product?”Measure what changes for product teams, not what the platform team produces: skills shipped is an output, these five are outcomes. Canonical metric definitions for the wider organization are in metrics frameworks.
| Metric | Definition | Why it matters |
|---|---|---|
| Voluntary adoption | Share of active agent sessions that load the paved-road plugin where it is not force-enabled | The only honest signal that the paved road is better than the alternatives |
| Time to first accepted agent change | Days from a new engineer’s first session to their first merged, agent-authored change that passed all gates | Measures onboarding and environments together |
| Fork rate | Number of shared skills or hooks copied into product repositories and modified | Each fork is a missing feature or a service-level miss |
| Harness-caused incidents | Incidents whose postmortem names a shared skill, hook, policy or environment | Quality of the platform’s own releases |
| Cost per accepted change | Agent spend divided by merged agent-authored changes that passed all gates, per team | Links telemetry to outcomes; see cost governance |
Report per team, never per person: telemetry that ranks individuals becomes a performance target, then surveillance; career ladders and performance reviews covers how to keep it out of reviews.
Copy-paste prompts for the agent platform team
Section titled “Copy-paste prompts for the agent platform team”Run these in Claude Code, Codex or Cursor. They behave the same way in all three tools.
What breaks when you run an agent platform team, and how to recover
Section titled “What breaks when you run an agent platform team, and how to recover”The platform team becomes the review gate. Product pull requests wait on platform approval because “they own the agents”. Recovery: remove the platform team from product CODEOWNERS, restate the “we do not” list, and turn the check they were doing by hand into a hook or eval case.
Nobody adopts the paved road. The team built what it guessed teams needed. Recovery: run the first prompt above, rebuild the plugin around the three most-copied skills, and measure voluntary adoption for a month before building more.
A policy-floor change breaks teams without warning. An allowlist update refuses an MCP server two teams depend on. Recovery: roll back through the same channel, then add a canary group that receives floor changes a week early, and announce floor changes with a date.
The plugin eats the context window. Each team’s favourite skill joined the shared plugin, and every session now starts with thousands of tokens of descriptions. Recovery: publish the token cost per release (claude plugin details on Claude Code), set a budget, and move rarely used skills to an optional plugin.
A model or client update ships before the evals run. A new default arrives through auto-update, and agent behaviour changes across the organization overnight. Recovery: pin the channel and allowed models in managed policy, and make the version-channel service level the only route to promotion. Current defaults are on the models hub.
Telemetry turns into a leaderboard. Someone exports per-person token counts into a performance review. Recovery: aggregate to team level in the collector, restrict raw data to the platform team and security, and put that rule in the charter.
Where to go next with the agent platform team
Section titled “Where to go next with the agent platform team”In the CTO track this page follows managed policy, which defines the policy floor, and leads to cost governance, which puts budgets on the telemetry you now collect.