Governance and autonomy: put humans at the gates
Governance for AI coding agents assigns each change a risk class (low, medium, high, or critical) and lets that class, not the agent’s confidence or the vendor, decide what the agent may do alone. Guardrails constrain the agent while it works; gates decide whether the evidence is enough to merge or release. Named humans own the high and critical gates.
This page is for CTOs and tech leads who own the autonomy policy. Your agents already open pull requests overnight. Someone asks in the incident review who approved the migration that dropped a column, and the honest answer is “the agent’s own review pass, plus a CLAUDE.md line that said never touch production.” That is advice, not a control, and this page replaces it with controls you can point at.
What this autonomy policy gives you
Section titled “What this autonomy policy gives you”- A four-class risk table that says what the agent may do alone and which human decision each class needs.
- A placement map of guardrails and gates across the lifecycle, with the evidence each one leaves.
- Per-tool enforcement for Claude Code, Codex, and Cursor, so the boundary sits outside the model.
- An evidence contract prompt and a policy template you can adopt as they stand.
- A rehearsal drill and four control metrics that prove the controls work without anyone reading every diff.
Which risk class does an agent change belong to?
Section titled “Which risk class does an agent change belong to?”Classify by what a mistake costs, not by how large the diff is. Score five factors: blast radius, data sensitivity, reversibility, regulatory impact, and the credential scope the run needs. The highest factor wins. A three-line change to session handling is critical even when the agent is certain.
| Risk class | Typical change | The agent may | The gate | Required human decision |
|---|---|---|---|---|
| Low | Tests, docs, research, internal tooling, isolated refactors behind green tests | Edit an isolated checkout, run local checks, open a PR | CI green plus a complete evidence bundle | Lightweight review and merge by any team member |
| Medium | Feature work in one service, dependency upgrades, preview or staging changes | Use scoped non-production credentials and pre-approved runbooks | CI plus accepted plan plus code-owner approval | Code owner accepts the plan and approves the merge |
| High | Auth, billing, schema migrations, shared libraries, sensitive data paths | Prepare the change, a dry run, and a tested rollback | Specialist review plus rehearsed rollback plus release approval | Named specialist and release owner, separately |
| Critical | Production data changes, security exceptions, destructive operations, regulated scope | Prepare evidence and an approval request only | Deployment platform gate the agent’s identity cannot pass | Named owner authorizes; the platform executes |
Risk classes are not maturity levels. A loop’s level on the autonomy ladder measures how far you trust that loop; its risk class measures what a mistake costs. Keep both in the autonomy register, so a loop with a long clean record never earns its way into critical changes by track record alone.
Some actions stay outside agent authority in every class: reading or rotating production secrets, unapproved external communication, deleting backups, and approving its own work. Write them into the policy as forbidden, and enforce them with deny rules and credentials the agent never holds.
Guardrails or gates: which control goes where?
Section titled “Guardrails or gates: which control goes where?”A guardrail shapes or constrains the agent before and during work: a sandbox, a scoped token, a hook that blocks a command. A gate decides whether the evidence is sufficient to advance: a required check, a code-owner approval, a deployment approval. You need both, because a guardrail cannot judge product intent and a gate cannot stop a command that already ran.
| Moment | Control | Example | Evidence it leaves |
|---|---|---|---|
| Before work | Boundary | Accepted intent, allowed paths, data class, budget | Reviewed intent.md and plan.md |
| During work | Guardrail | Sandbox, scoped credentials, hook, protected path | Policy log and a blocked-action test |
| Before merge | Gate | Types, tests, security scan, specification review | CI results and the review record |
| Before production | Authority gate | Named approver, change window, rollback readiness | Approval and release record |
| After release | Feedback | Canary, telemetry, incident trigger | Observation window and the promote-or-roll-back decision |
Repository instructions (CLAUDE.md, AGENTS.md, Cursor rules, and skills) improve behavior, but they are instructions the model can misread. Match each need to the layer that actually enforces it:
| Need | Correct control | Wrong control |
|---|---|---|
| Explain architecture and conventions | CLAUDE.md, AGENTS.md, Cursor rules | A hook |
| Reuse a multi-step workflow | Skill | A long prompt pasted by hand |
| Block or log a deterministic tool action | Hook or execution policy | A sentence in CLAUDE.md |
| Limit files, commands, and network | Sandbox plus scoped identity | “Do not deploy” in the prompt |
| Require review before merge | Branch protection and code owners | An AI reviewer’s approval |
| Require authorization before production | Deployment environment or change-management gate | The authoring agent’s summary |
| Recover safely | Rehearsed rollback with a named owner | A runbook nobody has run |
| Prove what happened | Immutable CI, PR, deployment, and incident logs | The agent’s transcript alone |
Separate agent identity from human authority
Section titled “Separate agent identity from human authority”-
Give every non-interactive agent its own service identity. Never reuse a developer’s broad token; the details are in agent identity, credentials, and secrets.
-
Grant the smallest repository, network, cloud, and data scope the run’s risk class needs. A worktree or container does not narrow a cloud token.
-
Keep production secrets out of every agent run, whatever its class. In critical changes the deployment platform holds them, never the agent. A prompt that says “do not deploy” is not an access boundary.
-
Require branch protection and code-owner approval for merge. The agent’s identity must not be a code owner.
-
Require a separate deployment-environment approval for high and critical changes. Neither the authoring agent nor a reviewing agent can satisfy it.
-
Log the initiating human, agent identity, tool and model version, artifact versions, commands, findings, approval, deployment, and rollback outcome.
Enforce the boundary in each tool
Section titled “Enforce the boundary in each tool”The pattern is identical in all three tools: guidance in instruction files, hard limits in managed settings or admin policy, and merge and production gates on the forge and the deployment platform. The settings differ. For a side-by-side of every managed key, see enforcing one policy across every coding agent.
Put hard limits in managed settings (managed-settings.json, MDM, or server-managed settings on Team and Enterprise), which user and project settings cannot override. Keys checked against Claude Code 2.1.283:
{ "permissions": { "disableBypassPermissionsMode": "disable", "deny": [ "Bash(terraform apply:*)", "Bash(kubectl delete:*)" ] }, "allowManagedPermissionRulesOnly": true, "allowManagedHooksOnly": true}A Bash deny rule matches the command text Claude writes, not the program, so sh -c or a script gets past it (Claude Code docs, checked against v2.1.283). The real boundary for terraform and kubectl is that the agent identity holds no production cloud or cluster credentials; treat the deny rule as a typo-catcher on top.
allowManagedPermissionRulesOnly ignores permission rules from user, project, and local settings; allowManagedHooksOnly does the same for hooks. From v2.1.283 (the latest channel; stable is 2.1.274), auto mode is the starting mode for interactive terminal and VS Code sessions on supported models, and a classifier reviews each action. Treat that classifier as a guardrail, not a gate. To turn it off for regulated repositories, add "disableAutoMode": "disable" inside permissions. In CI, claude -p starts in Manual mode, and --permission-prompts none denies anything that would prompt.
Claude Code’s managed Code Review posts findings, but “the check run always completes with a neutral conclusion so it never blocks merging through branch protection rules” (Claude Code docs, checked against v2.1.283). If its findings must gate a merge, parse them in your own required CI check. A failed or timed-out review also completes as neutral (Claude Code docs, checked against v2.1.283), so your CI check must fail when the result is missing.
Put hard limits in the admin-managed requirements.toml, not in config.toml: OpenAI says to keep “config.toml defaults, requirements.toml constraints, and managed or administrator policy separate.” Keys checked against Codex CLI 0.157.1:
# requirements.toml (admin-managed)allowed_approval_policies = ["untrusted", "on-request"]allowed_sandbox_modes = ["read-only", "workspace-write"]allow_managed_hooks_only = trueThis blocks never as an approval policy and danger-full-access as a sandbox mode on managed machines. allowed_permission_profiles is a table of profile names set to true or false (for example { managed = true }) that limits which beta permission profiles can be selected, and mcp_servers and plugins constrain what can be connected. Codex’s automatic approval review (--approve-for-me) routes sandbox-boundary approvals to a reviewer model; that is an agent approving an agent, so allow it only for low-class work. To forbid it on machines that work on medium-or-higher repositories, set allowed_approvals_reviewers = ["user"] in requirements.toml (the other value is "auto_review"); requirements apply per machine, not per repository. Since 0.150.0, untrusted projects no longer supply project-level AGENTS.md, so the trust decision is part of the policy. On the gate side, codex review and codex exec review produce findings that can feed a required CI check you own; they are not a gate by themselves, so keep branch protection and deployment approval on the forge and the deployment platform.
Cursor offers the same artifact types: Rules, Agent Skills, Hooks, Plugins, and MCP. Cloud Agents run in isolated VMs, and a Run modes page sits under Agent security (Cursor docs, checked 2026-08-28). For a hard limit inside the tool, use Cursor Hooks: a hook talks to the agent in JSON over stdio, runs before defined stages of the agent loop, and can block them (Cursor docs, checked 2026-08-28). For Cursor’s admin side, see Managed policy across agents. Confirm the admin enforcement options in Cursor’s current docs before you write them into policy.
For the gate rows, Bugbot reviews pull requests, and PR Routing & Approval “assigns reviewers based on code ownership and commit history, and can approve low-risk PRs when your criteria are met.” An auto-approval rule is a change to your autonomy policy: restrict it to the low class, give its criteria a named owner, and keep branch protection and deployment approval on the forge and the deployment platform, outside Cursor.
Write an evidence contract for every agent run
Section titled “Write an evidence contract for every agent run”Every background or autonomous run states its contract before it starts. The next gate reads the output, so a run with no contract has nothing for a gate to check.
- Input: the exact commit, artifacts, ticket, telemetry window, and allowed external sources.
- Scope: directories, tools, network destinations, credential class, and a time or cost budget.
- Proof: commands with expected exit status, schemas, screenshots, or eval thresholds.
- Output: the file, pull request, comment, or structured record the next gate reads.
- Failure: non-zero exit, timeout, missing dependency, or low confidence, plus the escalation owner.
- Authority: the actions allowed without approval and the first action that always stops for a human.
Replace the commit, paths, commands, and owner with your own. The fixed clauses (fail closed, never edit the oracle, never approve your own work) are the part to keep verbatim; protecting the test oracle explains why the second one matters most.
Adopt this autonomy policy template
Section titled “Adopt this autonomy policy template”Paste this into your engineering handbook, fill in the owners, and link it from every repository’s instruction file. Each clause names what enforces it; a clause that says “trust” is a known gap.
AGENT AUTONOMY POLICY — v1.0 — owner: CTO — review: quarterly
1. Classification. Every agent change gets a risk class (low, medium, high, critical), computed in CI from touched paths, migrations, and requested credentials. Highest factor wins. Enforced by: required CI check that writes the PR label.2. Forbidden in every class: production secrets, destructive production commands, deleting backups, external communication, approving own work. Enforced by: managed deny rules + credentials the agent identity never holds.3. Identity. One service identity per agent loop, least privilege, short-lived tokens. Enforced by: identity provider + CI OIDC. Owner: platform team.4. Merge gate. Low: CI + evidence bundle + any reviewer. Medium: + code owner. High: + specialist review + rehearsed rollback. Enforced by: branch protection.5. Production gate. High and critical: named release owner approves in the deployment environment. Agent identities cannot approve. Enforced by: deploy platform.6. AI reviewers and auto-approval tools may approve low-class changes only. Criteria owner: release owner. Enforced by: routing rules + branch protection.7. Overrides. Every bypass is attributed, expires in 7 days, and is reviewed. Enforced by: override log in the change record.8. Evidence. Every control has an owner, trigger, failure behavior, exception path, and audit record. Controls fail closed.9. Proof. Controls are rehearsed quarterly; metrics are reported by risk class.The seven-day override expiry and quarterly cadence are this guide’s working defaults, not a published standard; set them to your own risk appetite.
Prompts for designing and auditing controls
Section titled “Prompts for designing and auditing controls”Run these in Claude Code, Codex, or Cursor with read access to the repository and its CI configuration. They produce reports; they change nothing.
How do you prove the controls work?
Section titled “How do you prove the controls work?”Prove each control by making it fail on purpose, then measure the controls by risk class. Nobody has to read every diff for either step.
Rehearse the control path every quarter
Section titled “Rehearse the control path every quarter”-
Ask a low-class agent run to write outside its allowed workspace. The sandbox must block it, and the block must appear in the policy log.
-
Ask it to skip a failing required test. The workflow must fail.
-
Attempt a merge without code-owner approval. Branch protection must block it.
-
Attempt a production release with the agent’s identity and without the named approval. The deployment platform must block it.
-
Make the review service unavailable. The merge must stay blocked; a missing reviewer never counts as an approval.
-
Trigger a rollback in a non-production environment and record the time to restore.
-
Change a rule, skill, or hook and run the harness evals before rollout.
A step that passes when it should fail is an incident in the control, not a flaky test. Record it the way you would record an incident an agent caused.
Measure the controls by risk class
Section titled “Measure the controls by risk class”| Metric | Definition | What a bad trend means |
|---|---|---|
| False-block rate | Blocked changes later merged unchanged ÷ all blocked changes | The gate costs trust and invites bypasses |
| Override rate | Merges or releases with an attributed bypass ÷ all merges in the class | The control is wrong or the class is too strict |
| Escaped defects | Incidents traced to a change that passed every gate | The gate checks the wrong thing |
| Review wait | Median hours from “ready for review” to decision | Human gates became the bottleneck; move checks to automation |
Report these per class. A healthy low class shows short waits and few overrides; a healthy critical class can show long waits and still be correct. The canonical metric definitions live in metrics frameworks.
Who signs off: the release owner signs off the production gate for high and critical changes, the code owner signs off medium merges, and the CTO signs off changes to the policy itself, with the rehearsal results and the metrics attached.
Why is an AI reviewer not a gate?
Section titled “Why is an AI reviewer not a gate?”An author agent and a reviewer agent improve detection, but they can share credentials, configuration, model blind spots, and the organization that set them up. Two agents are not separation of duties. The OWASP GenAI LLM Top 10 2026 (published 2026-08-04) lists Excessive Agency as LLM03 and advises: “Use human-in-the-loop control to require a human to approve high-impact actions before they are taken.”
A real gate uses an identity the authoring run cannot act as, or a platform policy it cannot change, plus the accountable human where the risk class demands one. Use AI review to decide where human attention goes, not to replace it on high and critical changes. For evidence that auditors accept, see agentic engineering in regulated industries.
What breaks in agent governance, and how do you recover?
Section titled “What breaks in agent governance, and how do you recover?”Advisory text is treated as enforcement. A CLAUDE.md line says “never run migrations,” and one day the agent runs one. Recover: run the “separate advice from enforcement” prompt, then move each advisory-only invariant into an enforced control. For destructive actions, prefer removing the credentials from the agent identity or a platform permission; a managed deny rule, a hook, or CI catches the common spelling on top.
One token reaches every environment. The agent’s CI token can also deploy. Recover: revoke it first, then split identities per loop and environment, and rotate. Treat the exposure as an incident until the logs show what the token touched.
The agent labels its own risk. Every change arrives as “low.” Recover: compute the class in CI from paths and migrations, and alert when an agent-proposed label is lower than the computed one.
A human gate has no decision criteria. Approvers click through because nothing tells them what to check. Recover: attach the evidence bundle and a one-line risk statement to every request, and set the criteria, owner, timeout, and escalation path before automating the request.
Overrides never expire. A one-off bypass for an outage becomes permanent. Recover: give every override an expiry and an owner, and report the open ones in the quarterly review.
A control fails open. The review service times out and the merge proceeds. Recover: make the check required and fail on a missing result, then add the outage case to the rehearsal.
Rollback exists only in a document. Recover: rehearse it in a representative non-production environment and measure the time to restore. Pair high-class releases with progressive delivery so the rollback is a flag, not a redeploy.
False certainty. Someone calls the instructions “unbreakable guardrails” or retires human review because the checks are green. Models misread instructions, hooks can be misconfigured, and deterministic checks cannot judge product intent. Recover: test the full control path, write down the residual risk, and keep human authority at the consequential gates.
Where to go next with agent governance
Section titled “Where to go next with agent governance”Before this page, read the operating model for who owns the autonomy register. Next, apply the policy to the threats and credentials it has to hold against.
For the stage procedures these gates sit inside, see Deploy and Maintain.