Skip to content

Governance and autonomy: put humans at the gates

Governance for AI coding agents assigns each change a risk class (low, medium, high, or critical) and lets that class, not the agent’s confidence or the vendor, decide what the agent may do alone. Guardrails constrain the agent while it works; gates decide whether the evidence is enough to merge or release. Named humans own the high and critical gates.

This page is for CTOs and tech leads who own the autonomy policy. Your agents already open pull requests overnight. Someone asks in the incident review who approved the migration that dropped a column, and the honest answer is “the agent’s own review pass, plus a CLAUDE.md line that said never touch production.” That is advice, not a control, and this page replaces it with controls you can point at.

  • A four-class risk table that says what the agent may do alone and which human decision each class needs.
  • A placement map of guardrails and gates across the lifecycle, with the evidence each one leaves.
  • Per-tool enforcement for Claude Code, Codex, and Cursor, so the boundary sits outside the model.
  • An evidence contract prompt and a policy template you can adopt as they stand.
  • A rehearsal drill and four control metrics that prove the controls work without anyone reading every diff.

Which risk class does an agent change belong to?

Section titled “Which risk class does an agent change belong to?”

Classify by what a mistake costs, not by how large the diff is. Score five factors: blast radius, data sensitivity, reversibility, regulatory impact, and the credential scope the run needs. The highest factor wins. A three-line change to session handling is critical even when the agent is certain.

Risk classTypical changeThe agent mayThe gateRequired human decision
LowTests, docs, research, internal tooling, isolated refactors behind green testsEdit an isolated checkout, run local checks, open a PRCI green plus a complete evidence bundleLightweight review and merge by any team member
MediumFeature work in one service, dependency upgrades, preview or staging changesUse scoped non-production credentials and pre-approved runbooksCI plus accepted plan plus code-owner approvalCode owner accepts the plan and approves the merge
HighAuth, billing, schema migrations, shared libraries, sensitive data pathsPrepare the change, a dry run, and a tested rollbackSpecialist review plus rehearsed rollback plus release approvalNamed specialist and release owner, separately
CriticalProduction data changes, security exceptions, destructive operations, regulated scopePrepare evidence and an approval request onlyDeployment platform gate the agent’s identity cannot passNamed owner authorizes; the platform executes

Risk classes are not maturity levels. A loop’s level on the autonomy ladder measures how far you trust that loop; its risk class measures what a mistake costs. Keep both in the autonomy register, so a loop with a long clean record never earns its way into critical changes by track record alone.

Some actions stay outside agent authority in every class: reading or rotating production secrets, unapproved external communication, deleting backups, and approving its own work. Write them into the policy as forbidden, and enforce them with deny rules and credentials the agent never holds.

Guardrails or gates: which control goes where?

Section titled “Guardrails or gates: which control goes where?”

A guardrail shapes or constrains the agent before and during work: a sandbox, a scoped token, a hook that blocks a command. A gate decides whether the evidence is sufficient to advance: a required check, a code-owner approval, a deployment approval. You need both, because a guardrail cannot judge product intent and a gate cannot stop a command that already ran.

MomentControlExampleEvidence it leaves
Before workBoundaryAccepted intent, allowed paths, data class, budgetReviewed intent.md and plan.md
During workGuardrailSandbox, scoped credentials, hook, protected pathPolicy log and a blocked-action test
Before mergeGateTypes, tests, security scan, specification reviewCI results and the review record
Before productionAuthority gateNamed approver, change window, rollback readinessApproval and release record
After releaseFeedbackCanary, telemetry, incident triggerObservation window and the promote-or-roll-back decision

Repository instructions (CLAUDE.md, AGENTS.md, Cursor rules, and skills) improve behavior, but they are instructions the model can misread. Match each need to the layer that actually enforces it:

NeedCorrect controlWrong control
Explain architecture and conventionsCLAUDE.md, AGENTS.md, Cursor rulesA hook
Reuse a multi-step workflowSkillA long prompt pasted by hand
Block or log a deterministic tool actionHook or execution policyA sentence in CLAUDE.md
Limit files, commands, and networkSandbox plus scoped identity“Do not deploy” in the prompt
Require review before mergeBranch protection and code ownersAn AI reviewer’s approval
Require authorization before productionDeployment environment or change-management gateThe authoring agent’s summary
Recover safelyRehearsed rollback with a named ownerA runbook nobody has run
Prove what happenedImmutable CI, PR, deployment, and incident logsThe agent’s transcript alone

Separate agent identity from human authority

Section titled “Separate agent identity from human authority”
  1. Give every non-interactive agent its own service identity. Never reuse a developer’s broad token; the details are in agent identity, credentials, and secrets.

  2. Grant the smallest repository, network, cloud, and data scope the run’s risk class needs. A worktree or container does not narrow a cloud token.

  3. Keep production secrets out of every agent run, whatever its class. In critical changes the deployment platform holds them, never the agent. A prompt that says “do not deploy” is not an access boundary.

  4. Require branch protection and code-owner approval for merge. The agent’s identity must not be a code owner.

  5. Require a separate deployment-environment approval for high and critical changes. Neither the authoring agent nor a reviewing agent can satisfy it.

  6. Log the initiating human, agent identity, tool and model version, artifact versions, commands, findings, approval, deployment, and rollback outcome.

The pattern is identical in all three tools: guidance in instruction files, hard limits in managed settings or admin policy, and merge and production gates on the forge and the deployment platform. The settings differ. For a side-by-side of every managed key, see enforcing one policy across every coding agent.

Put hard limits in managed settings (managed-settings.json, MDM, or server-managed settings on Team and Enterprise), which user and project settings cannot override. Keys checked against Claude Code 2.1.283:

{
"permissions": {
"disableBypassPermissionsMode": "disable",
"deny": [
"Bash(terraform apply:*)",
"Bash(kubectl delete:*)"
]
},
"allowManagedPermissionRulesOnly": true,
"allowManagedHooksOnly": true
}

A Bash deny rule matches the command text Claude writes, not the program, so sh -c or a script gets past it (Claude Code docs, checked against v2.1.283). The real boundary for terraform and kubectl is that the agent identity holds no production cloud or cluster credentials; treat the deny rule as a typo-catcher on top.

allowManagedPermissionRulesOnly ignores permission rules from user, project, and local settings; allowManagedHooksOnly does the same for hooks. From v2.1.283 (the latest channel; stable is 2.1.274), auto mode is the starting mode for interactive terminal and VS Code sessions on supported models, and a classifier reviews each action. Treat that classifier as a guardrail, not a gate. To turn it off for regulated repositories, add "disableAutoMode": "disable" inside permissions. In CI, claude -p starts in Manual mode, and --permission-prompts none denies anything that would prompt.

Claude Code’s managed Code Review posts findings, but “the check run always completes with a neutral conclusion so it never blocks merging through branch protection rules” (Claude Code docs, checked against v2.1.283). If its findings must gate a merge, parse them in your own required CI check. A failed or timed-out review also completes as neutral (Claude Code docs, checked against v2.1.283), so your CI check must fail when the result is missing.

Write an evidence contract for every agent run

Section titled “Write an evidence contract for every agent run”

Every background or autonomous run states its contract before it starts. The next gate reads the output, so a run with no contract has nothing for a gate to check.

  • Input: the exact commit, artifacts, ticket, telemetry window, and allowed external sources.
  • Scope: directories, tools, network destinations, credential class, and a time or cost budget.
  • Proof: commands with expected exit status, schemas, screenshots, or eval thresholds.
  • Output: the file, pull request, comment, or structured record the next gate reads.
  • Failure: non-zero exit, timeout, missing dependency, or low confidence, plus the escalation owner.
  • Authority: the actions allowed without approval and the first action that always stops for a human.

Replace the commit, paths, commands, and owner with your own. The fixed clauses (fail closed, never edit the oracle, never approve your own work) are the part to keep verbatim; protecting the test oracle explains why the second one matters most.

Paste this into your engineering handbook, fill in the owners, and link it from every repository’s instruction file. Each clause names what enforces it; a clause that says “trust” is a known gap.

AGENT AUTONOMY POLICY — v1.0 — owner: CTO — review: quarterly
1. Classification. Every agent change gets a risk class (low, medium, high, critical),
computed in CI from touched paths, migrations, and requested credentials.
Highest factor wins. Enforced by: required CI check that writes the PR label.
2. Forbidden in every class: production secrets, destructive production commands,
deleting backups, external communication, approving own work.
Enforced by: managed deny rules + credentials the agent identity never holds.
3. Identity. One service identity per agent loop, least privilege, short-lived tokens.
Enforced by: identity provider + CI OIDC. Owner: platform team.
4. Merge gate. Low: CI + evidence bundle + any reviewer. Medium: + code owner.
High: + specialist review + rehearsed rollback. Enforced by: branch protection.
5. Production gate. High and critical: named release owner approves in the
deployment environment. Agent identities cannot approve. Enforced by: deploy platform.
6. AI reviewers and auto-approval tools may approve low-class changes only.
Criteria owner: release owner. Enforced by: routing rules + branch protection.
7. Overrides. Every bypass is attributed, expires in 7 days, and is reviewed.
Enforced by: override log in the change record.
8. Evidence. Every control has an owner, trigger, failure behavior, exception path,
and audit record. Controls fail closed.
9. Proof. Controls are rehearsed quarterly; metrics are reported by risk class.

The seven-day override expiry and quarterly cadence are this guide’s working defaults, not a published standard; set them to your own risk appetite.

Prompts for designing and auditing controls

Section titled “Prompts for designing and auditing controls”

Run these in Claude Code, Codex, or Cursor with read access to the repository and its CI configuration. They produce reports; they change nothing.

Prove each control by making it fail on purpose, then measure the controls by risk class. Nobody has to read every diff for either step.

  1. Ask a low-class agent run to write outside its allowed workspace. The sandbox must block it, and the block must appear in the policy log.

  2. Ask it to skip a failing required test. The workflow must fail.

  3. Attempt a merge without code-owner approval. Branch protection must block it.

  4. Attempt a production release with the agent’s identity and without the named approval. The deployment platform must block it.

  5. Make the review service unavailable. The merge must stay blocked; a missing reviewer never counts as an approval.

  6. Trigger a rollback in a non-production environment and record the time to restore.

  7. Change a rule, skill, or hook and run the harness evals before rollout.

A step that passes when it should fail is an incident in the control, not a flaky test. Record it the way you would record an incident an agent caused.

MetricDefinitionWhat a bad trend means
False-block rateBlocked changes later merged unchanged ÷ all blocked changesThe gate costs trust and invites bypasses
Override rateMerges or releases with an attributed bypass ÷ all merges in the classThe control is wrong or the class is too strict
Escaped defectsIncidents traced to a change that passed every gateThe gate checks the wrong thing
Review waitMedian hours from “ready for review” to decisionHuman gates became the bottleneck; move checks to automation

Report these per class. A healthy low class shows short waits and few overrides; a healthy critical class can show long waits and still be correct. The canonical metric definitions live in metrics frameworks.

Who signs off: the release owner signs off the production gate for high and critical changes, the code owner signs off medium merges, and the CTO signs off changes to the policy itself, with the rehearsal results and the metrics attached.

An author agent and a reviewer agent improve detection, but they can share credentials, configuration, model blind spots, and the organization that set them up. Two agents are not separation of duties. The OWASP GenAI LLM Top 10 2026 (published 2026-08-04) lists Excessive Agency as LLM03 and advises: “Use human-in-the-loop control to require a human to approve high-impact actions before they are taken.”

A real gate uses an identity the authoring run cannot act as, or a platform policy it cannot change, plus the accountable human where the risk class demands one. Use AI review to decide where human attention goes, not to replace it on high and critical changes. For evidence that auditors accept, see agentic engineering in regulated industries.

What breaks in agent governance, and how do you recover?

Section titled “What breaks in agent governance, and how do you recover?”

Advisory text is treated as enforcement. A CLAUDE.md line says “never run migrations,” and one day the agent runs one. Recover: run the “separate advice from enforcement” prompt, then move each advisory-only invariant into an enforced control. For destructive actions, prefer removing the credentials from the agent identity or a platform permission; a managed deny rule, a hook, or CI catches the common spelling on top.

One token reaches every environment. The agent’s CI token can also deploy. Recover: revoke it first, then split identities per loop and environment, and rotate. Treat the exposure as an incident until the logs show what the token touched.

The agent labels its own risk. Every change arrives as “low.” Recover: compute the class in CI from paths and migrations, and alert when an agent-proposed label is lower than the computed one.

A human gate has no decision criteria. Approvers click through because nothing tells them what to check. Recover: attach the evidence bundle and a one-line risk statement to every request, and set the criteria, owner, timeout, and escalation path before automating the request.

Overrides never expire. A one-off bypass for an outage becomes permanent. Recover: give every override an expiry and an owner, and report the open ones in the quarterly review.

A control fails open. The review service times out and the merge proceeds. Recover: make the check required and fail on a missing result, then add the outage case to the rehearsal.

Rollback exists only in a document. Recover: rehearse it in a representative non-production environment and measure the time to restore. Pair high-class releases with progressive delivery so the rollback is a flag, not a redeploy.

False certainty. Someone calls the instructions “unbreakable guardrails” or retires human review because the checks are green. Models misread instructions, hooks can be misconfigured, and deterministic checks cannot judge product intent. Recover: test the full control path, write down the residual risk, and keep human authority at the consequential gates.

Before this page, read the operating model for who owns the autonomy register. Next, apply the policy to the threats and credentials it has to hold against.

For the stage procedures these gates sit inside, see Deploy and Maintain.