Skip to content

Shared hooks governance: reviewed runtime controls

Shared hooks governance treats every agent hook an organization distributes as production code: one narrow purpose, a written contract, reviewed and version-pinned source, fixtures that prove allow, deny, and failure behaviour, managed distribution, staged rollout, audit logs, and a rehearsed rollback. Distribution alone is not maturity, because an auto-installed hook spreads a defect as fast as a control.

This page is for the tech lead or CTO who owns question 18 of the CTO Scorecard and for the platform engineer who ships the hooks. The situation it prevents: your platform team pushed a secret-guard hook to every laptop, half the team had no jq, the hook crashed on every call, the agent treated the crash as a non-blocking error, and for three weeks the “enforced” control enforced nothing. Nobody noticed, because a failing hook looks exactly like a quiet one.

Q18 · Organization enablement: How do you govern shared agent hooks and their tool-specific adapters?

Max-score answer: versioned, reviewed, least-privilege hooks with fixtures, staged rollout, audit logs, and a tested rollback.

  • A hook contract template that a reviewer can approve before any code exists
  • A fail-closed hook skeleton that works unchanged in Claude Code and Codex
  • A fixture runner that proves allow, deny, and failure behaviour in CI without anyone reading the script
  • Managed distribution for each tool, so a pinned version is the only version that runs
  • A rollout ladder, four metrics with owners, and a rollback drill you can run this quarter
  • Three copy-paste prompts: write the contract, attack the implementation, and generate the fixtures

The developer-level view of what to put in a hook is in hooks as deterministic guardrails. This page covers only the organizational layer: who ships hooks, how they are proven, and how they are withdrawn.

Why do shared hooks need stricter governance than rules or skills?

Section titled “Why do shared hooks need stricter governance than rules or skills?”

A rule or a skill is advice the model reads. A hook is a program that runs with the developer’s credentials at a fixed point in the agent loop, whatever the model decides. That makes a shared hook part of your execution boundary: a hook with a defect blocks work for everyone at once, and a compromised hook runs on every machine that installed it.

The second reason is less obvious. In both CLIs, most ways a hook can fail let the action through. The table records what each tool did when checked on 2026-09-26.

BehaviourClaude Code 2.1.283 (latest channel, checked 2026-09-26)Codex 0.157.1
What blocks a PreToolUse callExit code 2 (stderr becomes the reason), or JSON with permissionDecision: "deny"Exit code 2 with a non-empty stderr reason, or a JSON block decision
Exit code 1 or any other non-zero codeNon-blocking error; the tool call proceedsHook marked failed; the tool call proceeds
Script path missing or not executableNon-blocking error; the tool call proceeds, so “a mistyped path in settings.json leaves the gate silently disabled” (Anthropic docs)Same outcome: the run fails and does not block
TimeoutOutput discarded, no decision; “don’t count on a stalled hook to act as a gate”Hook marked failed; no block
Default command timeout600 seconds600 seconds
Changed hook definitionNo per-hook trust record; the merged settings decide what runsA non-managed hook whose definition changed is marked modified and does not run until trusted again

Sources: the Claude Code hooks reference and the Codex source at tag rust-v0.157.1 (codex-rs/hooks/src/events/pre_tool_use.rs, codex-rs/hooks/src/engine/discovery.rs). Cursor runs hooks as “spawned processes that communicate over stdio using JSON in both directions” that “can observe, block, or modify behavior” (Cursor hooks docs, checked 2026-08-28); its current failure semantics were not re-verified for this page, so prove them with the fixture runner below before you rely on them.

So “fail closed” is never a setting. It is something the hook’s own code must do: catch every failure it can detect and exit 2 with a reason. And because a timeout or a crash still fails open, a hook is a guardrail, not the security boundary. The boundary is the permission and sandbox layer described in permissions and sandboxing.

Every shared hook starts as a reviewed contract. Commit it next to the code, and merge it through the same review as the hook. Copy this template as-is.

hooks/guard-env-files/CONTRACT.md
# guard-env-files — contract
Owner: platform team (#platform-hooks). Version: 3. Status: blocking.
Mode switch: GUARD_MODE=advisory logs would-be blocks and exits 0.
## Purpose (one sentence)
Stop agent shell commands that read or write `.env` files.
## Trigger
Event: PreToolUse. Matcher: the shell tool ("Bash" in Claude Code and Codex).
## Input
Reads only `tool_input.command`. Treats the whole input as untrusted text:
never evaluates it, never interpolates it into a shell.
## Authority
No network. No credentials. Reads nothing but stdin.
Writes one log line per run to a local file, never to stdout.
## Decision
Block (exit 2, reason on stderr) when the command references a `.env` file.
Allow (exit 0, no output) otherwise.
## Failure mode: fail closed
Missing jq, unreadable input, or a missing command field → exit 2 with a reason.
Timeout: 10 s. A timeout fails open in every supported tool, so the
permission deny rule on `.env` files stays the real boundary.
## User-visible message
Names the hook, its version, why it blocked, and the doc to read.
## Logs
One JSON line per run in $HOOK_LOG (default
~/.local/state/company-hooks/hooks.jsonl): hook, version, event, mode,
decision, reason_code, session_id, duration_ms. Never the command text. Retention: 30 days.
## Rollback
Managed settings point back to v2. Drill: quarterly. Last drill: (date).

The contract answers the questions a reviewer would otherwise have to reverse-engineer from the script: what the hook may touch, what happens when it breaks, and how to turn it off.

Both CLIs pass the same JSON on stdin for a PreToolUse shell call, with the command at tool_input.command, and both block on exit 2 with a stderr reason. One script therefore serves both tools, and only the registration differs. This is the pattern from shared agent rules: one core, thin tested adapters.

#!/usr/bin/env bash
# hooks/guard-env-files/hook.sh — v3 · PreToolUse · shell tool · owner: platform team
set -uo pipefail
MODE=${GUARD_MODE:-blocking} # "advisory" logs would-be blocks and exits 0
LOG=${HOOK_LOG:-$HOME/.local/state/company-hooks/hooks.jsonl}
ms() { local t=${EPOCHREALTIME:-}; if [ -n "$t" ]; then t=${t/[.,]/}; echo $((t / 1000)); else echo $(($(date +%s) * 1000)); fi; }
start=$(ms); sid=unknown
# One JSON line per run, appended to a file: never stdout, never the command text.
log() {
mkdir -p "${LOG%/*}" 2>/dev/null
printf '{"hook":"guard-env-files","version":3,"event":"PreToolUse","mode":"%s","decision":"%s","reason_code":"%s","session_id":"%s","duration_ms":%s}\n' \
"$MODE" "$1" "$2" "$sid" "$(($(ms) - start))" >>"$LOG" 2>/dev/null || true
}
block() {
log block "$1"
printf 'guard-env-files v3: %s See docs/hooks/guard-env-files.md\n' "$2" >&2
[ "$MODE" = advisory ] && exit 0
exit 2
}
command -v jq >/dev/null 2>&1 || block missing_jq "jq is missing, so the check cannot run."
input=$(cat)
sid=$(jq -r '.session_id // "unknown"' <<<"$input" 2>/dev/null | tr -cd '[:alnum:]_-'); sid=${sid:-unknown}
jq -e . >/dev/null 2>&1 <<<"$input" || block malformed_input "the hook input is not valid JSON."
cmd=$(jq -er '.tool_input.command // empty' <<<"$input" 2>/dev/null) || block no_command "the hook input has no command field."
env_re='(^|[[:space:]/="'\''(<>;|&])\.env(\.[[:alnum:]_-]+)?([[:space:]"'\'';|&)<>]|$)'
if printf '%s' "$cmd" | grep -Eq "$env_re"; then
block env_file "a shell command touching a .env file was blocked. Read secrets through the secrets manager instead."
fi
log allow none
exit 0

Four properties make this a governed hook rather than a script someone wrote once. Every detectable failure exits 2 with a message, so the missing-jq laptop blocks loudly instead of passing silently. The input is only ever read as data; nothing from it reaches eval or an unquoted expansion, and the pattern also matches a .env path followed by a shell separator, so cat .env; ls and cat .env&&ls block too. Every run appends one JSON line to a local log file, never to stdout and never with the command text, which is what the rollout metrics below are computed from. And the message tells the developer what happened and where to go, which is what stops people from hunting for a way around it. GUARD_MODE=advisory keeps the logging and the message but exits 0, which is the first rollout stage.

Prove the hook with fixtures, not by reading it

Section titled “Prove the hook with fixtures, not by reading it”

A reviewer should not have to read hook.sh to know it works. Each fixture is a directory with the JSON input, the expected exit code, optionally a phrase the stderr reason must contain, and optionally an expect-stdout-empty marker for fixtures that must print nothing. The runner feeds each input to the hook exactly as the agent would, kills it at the contract’s 10-second timeout, and sends the log to a temporary file so CI never writes to a home directory.

#!/usr/bin/env bash
# hooks/run-fixtures.sh — usage: hooks/run-fixtures.sh hooks/guard-env-files
# Each fixture dir holds input.json, expect-exit, and optionally expect-stderr
# (a phrase the reason must contain) and expect-stdout-empty (a marker file).
set -u
hook_dir=$1; failed=0
tmp=$(mktemp -d); trap 'rm -rf "$tmp"' EXIT
export GUARD_MODE=blocking HOOK_LOG="$tmp/hooks.jsonl"
for f in "$hook_dir"/fixtures/*/; do
name=$(basename "$f")
# 10 s is the contract timeout: a hook that needs longer fails here with exit 124.
out=$(timeout 10 "$hook_dir/hook.sh" < "$f/input.json" 2>"$tmp/stderr"); code=$?
want=$(cat "$f/expect-exit")
if [ "$code" != "$want" ]; then
echo "FAIL $name: exit $code, expected $want"; failed=1; continue
fi
if [ -f "$f/expect-stderr" ] && ! grep -qF "$(cat "$f/expect-stderr")" "$tmp/stderr"; then
echo "FAIL $name: stderr did not contain the expected reason"; failed=1; continue
fi
if [ -f "$f/expect-stdout-empty" ] && [ -n "$out" ]; then
echo "FAIL $name: the hook wrote to stdout"; failed=1; continue
fi
echo "ok $name"
done
exit $failed

Run it as a required CI check on the hooks repository. The minimum fixture set for any blocking hook:

FixtureInputExpected
allow-commonAn ordinary command such as ls -la srcExit 0, no output (expect-stdout-empty)
deny-targetThe case the hook exists for, such as cat .env.productionExit 2, stderr names the hook and the reason
deny-chainedThe target followed by a shell separator, such as cat .env && lsExit 2
allow-near-missSomething that looks similar but is fine, such as cat .envrcExit 0 (catches over-blocking)
malformed-inputnot jsonExit 2 (proves fail-closed)
missing-field{"tool_input":{}}Exit 2
injectionA command containing $(…), backticks and quotesSame decision as the plain form; nothing executed
slow-dependencyRun with a dependency stubbed to hangFinishes within the contract’s 10-second timeout; the runner reports exit 124 if it does not

The runner proves the core. To prove each adapter, run the same fixtures once through each tool in a scratch repository and confirm the block reaches the agent: that is the only way to catch a wrong matcher or a registration that never loads.

Ship hooks so only the pinned version runs

Section titled “Ship hooks so only the pinned version runs”

Where a hook is installed decides who can change it. A hook in a repository’s settings file changes whenever someone merges a change to that file; a managed hook changes only when the platform team publishes one. For controls that must run, use the managed channel and a versioned path, so moving from v2 to v3 is a one-line, reviewable change and moving back is the same line.

Deploy through managed settings: a managed-settings.json file (/etc/claude-code/ on Linux and WSL, /Library/Application Support/ClaudeCode/ on macOS, C:\Program Files\ClaudeCode\ on Windows), MDM, or server-managed settings from the claude.ai console on Team and Enterprise.

{
"allowManagedHooksOnly": true,
"hooks": {
"PreToolUse": [
{
"matcher": "Bash",
"hooks": [
{ "type": "command", "command": "/opt/company-hooks/guard-env-files/v3/hook.sh", "timeout": 10 }
]
}
]
}
}

allowManagedHooksOnly blocks user, project, local, and plugin hooks, except plugins that managed settings force-enable. A disableAllHooks set in user or project settings cannot switch managed hooks off; only managed settings can. Before you approve a plugin that ships hooks, list what it brings with claude plugin details <name>. To see hooks load and fire, start a session with claude --debug hooks.

Roll out in stages and measure four things

Section titled “Roll out in stages and measure four things”

A hook that blocks work on day one for the whole company teaches people to route around it. Promote it through stages, and let the metrics decide each promotion, not the calendar.

  1. Advisory, one team. Run the hook with GUARD_MODE=advisory, for example through a two-line wrapper at a versioned path that exports it and calls hook.sh, so it works the same in every tool. The hook logs what it would have blocked and exits 0. Run for one or two weeks and read every would-be block.
  2. Blocking, one team. Flip to exit 2 for the volunteer team. Publish the exception channel before you flip.
  3. Blocking, half the org. Expand when the false-block rate and latency stay under the thresholds you set for two consecutive weeks.
  4. Blocking, everyone. Set allowManagedHooksOnly (Claude Code) or allow_managed_hooks_only (Codex) once all controls you need are managed, so locally added hooks cannot shadow or duplicate them.
  5. Rollback drill. Before general availability, and then quarterly, point the managed setting back to the previous version and confirm clients pick it up.

Each metric has an owner and a threshold you choose; these are starting definitions, not benchmarks.

MetricDefinitionSignal of trouble
False-block rateBlocks overturned by an approved exception ÷ all blocks, per weekRising rate means the rule is too broad; people will start bypassing it
Hook error rateRuns that exited non-zero other than 2, or timed out ÷ all runsAny sustained errors mean the control is silently off somewhere
p95 latency95th percentile of duration_ms per eventThe hook runs on every matching tool call, so latency multiplies across a session
Bypass findingsViolations found later by CI, review, or secret scanning that the hook should have blockedEach one is a missing fixture; add it before changing the rule

The log line from the contract is what makes these computable without anyone reading transcripts: hook, version, event, mode, decision, reason_code, session_id, and duration_ms, with no command text.

Who signs off, and what counts as evidence?

Section titled “Who signs off, and what counts as evidence?”

Nobody approves a hook by reading its source alone. The approval rests on artifacts a reviewer can check in minutes:

  • The contract is merged, with an owner and a version.
  • The fixture runner passes in CI for the core, and each adapter has a recorded run through its real tool.
  • The managed configuration pins a versioned path, and the change to it went through code-owner review.
  • The advisory-stage log shows the would-be block list, read and signed off by the owning team.
  • The last rollback drill has a date, and the time from decision to recovered clients is written down.

The hook owner signs off the contract and fixtures; the security or platform lead signs off promotion to blocking and to managed-only. That split matches the governance and autonomy model: the hook is a guardrail in the agent loop, and CI remains the gate that nothing merges without.

What breaks when you share hooks across a team?

Section titled “What breaks when you share hooks across a team?”

The gate is silently off. A wrong path, a missing dependency, or exit code 1 produces a non-blocking error in both CLIs, and work continues. Recovery: add the malformed-input and missing-field fixtures, make the hook’s own failures exit 2, and alert on the hook error rate instead of waiting for someone to notice.

A bad release blocks the whole company. A v4 release widens the pattern so it also matches src/environment.ts, and every agent session stalls. Recovery: point the managed setting back to v3, which is why the path is versioned. A disableAllHooks in a developer’s own settings cannot switch a managed hook off in Claude Code, so the managed channel is the only fast rollback; rehearse it before you need it.

The Codex hook went dark after an upgrade. A version bump in a repository hooks.json changed the definition hash, and nobody re-trusted it. Recovery: move must-run hooks to managed_hooks in requirements.toml; keep repository hooks for conveniences whose absence costs nothing.

The hook ran untrusted code in CI. An agent job checked out a contributor’s branch, and that branch’s .claude/settings.json hooks ran with the job’s secrets. Recovery: check out trusted configuration from the default branch, run the agent against the contributor’s code as read-only data, and start Claude Code with --bare --setting-sources "" --strict-mcp-config for that run, so no hooks, settings or MCP servers from the checkout load. --settings '{"disableAllHooks": true}' alone is not enough: it leaves the checkout’s MCP servers and permission rules in play. The full pattern is in AI in CI/CD.

One hook grows to do everything. A single “policy” hook accretes secret checks, formatting, and ticket lookups, gains network access and a token, and can no longer be tested or rolled back in parts. Recovery: split it by purpose, one contract per hook, and move anything that needs a persistent connection or credentials into a reviewed internal MCP server.

Put the hooks in context with the rest of the shared harness: shared agent rules for instructions and shared skills for procedures. For a single policy across every agent you run, see managed policy across coding agents, and for the threats hooks help contain, the agent threat model.