Skip to content

Security Gates for Agent-Written Code: Semgrep, Gitleaks, TruffleHog, Snyk and Claude Security Review

Security gates for agent-written code are deterministic scanners and review agents placed at four points: inside the agent session, at git commit, in CI, and on the pull request. Gitleaks and TruffleHog stop secrets, Semgrep and Snyk find vulnerable code and dependencies, and Claude’s security review adds a model’s judgement. Only CI can be made mandatory for every contributor.

Your agent wires up Stripe checkout in 12 minutes, the tests pass, and the diff contains new Stripe("sk_live_…") because the key was in the .env file it read for context. Nobody reads line 40 of a 600-line diff, and the agent commits every few minutes. This page is for developers who want the agent to be stopped before a secret or an injection reaches main, and for tech leads who need that stop to hold for eight people and three different agents.

What you’ll set up with these security gates

Section titled “What you’ll set up with these security gates”
  • A Gitleaks pre-commit hook that blocks an agent’s commit containing a secret.
  • A four-layer workflow with the tool for each layer in Claude Code, Codex and Cursor.
  • Two tested Claude Code hooks: one scans every written file, one stops the agent skipping commit hooks.
  • CI workflows for TruffleHog, Semgrep and a hardened claude-code-security-review.
  • Three copy-paste prompts, the metrics that prove the gates work, and the traps.

This page covers the scanners. The review bots that read a diff for logic bugs are compared in AI code review bots; the procedure a human follows on an agent’s pull request is in reviewing an agent’s pull request.

Every tool below was checked on 2026-09-26 against its own README, the npm or PyPI registry, Anthropic’s plugin marketplace file, or the installed CLIs (Claude Code 2.1.283, Codex CLI 0.157.1).

LayerToolCatchesRuns whereCan it block?Cost
1. In sessionsecurity-guidance plugin (Anthropic)~25 dangerous patterns on edit; LLM diff review at end of turn; agentic review on git commitClaude Code hooksFeeds findings back to Claude; you can switch it offModel tokens per review
1. In sessionSemgrep Guardian (plugin: MCP server, hooks, skills)Code, Supply Chain and Secrets findings on every file the agent writesClaude Code, Cursor; semgrep mcp in any MCP clientPrompts the agent to regenerate until cleanFree CLI; Pro rules need semgrep login
1. In session/security-review (built into Claude Code)Injection, auth and data-exposure risks in the branch diffYour sessionNo, it reportsNormal usage
2. CommitGitleaks pre-commit hookHardcoded secrets in staged changesYour machine, any agentYes, unless skippedFree
3. CITruffleHogSecrets, verified live against the issuerGitHub Actions or any CIYes, --fail exits 183Free
3. CISemgrepVulnerable code patterns, new findings onlyAny CIYes, --error exits 1Free CLI
3. CISnyk MCP and CLI; SonarQube MCPVulnerable dependencies and code; quality gatesAgent session and CIThrough your Snyk or SonarQube gateVendor plan
4. Pull requestanthropics/claude-code-security-reviewSemantic vulnerabilities, with false-positive filteringGitHub ActionsOnly through a check you buildAPI tokens
ToolingSnyk Agent ScanPrompt injection and tool poisoning in MCP configs and skillsYour machine or CI--ci exits non-zeroSnyk token

The split that matters: layers 1 and 2 run on the developer’s machine, where the agent or the developer can turn them off. Layer 3 runs where nobody can, so it is the only one you make a required check. Layers 1 and 2 exist to make layer 3 boring.

Popularity on 2026-09-26: gitleaks/gitleaks had 29.5k GitHub stars, trufflesecurity/trufflehog 28.1k and semgrep/semgrep 16.8k; in Anthropic’s plugin directory (claude.com/plugins), security-guidance showed 241,800 installs. Stars and installs measure attention, not detection quality.

Example: a Gitleaks pre-commit hook stops an agent’s commit

Section titled “Example: a Gitleaks pre-commit hook stops an agent’s commit”

This is the smallest gate with the largest payoff, and it works for every agent, because every agent commits through git.

  1. Install Gitleaks and the pre-commit framework (PyPI pre-commit 4.6.2 on 2026-09-26):

    Terminal window
    # Terminal
    brew install gitleaks
    pipx install pre-commit
  2. Add .pre-commit-config.yaml at the repository root. The rev below is the one Gitleaks’ README pins; pre-commit autoupdate moves it to the latest release.

    repos:
    - repo: https://github.com/gitleaks/gitleaks
    rev: v8.24.2
    hooks:
    - id: gitleaks

    The gitleaks hook runs gitleaks git --pre-commit --redact --staged --verbose, so it scans only what is staged and never prints the secret. If Gitleaks is already installed through Homebrew, the gitleaks-system hook ID uses that binary instead of building one.

  3. Install the hook and pin it to the latest release:

    Terminal window
    pre-commit autoupdate
    pre-commit install

    pre-commit install writes .git/hooks/pre-commit, which is per clone. Every developer runs it once, and every new worktree shares it through the common .git directory.

  4. Let the agent commit. When the staged diff contains a live Stripe key, the commit fails and the agent reads this in its shell output (the finding block is real Gitleaks 8.24.2 output):

    Detect hardcoded secrets.................................................Failed
    - hook id: gitleaks
    - exit code: 1
    Finding: ...tripe = new Stripe("REDACTED");
    Secret: REDACTED
    RuleID: stripe-access-token
    Entropy: 5.280395
    File: src/billing.ts
    Line: 1
    Fingerprint: src/billing.ts:stripe-access-token:1
  5. Check what the agent does next. The right recovery is to replace the literal with process.env.STRIPE_SECRET_KEY, unstage, and commit again. The wrong recoveries are git commit --no-verify, SKIP=gitleaks git commit, or a #gitleaks:allow comment, all of which Gitleaks documents. The second Claude Code hook in the next section denies the first two; the CI layer catches the first two, and a #gitleaks:allow comment too when TruffleHog can verify the key or when you add Gitleaks to CI with --ignore-gitleaks-allow (Layer 3).

  6. Rotate the key anyway. The secret was in a file the agent read, so it was in the model’s context and in the session transcript. A blocked commit means it did not reach Git history, not that it stayed private.

Each layer is cheaper than the next and catches what the previous one missed:

  1. Hook, in the session. Each written file is scanned and findings go back to the agent before the turn ends.
  2. Pre-commit, on the machine. Gitleaks blocks the commit, whichever agent, IDE or human made it.
  3. CI, on every push and pull request. TruffleHog and Semgrep run where no agent can skip them. This is the required check.
  4. Security review, on the pull request. A model reads the diff for flaws no pattern matches, such as an authorization check on the wrong object, and a human on your escalation list reads its findings.

The in-session layer differs most between the tools, so pick your tab.

Anthropic’s security-guidance plugin runs three layers as hooks: regex warnings on Edit and Write for about 25 dangerous patterns (yaml.load, pickle.load, raw innerHTML, hardcoded secrets), an LLM review of the diff when Claude finishes a turn, and an agentic reviewer on git commit that reads related files to trace data flow. It needs Claude Code 2.1.144 or later and Python 3.8 or later.

Terminal window
# Terminal
claude plugin install security-guidance@claude-plugins-official

Put rules your codebase needs in .claude/claude-security-guidance.md and commit it. The combined user, project and local files have an 8 KB budget, and the commit reviewer does not read them. In multi-agent setups that share one worktree, set ENABLE_STOP_REVIEW=0, because another agent can move HEAD between turns.

Semgrep Guardian bundles the Semgrep MCP server, hooks and skills, and rescans every file the agent writes with Semgrep Code, Supply Chain and Secrets:

# Inside Claude Code
/plugin install semgrep@claude-plugins-official
/setup-semgrep-plugin

/setup-semgrep-plugin also installs the Semgrep CLI. Before you push, run /security-review, which reviews the diff between your branch and the default branch on origin.

Two hooks of your own. The first scans each written file with Gitleaks, so the agent hears about a secret seconds after writing it rather than at commit time. Save it as .claude/hooks/scan-secrets.sh and make it executable:

#!/usr/bin/env bash
# .claude/hooks/scan-secrets.sh (PostToolUse, matcher "Edit|Write")
set -uo pipefail
command -v gitleaks >/dev/null || { echo "gitleaks not installed; secret scan skipped" >&2; exit 0; }
file=$(jq -r '.tool_input.file_path // empty')
[ -z "$file" ] || [ ! -f "$file" ] && exit 0
if ! out=$(gitleaks dir "$file" --redact --no-banner --verbose 2>&1); then
echo "Gitleaks found a hardcoded secret in $file:" >&2
echo "$out" | grep -E '^(RuleID|Line):' >&2
echo "Remove the literal. Read the value from an environment variable instead, and do not add gitleaks:allow." >&2
exit 2
fi
exit 0

Exit code 2 puts the message in front of Claude, which fixes the file on its next step. The second hook denies the commands that skip the pre-commit layer. Save it as .claude/hooks/guard-commit.sh:

#!/usr/bin/env bash
# .claude/hooks/guard-commit.sh (PreToolUse, matcher "Bash")
set -uo pipefail
cmd=$(jq -r '.tool_input.command // empty')
deny() {
jq -nc --arg r "$1" '{hookSpecificOutput: {hookEventName: "PreToolUse",
permissionDecision: "deny", permissionDecisionReason: $r}}'
exit 0
}
[[ $cmd =~ git[[:space:]] ]] || exit 0
[[ $cmd =~ --no-verify|(^|[[:space:]])-[a-zA-Z]*n[a-zA-Z]*([[:space:]]|$)|SKIP=|core\.hooksPath ]] &&
[[ $cmd =~ git[[:space:]]+(-c[[:space:]]+[^[:space:]]+[[:space:]]+)*(commit|push|config) ]] &&
deny "Commit hooks must run. Fix what the hook reported and commit again without --no-verify, -n, SKIP= or a hooksPath override."
exit 0

Tested against eight commands: it denies git commit --no-verify, git commit -nm, SKIP=gitleaks git commit, git -c core.hooksPath=/dev/null commit, git push --no-verify and git config core.hooksPath, and allows git commit -m "fix n+1 query" and git log -n 5. It can false-positive on a commit message that contains a standalone -n word or on a compound command; the agent then rephrases, which is the safe failure direction. Register both in .claude/settings.json:

{
"hooks": {
"PreToolUse": [
{ "matcher": "Bash",
"hooks": [{ "type": "command", "command": "${CLAUDE_PROJECT_DIR}/.claude/hooks/guard-commit.sh", "args": [] }] }
],
"PostToolUse": [
{ "matcher": "Edit|Write",
"hooks": [{ "type": "command", "command": "${CLAUDE_PROJECT_DIR}/.claude/hooks/scan-secrets.sh", "args": [], "timeout": 30 }] }
]
}
}

The hook format and the deny semantics are covered in Claude Code hooks for automation.

Layer 2 is the Gitleaks hook from the example above; add pre-commit install to the repository’s setup script so a fresh clone or agent sandbox gets it too.

Layer 3: scan in CI where no agent can skip it

Section titled “Layer 3: scan in CI where no agent can skip it”

This workflow runs on every pull request. It holds no secrets, uses a read-only token, and keeps the token out of .git/config. Save it as .github/workflows/security-gates.yml:

name: Security gates
on:
pull_request:
push:
branches: [main]
permissions:
contents: read
jobs:
secrets:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
persist-credentials: false
- uses: trufflesecurity/trufflehog@main # pin to a release commit SHA
with:
extra_args: --results=verified,unknown
sast:
if: github.event_name == 'pull_request'
runs-on: ubuntu-latest
container: semgrep/semgrep # pin a version tag or a sha256 digest
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
persist-credentials: false
- name: Semgrep, new findings only
env:
BASE_SHA: ${{ github.event.pull_request.base.sha }}
run: semgrep scan --config p/default --error --metrics=off --baseline-commit "$BASE_SHA"

The TruffleHog action passes --fail itself, so a verified live secret turns the job red with exit code 183. --results=verified,unknown drops unverified candidates to keep noise down. Semgrep’s --error exits 1 on findings, and --baseline-commit reports only those the pull request introduced. To run Gitleaks in CI as well, gitleaks git --redact --ignore-gitleaks-allow --log-opts="$BASE_SHA..HEAD" scans the pull request’s commit range and ignores #gitleaks:allow comments, so it also catches a suppressed key that TruffleHog cannot verify, such as a revoked or test-format key.

Then mark secrets and sast as required status checks; a failing job that is not required blocks nothing.

Layer 4: security review on the pull request

Section titled “Layer 4: security review on the pull request”

anthropics/claude-code-security-review runs Claude Code over the pull request’s diff and posts findings as review comments. Its README says it “is not hardened against prompt injection attacks and should only be used to review trusted PRs”. The workflow below, adapted from that README, adds four protections: it skips pull requests from forks, keeps the token out of .git/config, removes the pull request’s own agent configuration before Claude Code starts, and sets the model explicitly:

name: Security review
on:
pull_request:
permissions:
contents: read
pull-requests: write
jobs:
security:
if: github.event.pull_request.head.repo.full_name == github.repository
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
ref: ${{ github.event.pull_request.head.sha }}
fetch-depth: 2
persist-credentials: false
- name: Drop the pull request's agent config so the reviewer loads none of it
run: rm -rf .claude .mcp.json CLAUDE.md CLAUDE.local.md
- uses: anthropics/claude-code-security-review@main # pin to a commit SHA
with:
comment-pr: true
claude-api-key: ${{ secrets.CLAUDE_API_KEY }}
claude-model: claude-opus-5-5

The rm -rf step matters because the action runs the claude CLI in the checked-out tree while the API key is in the environment as ANTHROPIC_API_KEY, so hooks or MCP servers committed in the pull request would run with it. The action reads the diff from the GitHub API, so the review still sees changes to those files. Set claude-model: the action’s code falls back to claude-opus-4-1-20250805, and Anthropic retired Claude Opus 4.1 on the Claude API in 2026. Model choice and prices are on the models hub.

To gate on the action’s findings-count output, add a step that fails when it is above zero and mark the job required; otherwise keep it advisory. run-every-commit: true reviews every push instead of once per pull request, at a higher cost and, per its README, with more false positives. Denial of service, rate limiting, generic input validation and open redirects are excluded by default; add what your system cares about in the file named by custom-security-scan-instructions.

Scan the agent’s own tools with Snyk Agent Scan

Section titled “Scan the agent’s own tools with Snyk Agent Scan”

The agent itself is attack surface. An MCP server’s tool description or a skill’s SKILL.md can carry instructions the agent obeys. Snyk Agent Scan (formerly mcp-scan; PyPI snyk-agent-scan 0.6.4 on 2026-09-26) finds the agent harnesses, MCP servers and skills on a machine and scans them for prompt injection, tool poisoning and toxic flows. It needs a Snyk API token and is Python only:

Terminal window
# Terminal, with SNYK_TOKEN exported from your secret store
uvx snyk-agent-scan@0.6.4 ~/.claude/skills
uvx snyk-agent-scan@0.6.4 ./vendor-skill/SKILL.md
uvx snyk-agent-scan@0.6.4 --ci

Run it before you install a marketplace skill or plugin, and on a schedule across the team’s machines. --dangerously-run-mcp-servers skips the per-server consent prompt and starts every configured stdio server, so use it only where you have checked every command. The wider procedure is in MCP security and skill security.

If your organization already gates on SonarQube, its official MCP server (the Docker image sonarsource/sonarqube-mcp) lets the agent read the quality gate and fix new issues in the session; the setup for Claude Code, Codex and Cursor is in the server README.

How much context do the security tools cost?

Section titled “How much context do the security tools cost?”

The security-guidance plugin runs as hooks, so it adds nothing to the context window until a finding comes back, but each end-of-turn and commit review costs model tokens. The Semgrep plugin and the Snyk and SonarQube servers add MCP tool definitions, which Claude Code and Codex load through tool search. Measure instead of guessing: claude plugin details semgrep prints the plugin’s components and projected token cost, and /context before and after shows the difference. Pre-commit, CI and the pull-request review cost no session context, which is one more reason to put the mandatory gates there.

A gate you have never seen fire is an assumption. Prove each layer with a drill, then watch four numbers.

Run a canary drill each quarter. On a throwaway branch, have the agent commit a fake credential in a format the scanners detect, and add one known-vulnerable pattern, such as a shell command built from request input. Record which layer stopped each. If CI stops something pre-commit should have, a developer’s hook is missing; if nothing stops it, you found the gap before an attacker did. For a live-format test key, use one your provider issues for testing and revoke it after the drill.

MetricDefinitionAct when
Layer of first catchFor each finding, the earliest layer that caught itFindings move later (CI instead of session); check that hooks and pre-commit are installed on every machine
Escaped secretsSecrets found in main history or reported by GitHub secret scanning after mergeAny non-zero value: rotate, then add the missed format to the Gitleaks config
Suppressions addedNew #gitleaks:allow, nosemgrep, .gitleaksignore or .semgrepignore entries per monthSuppressions grow faster than findings; review who added them and why
Security-review precisionSecurity-review comments that led to a code change, divided by comments postedPrecision falls; tighten custom-security-scan-instructions or the false-positive file

Who signs off. CI owns the gate: the required checks decide whether a pull request can merge. The agent owns fixing findings. A named security owner in CODEOWNERS approves any change to .gitleaksignore, .gitleaks.toml, .semgrepignore, .pre-commit-config.yaml, the workflow files and .claude/, because those are the files that switch the gates off. The same owner reads every pull request on your escalation list (authentication, payments, data access), whatever the scanners said. The escalation list lives in reviewing an agent’s pull request.

What breaks when agents meet security gates?

Section titled “What breaks when agents meet security gates?”

The agent skips the hook with --no-verify. Recovery: guard-commit.sh in Claude Code and the required CI jobs everywhere. A regex hook is a typo-catcher; a command wrapped in a script gets past it. CI is the boundary.

The agent suppresses instead of fixing with #gitleaks:allow, nosemgrep or a .gitleaksignore line. Recovery: CODEOWNERS on the ignore files, the “suppressions added” metric, and a CI Gitleaks run with --ignore-gitleaks-allow (Layer 3).

The secret reached the model even though the commit was blocked. The agent read .env for context, so the value sits in the transcript and possibly in telemetry. Recovery: rotate, then deny reads of .env* in the agent’s permissions; the deny rules and sandboxing options are in permissions and sandboxing.

Semgrep fails every pull request on an old codebase, and people start merging with admin rights. Recovery: --baseline-commit, then burn down the backlog separately.

The pull-request reviewer follows instructions planted in the diff, such as a comment saying the file is pre-approved. Recovery: the fork condition and config-removal step above, with deterministic scanners as the required checks and the LLM review as advice.

The in-session review fights a parallel agent in the same worktree, reporting its changes to the wrong session. Recovery: ENABLE_STOP_REVIEW=0, or one worktree per agent, as in running agents in parallel.

A hallucinated package slips past every scanner. Secret and code scanners do not check that a dependency is real and trusted. Recovery: the lockfile-diff gate in dependency checks for agent changes.