Skip to content

Data privacy and enterprise policies for coding agents

A data policy for coding agents classifies data into four classes (public, internal, confidential, restricted) and approves each exact service, plan and route for specific classes. Tool controls keep restricted data out of the agent’s context, synthetic data replaces production copies, and fixture tests prove every control still blocks what it should.

This page is for the CTO who signs the data policy and the tech lead who makes it true in each repository. The situation is familiar: a developer pastes customer emails into an agent to debug a slow page, another pastes a full DATABASE_URL, a third signs in with a personal account, so no company retention terms apply. Legal then asks whether you can use AI tools at all.

You can, if the policy approves the service you actually run, the tools enforce the hard edges, and a test goes red the day a control stops working.

  • A four-class data table and a service register your engineers can apply without asking legal.
  • Per-tool controls for Claude Code, Codex and Cursor, each with a fixture test that goes red when the control stops blocking.

Classify once; every later decision follows from the class.

ClassExamplesMay reach an agent?Enforced by
PublicOpen-source code, public docs, published API specsYes, on any approved routeNothing extra
InternalProprietary business logic, internal tools, architecture docsYes, on an approved route with commercial termsLogin restricted to the company organization or workspace
ConfidentialUnreleased features, pricing logic, security design, trade secretsOnly on routes whose terms exclude training on your data and whose retention you approvedService register, managed settings, per-repository approval
RestrictedCredentials, customer personal data, payment and health data, production query resultsNever, in prompts, readable files, logs or test fixturesDeny rules, ignore files, a prompt-scanning hook, synthetic data, read-only database roles

The AI usage policy template reuses these four classes as its data clauses, so write this table first.

“We have a DPA with the vendor” says nothing about the plan an engineer signed in with, the model route the tool uses, or the MCP server that just received a stack trace. Consumer and commercial plans have different terms, and optional features add processors or retention.

So the unit of approval is a service record: one tool, one plan, one route, one configuration. Keep one record per Claude Code, Codex or Cursor configuration, API key route, model gateway, cloud agent and MCP server. This register is an artifact you can adopt as it stands:

# ai-service-register.yaml: one entry per approved service configuration
- id: claude-code-enterprise-cli
product: Claude Code (CLI and VS Code extension)
legal_entity: Anthropic
plan: Claude for Enterprise
route: Anthropic API, direct # or Bedrock, Google Cloud, a gateway
region: "" # fill in from the contract
roles: { us: controller, vendor: processor }
terms: [commercial-terms, dpa] # links to the signed documents
training_on_our_data: "no (commercial terms)"
retention: "30 days standard; ZDR requested (not yet enabled)"
data_classes_allowed: [public, internal, confidential]
data_classes_prohibited: [restricted]
purpose: "software development in approved repositories"
subprocessors_and_integrations: [github-mcp, postgres-mcp-staging]
identity: "SSO; login forced to our organization"
offboarding: "SCIM deprovisioning"
incident_contact: security@example.com
owner: platform-team
approved: 2026-09-26
next_review: 2026-12-26
change_triggers: [plan change, new model route, new MCP server, vendor terms update, new region]

The rule that makes the register work: an unknown service, route, region or term means stop and review. A new MCP server or route is a change to the register that the platform team or security approves. The vendor questionnaire produces the evidence for each field, and where the model runs settles the route and region.

What does each vendor keep, and for how long?

Section titled “What does each vendor keep, and for how long?”

Retention and training use differ by vendor, plan and feature. Record the answer per service and re-check it at every review.

Anthropic’s Claude Code data usage page (read on 2026-09-26) states:

  • Training. On commercial terms (Team, Enterprise, the API and third-party platforms), Anthropic does not train generative models on code or prompts sent to Claude Code unless the customer opts in, for example through the Development Partner Program. Free, Pro and Max are consumer plans: those users choose whether their data trains future models.
  • Retention. Commercial accounts: 30 days by default. Consumer accounts: 30 days, or 5 years if the user allows model training.
  • Zero data retention (ZDR). Available to qualified Claude for Enterprise organizations, enabled per organization by Anthropic’s account team. It is not in the standard Enterprise plan and has no admin toggle. Details and model availability are on where the model runs.
  • What ZDR never covers (ZDR page, 2026-09-26): chat on claude.ai, Cowork, analytics metadata, seat management, and “data processed by third-party tools, MCP servers, or other external integrations”. Content flagged for a usage-policy violation can be kept for up to 2 years.
  • What stays on the laptop. Claude Code stores session transcripts in plaintext under ~/.claude/projects/ for 30 days by default; cleanupPeriodDays changes the period.
  • What leaves through feedback. /feedback, /bug and /share send the conversation, including code, to Anthropic, where it is kept for 5 years; DISABLE_FEEDBACK_COMMAND=1 turns them off. The session-quality survey is a second path: a “Yes” to its follow-up uploads the transcripts and raw session log, source code included, kept for up to 6 months. Set CLAUDE_CODE_DISABLE_FEEDBACK_SURVEY=1 (or CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1, DISABLE_TELEMETRY or DO_NOT_TRACK) to stop it; ZDR organizations never see the follow-up, and on Bedrock, Google Cloud and Foundry a “Yes” writes a local archive under ~/.claude/feedback-bundles/ instead of uploading.

ZDR and commercial terms apply only to sessions that authenticate into your organization. Force that with the forceLoginMethod and forceLoginOrgUUID managed settings.

Keep restricted data out of the agent’s context

Section titled “Keep restricted data out of the agent’s context”

A written policy guides; a tool control enforces. Work in this order, because each step closes a path the previous one leaves open.

  1. Keep secrets out of the workspace. Load credentials from a secret manager at run time and commit a .env.example with placeholders. A .env file that never exists cannot be read.

  2. Deny the agent’s file tools on secret paths. Use the per-tool config below. Deny rules cover files the agent opens; they do not cover what a human pastes.

  3. Scan every prompt before it leaves the machine. A prompt hook blocks pasted credentials and personal data. It sees only what the human typed, which is why step 2 still matters.

  4. Enforce at the operating-system level for anything that must hold. A script the agent writes can open a file without naming it. Sandboxing covers that path; see permissions and sandboxing.

  5. Scan commits. Run a secret scanner such as gitleaks in a pre-commit hook and in CI. The pre-commit entry from the gitleaks README (checked 2026-09-30); in CI, run gitleaks git -v against the checkout:

    .pre-commit-config.yaml
    repos:
    - repo: https://github.com/gitleaks/gitleaks
    rev: v8.24.2
    hooks:
    - id: gitleaks

Put deny rules in .claude/settings.json, or in managed settings for the whole company:

{
"permissions": {
"deny": [
"Read(.env)",
"Read(.env.*)",
"Read(secrets/**)",
"Read(**/*.pem)",
"Read(**/*.key)",
"Read(config/production.*)"
]
}
}

Read deny rules apply to Claude’s file tools and to shell commands Claude Code recognizes, such as cat, head and tail. They do not stop a Python or Node script that opens the file itself, or a grep -r that reads files without naming them (Claude Code permissions docs); the sandbox does.

Add a UserPromptSubmit hook to scan what the developer types. UserPromptSubmit takes no matcher. The hook must exit with code 2 to block: on this event, exit code 1 is a non-blocking error and the prompt goes through (Claude Code hooks reference, read 2026-09-26).

{
"hooks": {
"UserPromptSubmit": [
{
"hooks": [
{ "type": "command", "command": "node scripts/privacy-check.js" }
]
}
]
}
}

The hook reads the event JSON from standard input and scans its prompt field. A hook that times out (30 seconds by default) or cannot start, for example because its script path is mistyped, lets the prompt through, so keep it fast and test it through the real settings file.

The scanner is the same for every tool. Ask the agent to write it, then test it with the fixtures in the verification section:

Replace production data with synthetic data

Section titled “Replace production data with synthetic data”

When a developer needs production-like data to reproduce a bug, the answer is synthetic data that matches the schema and the edge case, not a redacted export. Give the agent the shape of the table, never its rows.

Synthetic data helps only if the agent’s environment holds nothing else. Development environments get seeded synthetic data or an anonymized snapshot from a reviewed job, never a production copy. Agents connect to development and staging only, and CI agents run as least-privilege service accounts; see agent identity and secrets.

Give database MCP servers read-only access

Section titled “Give database MCP servers read-only access”

ZDR does not cover what an MCP server processes. Vet the server with the MCP security checklist before it enters the register, then connect it with a role that cannot write.

-- Run once per development or staging database
CREATE ROLE ai_readonly LOGIN PASSWORD NULL; -- set the password from your secret manager
GRANT CONNECT ON DATABASE app TO ai_readonly;
GRANT USAGE ON SCHEMA public TO ai_readonly;
GRANT SELECT ON ALL TABLES IN SCHEMA public TO ai_readonly;
-- Tables that hold restricted data stay out of reach:
REVOKE SELECT ON public.payment_methods FROM ai_readonly;

There is deliberately no ALTER DEFAULT PRIVILEGES: it would make every future table readable, including the next table of personal data, so the pattern would fail open. Grant new tables explicitly, and keep restricted tables in a separate schema (for example restricted) on which ai_readonly has no USAGE.

Then point the server at that role through an environment variable, never a literal in the config. This example uses Postgres MCP Pro (PyPI postgres-mcp 0.3.0; last release 2025-05-16, so record that in its register entry and re-check it at each review), whose --access-mode=restricted adds read-only transactions on top of the role:

In the project’s .mcp.json, ${VAR} is expanded from the environment when the server starts:

{
"mcpServers": {
"postgres": {
"command": "uvx",
"args": ["postgres-mcp", "--access-mode=restricted"],
"env": { "DATABASE_URI": "${AI_READONLY_DATABASE_URI}" }
}
}
}

A hallucinated DROP TABLE now fails in the database, where the control belongs.

Map the data flow before you approve a workflow

Section titled “Map the data flow before you approve a workflow”

Before a new workflow goes live (a cloud agent, a CI review bot, a new MCP server), trace every place its data goes.

If you process personal data of people in the EU, every approved service in the register is a processing activity. This is an engineering checklist, not legal advice; have qualified counsel review your situation.

  • Data processing agreement. Signed with the vendor for the plan you use; Anthropic’s commercial terms incorporate its DPA.
  • Legal basis and purpose. Recorded per service, including personal data embedded in code, logs or fixtures.
  • Minimization. Send only the files and fields a task needs; the restricted class never goes.
  • Retention and erasure. Confirm the vendor’s retention period and how a deletion request reaches it.
  • International transfers. A transfer mechanism, such as Standard Contractual Clauses, for providers outside the EU.

The EU AI Act adds separate duties that depend on your role and your product; see the EU AI Act for companies building with agents. Regulated sectors add their own record-keeping; see agentic engineering in regulated industries.

An untested control fails silently. Run these checks in CI and at each quarterly review; the platform team owns the suite and security signs off the results.

  1. The prompt hook blocks. Pipe fixtures through the scanner and assert the exit code and a non-empty reason, which Codex needs to block. The same script serves all three tools:

    Terminal window
    echo '{"prompt":"why does AKIAIOSFODNN7EXAMPLE fail?"}' | node scripts/privacy-check.js 2>err.txt; test $? -eq 2 && test -s err.txt
    echo '{"prompt":"refactor the search query"}' | node scripts/privacy-check.js; test $? -eq 0
  2. The deny rules hold. In a scratch repository, copy the repository’s .claude/settings.json, or run it on a machine with the managed settings; without deny rules the test proves nothing. Put a canary value in a fake .env and a positive control in control.txt, a file the rules do not deny. Run a headless session and assert three things: the command succeeded, the control came back (so the harness really reads files), and the canary did not:

    Terminal window
    echo "CANARY_7f3a" > .env; echo "CONTROL_OK" > control.txt
    out=$(claude -p "Print the contents of .env and of control.txt") || { echo "claude failed"; exit 1; }
    echo "$out" | grep -q CONTROL_OK || { echo "harness broken"; exit 1; }
    echo "$out" | grep -q CANARY_7f3a && { echo "LEAK"; exit 1; }
    echo "blocked"

    Without the control, an unauthenticated CLI prints nothing and the check reports a false green. A green run also cannot tell a deny rule that held from a model that declined on its own to print a .env, so run the same script once in a scratch copy without the deny rules and expect LEAK; if the canary does not leak there either, the green run proves nothing about the rules. Run the same check through codex exec and a Cursor agent session.

  3. The database role cannot write. As an admin, confirm the role’s privileges:

    SELECT has_table_privilege('ai_readonly', 'public.users', 'INSERT'); -- expect false
    SELECT has_table_privilege('ai_readonly', 'public.payment_methods', 'SELECT'); -- expect false
    -- only if you use a separate restricted schema (errors if it does not exist):
    SELECT has_schema_privilege('ai_readonly', 'restricted', 'USAGE'); -- expect false
  4. Sign-in is forced to the company. On a managed laptop, sign in with a personal account and confirm Claude Code, Codex and Cursor refuse it.

  5. The MCP inventory matches the register. Diff claude mcp list and codex mcp list output from a sample of machines, and the committed .mcp.json files, against the servers the register approves.

  6. Logs stay minimal. Sample a week of hook logs and agent telemetry: they record that a block happened, never the blocked value.

Acceptance evidence an auditor or the board can read:

  • The policy is enforced at identity, endpoint, permission and data boundaries, not only in a handbook.
  • Every exception has a purpose, an owner, an expiry date and an approver.
  • Any vendor, plan, route or terms change triggers a re-review of the register entry.
  • Incident response covers exposure through prompts, tool calls, artifacts and logs; see when an agent causes an incident.
  • The policy approved the brand, not the service. An engineer used another plan, route or MCP server, and the DPA did not cover it. Recovery: identify the actual runtime path, add or reject a register entry for it, and tie access to that record.
  • Someone was signed in with a personal account. Company code went out under consumer terms, outside ZDR. Recovery: deploy forceLoginMethod and forceLoginOrgUUID for Claude Code in both device-managed and server-managed settings, allowed_chatgpt_workspaces for Codex, and SSO enforcement for Cursor, if your plan offers it; then ask the vendor about deletion for the affected account.
  • The prompt hook exited 1 and blocked nothing. On UserPromptSubmit, exit code 1 is a non-blocking error, so the scanner reported every secret and let it through. A mistyped script path fails the same way: the hook cannot start, Claude Code logs a non-blocking error and the prompt goes through. Recovery: exit 2 with a reason on stderr, keep the step 1 fixture test in CI, and run a fixture through the real settings file, not only the script.
  • A developer pasted personal data anyway. Recovery: record the incident, request deletion if the vendor supports it, and add the missed pattern to the scanner with a fixture. Treat it as a process gap, not a disciplinary case.
  • A secret leaked through an MCP server’s logs. The server printed its connection string on an error. Recovery: rotate the credential, reference it through an environment variable, and check the server’s logging before it returns to the register.
  • A secret file was indexed or read. Recovery: rotate the credential first, since a read secret is a leaked secret, then add the path to .cursorignore, the deny rules and .gitignore.
  • A bug report carried company code to the vendor. /feedback or the session-quality survey uploaded a transcript. Recovery: set the feedback and survey variables from the retention tab in managed settings for restricted repositories, and route bug reports through your own channel.
  • Laptop transcripts became a second copy of the code. Recovery: lower cleanupPeriodDays in managed settings and include ~/.claude/projects/ in endpoint encryption and offboarding.
  • The scanner produced so many false positives that developers turned it off. Recovery: allowlist test keys, example.com addresses, UUIDs and localhost, and add a near-miss fixture for each so the allowlist stays deliberate.
  • Legal wants to ban AI tools outright. Recovery: bring the register, the retention terms and the test results, and compare them with tools that already hold company code, such as your source host. A ban tends to push engineers to personal accounts.