Data privacy and enterprise policies for coding agents
A data policy for coding agents classifies data into four classes (public, internal, confidential, restricted) and approves each exact service, plan and route for specific classes. Tool controls keep restricted data out of the agent’s context, synthetic data replaces production copies, and fixture tests prove every control still blocks what it should.
This page is for the CTO who signs the data policy and the tech lead who makes it true in each repository. The situation is familiar: a developer pastes customer emails into an agent to debug a slow page, another pastes a full DATABASE_URL, a third signs in with a personal account, so no company retention terms apply. Legal then asks whether you can use AI tools at all.
You can, if the policy approves the service you actually run, the tools enforce the hard edges, and a test goes red the day a control stops working.
What this data policy gives you
Section titled “What this data policy gives you”- A four-class data table and a service register your engineers can apply without asking legal.
- Per-tool controls for Claude Code, Codex and Cursor, each with a fixture test that goes red when the control stops blocking.
Which data may reach a coding agent?
Section titled “Which data may reach a coding agent?”Classify once; every later decision follows from the class.
| Class | Examples | May reach an agent? | Enforced by |
|---|---|---|---|
| Public | Open-source code, public docs, published API specs | Yes, on any approved route | Nothing extra |
| Internal | Proprietary business logic, internal tools, architecture docs | Yes, on an approved route with commercial terms | Login restricted to the company organization or workspace |
| Confidential | Unreleased features, pricing logic, security design, trade secrets | Only on routes whose terms exclude training on your data and whose retention you approved | Service register, managed settings, per-repository approval |
| Restricted | Credentials, customer personal data, payment and health data, production query results | Never, in prompts, readable files, logs or test fixtures | Deny rules, ignore files, a prompt-scanning hook, synthetic data, read-only database roles |
The AI usage policy template reuses these four classes as its data clauses, so write this table first.
Approve the exact service, not the vendor
Section titled “Approve the exact service, not the vendor”“We have a DPA with the vendor” says nothing about the plan an engineer signed in with, the model route the tool uses, or the MCP server that just received a stack trace. Consumer and commercial plans have different terms, and optional features add processors or retention.
So the unit of approval is a service record: one tool, one plan, one route, one configuration. Keep one record per Claude Code, Codex or Cursor configuration, API key route, model gateway, cloud agent and MCP server. This register is an artifact you can adopt as it stands:
# ai-service-register.yaml: one entry per approved service configuration- id: claude-code-enterprise-cli product: Claude Code (CLI and VS Code extension) legal_entity: Anthropic plan: Claude for Enterprise route: Anthropic API, direct # or Bedrock, Google Cloud, a gateway region: "" # fill in from the contract roles: { us: controller, vendor: processor } terms: [commercial-terms, dpa] # links to the signed documents training_on_our_data: "no (commercial terms)" retention: "30 days standard; ZDR requested (not yet enabled)" data_classes_allowed: [public, internal, confidential] data_classes_prohibited: [restricted] purpose: "software development in approved repositories" subprocessors_and_integrations: [github-mcp, postgres-mcp-staging] identity: "SSO; login forced to our organization" offboarding: "SCIM deprovisioning" incident_contact: security@example.com owner: platform-team approved: 2026-09-26 next_review: 2026-12-26 change_triggers: [plan change, new model route, new MCP server, vendor terms update, new region]The rule that makes the register work: an unknown service, route, region or term means stop and review. A new MCP server or route is a change to the register that the platform team or security approves. The vendor questionnaire produces the evidence for each field, and where the model runs settles the route and region.
What does each vendor keep, and for how long?
Section titled “What does each vendor keep, and for how long?”Retention and training use differ by vendor, plan and feature. Record the answer per service and re-check it at every review.
Anthropic’s Claude Code data usage page (read on 2026-09-26) states:
- Training. On commercial terms (Team, Enterprise, the API and third-party platforms), Anthropic does not train generative models on code or prompts sent to Claude Code unless the customer opts in, for example through the Development Partner Program. Free, Pro and Max are consumer plans: those users choose whether their data trains future models.
- Retention. Commercial accounts: 30 days by default. Consumer accounts: 30 days, or 5 years if the user allows model training.
- Zero data retention (ZDR). Available to qualified Claude for Enterprise organizations, enabled per organization by Anthropic’s account team. It is not in the standard Enterprise plan and has no admin toggle. Details and model availability are on where the model runs.
- What ZDR never covers (ZDR page, 2026-09-26): chat on claude.ai, Cowork, analytics metadata, seat management, and “data processed by third-party tools, MCP servers, or other external integrations”. Content flagged for a usage-policy violation can be kept for up to 2 years.
- What stays on the laptop. Claude Code stores session transcripts in plaintext under
~/.claude/projects/for 30 days by default;cleanupPeriodDayschanges the period. - What leaves through feedback.
/feedback,/bugand/sharesend the conversation, including code, to Anthropic, where it is kept for 5 years;DISABLE_FEEDBACK_COMMAND=1turns them off. The session-quality survey is a second path: a “Yes” to its follow-up uploads the transcripts and raw session log, source code included, kept for up to 6 months. SetCLAUDE_CODE_DISABLE_FEEDBACK_SURVEY=1(orCLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1,DISABLE_TELEMETRYorDO_NOT_TRACK) to stop it; ZDR organizations never see the follow-up, and on Bedrock, Google Cloud and Foundry a “Yes” writes a local archive under~/.claude/feedback-bundles/instead of uploading.
ZDR and commercial terms apply only to sessions that authenticate into your organization. Force that with the forceLoginMethod and forceLoginOrgUUID managed settings.
OpenAI publishes its business data commitments on openai.com, which could not be read from the writing environment on 2026-09-26. Get the retention, training and residency terms for your ChatGPT plan in writing.
The open-source Codex CLI (0.157.1, config_requirements.rs) shows the enforcement: an administrator’s requirements.toml pins sign-in to company workspaces and traffic to a residency region:
# requirements.toml (admin-managed constraints; users cannot override)allowed_login_methods = ["chatgpt"]allowed_chatgpt_workspaces = ["YOUR_WORKSPACE_ID"]enforce_residency = "us" # the only value the 0.157.1 source acceptsReplace YOUR_WORKSPACE_ID with your ChatGPT workspace ID. Drop enforce_residency if your contract does not include US residency. Enforcing one policy across every agent covers how to distribute the file.
Cursor’s retention and training terms could not be re-verified on 2026-09-26 (cursor.com was unreachable from the writing environment). From Cursor’s current privacy pages for your plan, record what Privacy Mode guarantees, which features require retention, and how long Cloud Agent data is kept.
What is confirmed from Cursor’s own SDK (@cursor/sdk 1.0.32, 2026-09-22): Privacy Mode exists and can be forced on by the team, and Cursor honours a .cursorignore file. Force Privacy Mode at the team level. Whether you can force sign-in to your team (SSO enforcement) is not confirmed from the SDK; record the answer from Cursor’s enterprise docs, because a personal account bypasses team settings.
Keep restricted data out of the agent’s context
Section titled “Keep restricted data out of the agent’s context”A written policy guides; a tool control enforces. Work in this order, because each step closes a path the previous one leaves open.
-
Keep secrets out of the workspace. Load credentials from a secret manager at run time and commit a
.env.examplewith placeholders. A.envfile that never exists cannot be read. -
Deny the agent’s file tools on secret paths. Use the per-tool config below. Deny rules cover files the agent opens; they do not cover what a human pastes.
-
Scan every prompt before it leaves the machine. A prompt hook blocks pasted credentials and personal data. It sees only what the human typed, which is why step 2 still matters.
-
Enforce at the operating-system level for anything that must hold. A script the agent writes can open a file without naming it. Sandboxing covers that path; see permissions and sandboxing.
-
Scan commits. Run a secret scanner such as gitleaks in a pre-commit hook and in CI. The pre-commit entry from the gitleaks README (checked 2026-09-30); in CI, run
gitleaks git -vagainst the checkout:.pre-commit-config.yaml repos:- repo: https://github.com/gitleaks/gitleaksrev: v8.24.2hooks:- id: gitleaks
Put deny rules in .claude/settings.json, or in managed settings for the whole company:
{ "permissions": { "deny": [ "Read(.env)", "Read(.env.*)", "Read(secrets/**)", "Read(**/*.pem)", "Read(**/*.key)", "Read(config/production.*)" ] }}Read deny rules apply to Claude’s file tools and to shell commands Claude Code recognizes, such as cat, head and tail. They do not stop a Python or Node script that opens the file itself, or a grep -r that reads files without naming them (Claude Code permissions docs); the sandbox does.
Add a UserPromptSubmit hook to scan what the developer types. UserPromptSubmit takes no matcher. The hook must exit with code 2 to block: on this event, exit code 1 is a non-blocking error and the prompt goes through (Claude Code hooks reference, read 2026-09-26).
{ "hooks": { "UserPromptSubmit": [ { "hooks": [ { "type": "command", "command": "node scripts/privacy-check.js" } ] } ] }}The hook reads the event JSON from standard input and scans its prompt field. A hook that times out (30 seconds by default) or cannot start, for example because its script path is mistyped, lets the prompt through, so keep it fast and test it through the real settings file.
Codex reads AGENTS.md before every task. State the policy there:
## Data handling- Do not open .env files, secrets/, *.pem or *.key. Use .env.example.- Never print credentials, tokens or personal data in code, comments or output.- For debugging, generate synthetic data from the schema; never ask for production rows.- Connection strings come from environment variables, never literals.AGENTS.md is guidance. For enforcement, keep secret files out of the checkout and stop the shell tool from inheriting your environment’s credentials. In ~/.codex/config.toml:
[shell_environment_policy]inherit = "core" # only core variables such as HOME, PATH and USER reach commandsCodex 0.157.1 also runs UserPromptSubmit hooks (the hooks feature is stable). An exit code of 2 with a reason on stderr blocks the prompt, so the same scripts/privacy-check.js works unchanged. Codex reads hooks.json from a config folder such as .codex/ in a trusted project, in the same JSON shape as Claude Code (hooks/src/engine/discovery.rs at rust-v0.157.1):
{ "hooks": { "UserPromptSubmit": [ { "hooks": [ { "type": "command", "command": "node scripts/privacy-check.js" } ] } ] }}Two traps. On Codex, exit 2 with empty stderr does not block: the hook is recorded as failed and the prompt goes through, so the scanner must always print a reason. And project hooks do not run until you trust them in /hooks, again after every change to hooks.json. Administrators can require managed hooks and constrain sandbox settings in requirements.toml; see enforcing one policy across every agent.
Commit a .cursorignore so sensitive files are neither indexed nor sent as context:
.env.env.***/secrets/****/*.pem**/*.keyconfig/production.*database/seeds/production/**State the same policy as a project rule under .cursor/rules/. Cursor also runs hooks, spawned processes that “can observe, block, or modify behavior” (Cursor hooks documentation, 2026-08-28). A hook can call the same scripts/privacy-check.js scanner; check Cursor’s hooks page for the event that fires before a prompt is sent and for how that event signals a block, because neither could be re-verified on 2026-09-26.
The scanner is the same for every tool. Ask the agent to write it, then test it with the fixtures in the verification section:
Replace production data with synthetic data
Section titled “Replace production data with synthetic data”When a developer needs production-like data to reproduce a bug, the answer is synthetic data that matches the schema and the edge case, not a redacted export. Give the agent the shape of the table, never its rows.
Synthetic data helps only if the agent’s environment holds nothing else. Development environments get seeded synthetic data or an anonymized snapshot from a reviewed job, never a production copy. Agents connect to development and staging only, and CI agents run as least-privilege service accounts; see agent identity and secrets.
Give database MCP servers read-only access
Section titled “Give database MCP servers read-only access”ZDR does not cover what an MCP server processes. Vet the server with the MCP security checklist before it enters the register, then connect it with a role that cannot write.
-- Run once per development or staging databaseCREATE ROLE ai_readonly LOGIN PASSWORD NULL; -- set the password from your secret managerGRANT CONNECT ON DATABASE app TO ai_readonly;GRANT USAGE ON SCHEMA public TO ai_readonly;GRANT SELECT ON ALL TABLES IN SCHEMA public TO ai_readonly;-- Tables that hold restricted data stay out of reach:REVOKE SELECT ON public.payment_methods FROM ai_readonly;There is deliberately no ALTER DEFAULT PRIVILEGES: it would make every future table readable, including the next table of personal data, so the pattern would fail open. Grant new tables explicitly, and keep restricted tables in a separate schema (for example restricted) on which ai_readonly has no USAGE.
Then point the server at that role through an environment variable, never a literal in the config. This example uses Postgres MCP Pro (PyPI postgres-mcp 0.3.0; last release 2025-05-16, so record that in its register entry and re-check it at each review), whose --access-mode=restricted adds read-only transactions on top of the role:
In the project’s .mcp.json, ${VAR} is expanded from the environment when the server starts:
{ "mcpServers": { "postgres": { "command": "uvx", "args": ["postgres-mcp", "--access-mode=restricted"], "env": { "DATABASE_URI": "${AI_READONLY_DATABASE_URI}" } } }}In ~/.codex/config.toml, env_vars passes the named variable from your environment to the server:
[mcp_servers.postgres]command = "uvx"args = ["postgres-mcp", "--access-mode=restricted"]env_vars = ["DATABASE_URI"]Export DATABASE_URI from your secret manager with the ai_readonly credentials before you start Codex.
Cursor reads the same mcpServers shape from .cursor/mcp.json. Use the Claude Code block, and confirm in Cursor’s MCP documentation that your version expands environment variables in env; if it does not, start Cursor from a shell where DATABASE_URI is already set and omit the env key.
A hallucinated DROP TABLE now fails in the database, where the control belongs.
Map the data flow before you approve a workflow
Section titled “Map the data flow before you approve a workflow”Before a new workflow goes live (a cloud agent, a CI review bot, a new MCP server), trace every place its data goes.
What GDPR asks of agent data flows
Section titled “What GDPR asks of agent data flows”If you process personal data of people in the EU, every approved service in the register is a processing activity. This is an engineering checklist, not legal advice; have qualified counsel review your situation.
- Data processing agreement. Signed with the vendor for the plan you use; Anthropic’s commercial terms incorporate its DPA.
- Legal basis and purpose. Recorded per service, including personal data embedded in code, logs or fixtures.
- Minimization. Send only the files and fields a task needs; the restricted class never goes.
- Retention and erasure. Confirm the vendor’s retention period and how a deletion request reaches it.
- International transfers. A transfer mechanism, such as Standard Contractual Clauses, for providers outside the EU.
The EU AI Act adds separate duties that depend on your role and your product; see the EU AI Act for companies building with agents. Regulated sectors add their own record-keeping; see agentic engineering in regulated industries.
How do you prove the data controls hold?
Section titled “How do you prove the data controls hold?”An untested control fails silently. Run these checks in CI and at each quarterly review; the platform team owns the suite and security signs off the results.
-
The prompt hook blocks. Pipe fixtures through the scanner and assert the exit code and a non-empty reason, which Codex needs to block. The same script serves all three tools:
Terminal window echo '{"prompt":"why does AKIAIOSFODNN7EXAMPLE fail?"}' | node scripts/privacy-check.js 2>err.txt; test $? -eq 2 && test -s err.txtecho '{"prompt":"refactor the search query"}' | node scripts/privacy-check.js; test $? -eq 0 -
The deny rules hold. In a scratch repository, copy the repository’s
.claude/settings.json, or run it on a machine with the managed settings; without deny rules the test proves nothing. Put a canary value in a fake.envand a positive control incontrol.txt, a file the rules do not deny. Run a headless session and assert three things: the command succeeded, the control came back (so the harness really reads files), and the canary did not:Terminal window echo "CANARY_7f3a" > .env; echo "CONTROL_OK" > control.txtout=$(claude -p "Print the contents of .env and of control.txt") || { echo "claude failed"; exit 1; }echo "$out" | grep -q CONTROL_OK || { echo "harness broken"; exit 1; }echo "$out" | grep -q CANARY_7f3a && { echo "LEAK"; exit 1; }echo "blocked"Without the control, an unauthenticated CLI prints nothing and the check reports a false green. A green run also cannot tell a deny rule that held from a model that declined on its own to print a
.env, so run the same script once in a scratch copy without the deny rules and expectLEAK; if the canary does not leak there either, the green run proves nothing about the rules. Run the same check throughcodex execand a Cursor agent session. -
The database role cannot write. As an admin, confirm the role’s privileges:
SELECT has_table_privilege('ai_readonly', 'public.users', 'INSERT'); -- expect falseSELECT has_table_privilege('ai_readonly', 'public.payment_methods', 'SELECT'); -- expect false-- only if you use a separate restricted schema (errors if it does not exist):SELECT has_schema_privilege('ai_readonly', 'restricted', 'USAGE'); -- expect false -
Sign-in is forced to the company. On a managed laptop, sign in with a personal account and confirm Claude Code, Codex and Cursor refuse it.
-
The MCP inventory matches the register. Diff
claude mcp listandcodex mcp listoutput from a sample of machines, and the committed.mcp.jsonfiles, against the servers the register approves. -
Logs stay minimal. Sample a week of hook logs and agent telemetry: they record that a block happened, never the blocked value.
Acceptance evidence an auditor or the board can read:
- The policy is enforced at identity, endpoint, permission and data boundaries, not only in a handbook.
- Every exception has a purpose, an owner, an expiry date and an approver.
- Any vendor, plan, route or terms change triggers a re-review of the register entry.
- Incident response covers exposure through prompts, tool calls, artifacts and logs; see when an agent causes an incident.
When data privacy controls break down
Section titled “When data privacy controls break down”- The policy approved the brand, not the service. An engineer used another plan, route or MCP server, and the DPA did not cover it. Recovery: identify the actual runtime path, add or reject a register entry for it, and tie access to that record.
- Someone was signed in with a personal account. Company code went out under consumer terms, outside ZDR. Recovery: deploy
forceLoginMethodandforceLoginOrgUUIDfor Claude Code in both device-managed and server-managed settings,allowed_chatgpt_workspacesfor Codex, and SSO enforcement for Cursor, if your plan offers it; then ask the vendor about deletion for the affected account. - The prompt hook exited 1 and blocked nothing. On
UserPromptSubmit, exit code 1 is a non-blocking error, so the scanner reported every secret and let it through. A mistyped script path fails the same way: the hook cannot start, Claude Code logs a non-blocking error and the prompt goes through. Recovery: exit 2 with a reason on stderr, keep the step 1 fixture test in CI, and run a fixture through the real settings file, not only the script. - A developer pasted personal data anyway. Recovery: record the incident, request deletion if the vendor supports it, and add the missed pattern to the scanner with a fixture. Treat it as a process gap, not a disciplinary case.
- A secret leaked through an MCP server’s logs. The server printed its connection string on an error. Recovery: rotate the credential, reference it through an environment variable, and check the server’s logging before it returns to the register.
- A secret file was indexed or read. Recovery: rotate the credential first, since a read secret is a leaked secret, then add the path to
.cursorignore, the deny rules and.gitignore. - A bug report carried company code to the vendor.
/feedbackor the session-quality survey uploaded a transcript. Recovery: set the feedback and survey variables from the retention tab in managed settings for restricted repositories, and route bug reports through your own channel. - Laptop transcripts became a second copy of the code. Recovery: lower
cleanupPeriodDaysin managed settings and include~/.claude/projects/in endpoint encryption and offboarding. - The scanner produced so many false positives that developers turned it off. Recovery: allowlist test keys,
example.comaddresses, UUIDs and localhost, and add a near-miss fixture for each so the allowlist stays deliberate. - Legal wants to ban AI tools outright. Recovery: bring the register, the retention terms and the test results, and compare them with tools that already hold company code, such as your source host. A ban tends to push engineers to personal accounts.