Skip to content

Where the model runs: gateways, cloud routes, residency and zero retention

Model hosting for coding agents is the choice of where inference runs: the vendor’s own API; a cloud route such as Amazon Bedrock, Google Cloud’s Agent Platform, formerly Vertex AI, Microsoft Foundry or Claude Platform on AWS; an LLM gateway in front of either; or self-hosted open weights. The route decides data residency, retention, billing, and which agent features keep working.

Legal has just signed a customer contract that says source code is processed in the EU and not retained by subprocessors. Your teams use Claude Code on a Team plan, Codex through ChatGPT, and Cursor. Procurement asks per tool, “where does our code go, and for how long?”, and nobody can answer with evidence.

This page is for the CTO who owns that answer. The route also decides who bills you, so read it with AI usage cost governance. Per-tool setup lives in LLM gateway configuration for Claude Code; model versions and prices live on the models hub.

What you decide when you choose where the model runs

Section titled “What you decide when you choose where the model runs”
  • The five routes, who operates inference on each, and which features each costs you.
  • Region controls per route, including where “global” routing leaves the region you assumed.
  • Which models run under zero data retention (ZDR), and what ZDR switches off or never covers.
  • Managed configuration per tool, a decision record, and a quarterly proof of the route.

Facts were checked on 2026-09-26 against Anthropic’s documentation, Claude Code 2.1.283 and Codex 0.157.1. AWS and Google Cloud documentation and openai.com were unreachable that day, so their retention terms appear as questions, not facts.

Which hosting routes exist for coding agents?

Section titled “Which hosting routes exist for coding agents?”

Five routes cover every deployment in practice. The first four are places inference runs; the gateway is a layer you put in front of any of them.

RouteWho operates inferenceBillingWhat you gainWhat you give up
Vendor direct: Claude Team or Enterprise, Anthropic API, a ChatGPT workspace for Codex, Cursor’s backendThe vendorSeats or per token with the vendorEvery feature on release day, including cloud sessions and managed code reviewResidency choices are the vendor’s (Anthropic: us or global only)
Claude Platform on AWSAnthropic, with AWS authenticationAWS MarketplaceThe Claude API’s models and features on the same schedule, paid through AWSA separate Anthropic organization tied to your AWS account; credentials do not transfer
Partner clouds: Amazon Bedrock, Google Cloud’s Agent Platform, Microsoft FoundryAWS and Google operate theirs; Foundry is Anthropic-operated, hosted on Azure or on Anthropic infrastructureYour cloud billYour IAM, audit logs (CloudTrail, Cloud Audit Logs, Azure Monitor) and region choiceCloud sessions, Routines, Remote Control, Ultrareview and managed Code Review, which need a Claude subscription
LLM gateway (Anthropic’s Claude apps gateway, LiteLLM, your own)Whatever upstream the gateway forwards toThe upstream’s account, per tokenOne place for credentials, attribution, budgets, audit and provider switchingA service you upgrade with every agent release, and some features
Self-hosted open weightsYouYour hardwareNothing leaves your networkModel quality you must prove with your own evals; Anthropic does not support Claude Code on non-Claude models through any gateway

Two rows surprise people. Foundry deployments “hosted on Azure” keep prompts and completions within Azure; only usage metadata and safety-flagged content go to Anthropic. And a gateway credential replaces the developer’s claude.ai subscription, so usage is billed per token to the owner of the provider credential. Anthropic estimates enterprise Claude Code use at “around $13 per developer per active day and $150-250 per developer per month” (Claude Code costs page, checked 2026-09-26).

Which route keeps inference in your region?

Section titled “Which route keeps inference in your region?”

The defaults favour availability over geography. Decide the requirement first (EU, US, or none), then pick a route that can enforce it rather than merely allow it. Cursor has no region control of its own; see its tab under the configuration section.

Claude models: four routes, four different region controls

Section titled “Claude models: four routes, four different region controls”
RouteRegion controlEU-only inference possible?Trap
Anthropic API and Claude Platform on AWSinference_geo per request, or default_inference_geo and allowed_inference_geos per workspaceNo. Only "us" and "global" exist, and workspace geo (storage at rest) is "us" only"global" is the default: inference “may run in any available geography”
Amazon BedrockRegion plus cross-region inference profile prefix: Claude Code prefers eu. in eu-* regions, us. in us-*, apac. in ap-*, and global. everywhere elseYes, with eu. profilesThe prefix is a preference, not a guarantee
Google Cloud’s Agent PlatformCLOUD_ML_REGION: global, a multi-region such as eu, or a region such as europe-west1Yes, with eu or an EU regionA malformed value is treated as unset and falls back to us-east5; model availability differs per location
Microsoft FoundryDeployment type: Global Standard or US Data Zone StandardNo EU data zone is listed for ClaudeUS Data Zone covers only Azure-hosted deployments

On Bedrock, ANTHROPIC_BEDROCK_REGION_PREFIX (Claude Code 2.1.224 or later) is “a preference, not a guarantee”: a model with no eu. profile falls back to any matching profile, so enforce the region with pinned profile IDs and IAM. On Foundry, Fable 5.1 is available only hosted on Anthropic, so it cannot use the US Data Zone.

Residency costs money: US-only inference (inference_geo: "us", or a Foundry US Data Zone) is priced at 1.1x on Claude 4.6 and later models, and regional endpoints on Bedrock and Google Cloud carry a 10% premium over global ones (Anthropic pricing and data-residency pages, checked 2026-09-26).

OpenAI models in Codex: US residency is a managed switch

Section titled “OpenAI models in Codex: US residency is a managed switch”

In Codex 0.157.1, enforce_residency = "us" in /etc/codex/requirements.toml adds a US-residency header to every model request; any other value, such as "eu", fails at startup with unknown variant `eu`, expected `us` . The same file can pin model_provider, so the route does not depend on each laptop’s config.toml.

Codex 0.157.1 can enforce only US residency. OpenAI’s own material mentions EU data residency, so get EU terms from OpenAI in writing before you rely on them.

Which models can you use under zero data retention?

Section titled “Which models can you use under zero data retention?”

Zero data retention means the provider does not store prompts and responses after the response returns, except where law or abuse handling requires it. For Claude Code it is narrower than most contracts assume.

Who can get it. Qualified Claude for Enterprise accounts, enabled per organization by Anthropic’s account team, with no admin toggle and no inheritance by a new organization. API keys from a commercial organization under ZDR are covered too, and Claude Platform on AWS offers ZDR on request. Claude Code’s ZDR page covers only Anthropic’s direct platform. On Bedrock and Google Cloud the cloud provider is the data processor; for Foundry, Anthropic’s API retention page names Anthropic, so get Foundry’s terms for Claude Code in writing.

Which models. Everything your ZDR organization can reach stays available except the Covered Models, which require 30-day retention by default:

ModelUnder ZDR (Anthropic direct)How to use it anyway
Claude Opus 5.5 (the Claude Code default from v2.1.280, latest channel), Sonnet 5, Haiku 4.5 and older Opus and Sonnet modelsAvailable—
Claude Fable 5.1, Fable 5Only where Anthropic authorizes it; Anthropic says eligible customers can use Fable 5.1 under ZDR until its Enterprise Frontier Safeguards arriveTurn on 30-day retention for one workspace in Claude Console > Settings > Workspaces > Privacy controls; other workspaces keep ZDR
Claude Mythos 5.1, Mythos 5Same rule; invitation only in any case—

Where Fable is unavailable, the best alias resolves to Opus. A Covered Model called without retention returns “In order to access this model, your organization or workspace must have data retention enabled.”

What ZDR switches off. Cloud sessions (including Desktop-started ones), Claude Tag, Artifacts, /feedback, /bug, /share, Remote Control, managed Code Review and Ultrareview. Plan review around local /code-review and CI-run agents before you sign.

What ZDR never covers. Chat on claude.ai, Cowork, analytics metadata, seat management, and every MCP server or third-party integration the agent calls. Transcripts also stay on each laptop in plaintext under ~/.claude/projects/ for 30 days by default (cleanupPeriodDays), and content flagged for a usage-policy violation may be retained for up to two years.

Codex and Cursor retention terms could not be checked on 2026-09-26. Ask each vendor in writing which models or features require retention, and whether cloud agents follow the same terms.

What does an LLM gateway add, and what does it take away?

Section titled “What does an LLM gateway add, and what does it take away?”

A gateway is a proxy between the agent and the model provider: developers hold a gateway credential, and the gateway holds the provider credential, attribution, budgets, audit logging and provider switching.

Anthropic’s Claude apps gateway ships inside the claude binary (claude gateway --config gateway.yaml, Claude Code 2.1.195 or later). Developers sign in through your OIDC identity provider, IdP groups map to server-side model allowlists and managed-settings policies, and it exports OpenTelemetry metrics and fails over between cloud routes and the Anthropic API. Its data plane sends nothing to Anthropic unless the Anthropic API is an upstream. Its documented limits:

  • Browser SSO only, with no service-token flow, so CI pipelines authenticate to the provider directly.
  • OIDC only, one issuer per gateway, Linux servers only, no admin UI.
  • No server-side web search, no Remote Control, no /import on gateway sessions.
  • Prompt caching uses the 5-minute TTL only, so long sessions re-read more uncached context; watch cost per session after the switch.
  • Private-network address only: /login refuses a gateway that resolves to a public IP; list internally used public blocks in gatewayInternalNetworks (Claude Code 2.1.268 or later).

Other gateways work if they expose a supported API format and forward the headers and body fields Claude Code sends; one that lags Claude Code’s releases breaks the features whose fields it strips. Treat the gateway as supply chain: on 2026-03-24 the PyPI releases litellm 1.82.7 and 1.82.8 were malicious credential stealers, since removed (BerriAI/litellm issue 24518). Pin versions and verify hashes. Gateways and local models compares products; MCP registries and gateways covers the tool-side equivalent.

Self-hosting is a separate decision. Codex runs local models directly (codex --oss --local-provider ollama or lmstudio); Claude Code does not. Before an open-weight model carries real work, run your own evals and check the weights licence; see open-weight models.

Deliver the route through managed configuration, never a wiki page. The same files carry the rest of your managed agent policy.

Bedrock with EU inference, delivered in the env block of the managed settings file:

{
"env": {
"CLAUDE_CODE_USE_BEDROCK": "1",
"AWS_REGION": "eu-central-1",
"ANTHROPIC_BEDROCK_REGION_PREFIX": "eu",
"ANTHROPIC_DEFAULT_OPUS_MODEL": "eu.anthropic.claude-opus-5-5",
"ANTHROPIC_DEFAULT_SONNET_MODEL": "eu.anthropic.claude-sonnet-5"
}
}

The prefix only expresses a preference. Claude Code does not rewrite inference profile IDs you pin, so pin each model family you allow to an eu. profile that exists in your account: the block pins Opus and Sonnet, and Haiku needs the same ANTHROPIC_DEFAULT_HAIKU_MODEL entry if you allow it. Limit the IAM policy’s Resource to those inference profile ARNs so anything else fails closed.

Google Cloud’s Agent Platform in the EU multi-region:

{
"env": {
"CLAUDE_CODE_USE_VERTEX": "1",
"CLOUD_ML_REGION": "eu",
"ANTHROPIC_VERTEX_PROJECT_ID": "acme-ai-prod"
}
}

A gateway you run, with a per-developer key fetched from your secrets store:

{
"env": {
"ANTHROPIC_BASE_URL": "https://llm-gateway.internal.example.com"
},
"apiKeyHelper": "/usr/local/bin/get-gateway-key"
}

A managed ANTHROPIC_BASE_URL cannot be overridden by a developer’s shell export. Do not combine a gateway credential like this one with forceLoginMethod or forceLoginOrgUUID: either key blocks ANTHROPIC_API_KEY, ANTHROPIC_AUTH_TOKEN and apiKeyHelper, and nobody can start a session. Anthropic’s Claude apps gateway is the exception, selected with forceLoginMethod: "gateway" and forceLoginGatewayUrl. On cloud routes, pin all three ANTHROPIC_DEFAULT_*_MODEL variables so a new default arrives when you decide. /status shows the provider, region and base URL a session actually uses. Full setup is in LLM gateway configuration and corporate proxy configuration.

Fill one column per tool and have the security lead and data-protection officer sign it; it is what an auditor or customer asks for.

FieldClaude CodeCodexCursor
Data classes allowed (see privacy and data handling)
Route (vendor direct, Claude Platform on AWS, Bedrock, Google Cloud, Foundry, gateway, self-hosted)
Inference region and how it is enforced (setting name, managed file)
Retention: ZDR, 30-day, or provider policy; contract clause reference
Models allowed, and any Covered Models allowed with retention
Features lost on this route and their replacement
Who bills, and the budget owner
Evidence source (gateway logs, CloudTrail, Cloud Audit Logs, Azure Monitor)
Review date (at least quarterly, and on every model or plan change)

How do you prove traffic takes the route you chose?

Section titled “How do you prove traffic takes the route you chose?”

A route on paper is not a control. Prove it with machine-produced evidence on a schedule.

  1. Check the effective configuration on a sample of machines. In Claude Code, /status shows the provider, region, base URL and setting sources; a missing managed source means distribution failed. codex doctor prints the active provider, whether its credential is set and its endpoint reachable, and any overriding requirement.

  2. Read the provider-side record. Bedrock requests appear in CloudTrail, Google Cloud’s in Cloud Audit Logs, Foundry’s in Azure Monitor, and a gateway’s in its own logs or OpenTelemetry export. On the Anthropic API, the response’s usage.inference_geo field says where inference actually ran.

  3. Block the bypass at the network. Allow model endpoints only from the gateway or the approved cloud route. On every provider, Claude Code’s WebFetch tool still sends each hostname to api.anthropic.com for a safety check; only skipWebFetchPreflight: true turns that off (CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC does not). Record whether you allowlist the host or disable the check.

  4. Run a canary each quarter. One scripted request per tool and route, with the log line and region saved as evidence and fed to agent observability.

  5. Re-run your evals before switching route or model. Cloud routes lag the vendor’s defaults (on Foundry, Claude Code defaults to Sonnet 4.5 and opus is Opus 4.6); the new-model playbook covers the eval gate.

The platform team owns steps 1 to 4, the security lead signs the quarterly evidence, and the CTO signs any change of route (operating model).

Copy-paste prompts for auditing model routes

Section titled “Copy-paste prompts for auditing model routes”

What breaks when you move agents to a cloud route or gateway?

Section titled “What breaks when you move agents to a cloud route or gateway?”

Traffic leaves the region through a default. Each region-table trap sends traffic out without an error. Recovery: pin profile IDs or locations in managed settings, restrict IAM to them, and confirm with /status and the provider’s audit log.

Nobody can start a session after the gateway rollout. The error is “This machine’s managed settings require a first-party login”. Recovery: remove forceLoginMethod and forceLoginOrgUUID next to a third-party gateway credential, as the Claude Code tab explains.

Policy stops reaching developers. Server-managed settings need a direct connection to api.anthropic.com, so they miss gateway-routed sessions. Recovery: deliver the same keys through file-based managed settings.

Costs rise after the gateway switch. Billing moves to per-token rates and the Claude apps gateway forces the 5-minute cache TTL. Recovery: set per-group spend limits on the gateway and compare cost per accepted change for two weeks before and after, as described in cost governance.

A route change silently changes the model. Cloud providers enable and retire models on their own schedule, and auto mode there works only with Sonnet 5, Opus 4.7 and later, and Fable, so a default Foundry session on Sonnet 4.5 starts in Manual. Recovery: pin model versions per route and re-run your evals before each change.