Gateways and Local Models: LiteLLM, CC Switch, Claude Code Router, Ollama and LM Studio
A gateway or local model server lets Claude Code and Codex run on models other than the vendor’s own endpoint: Claude Code through ANTHROPIC_BASE_URL plus a credential variable, Codex through --oss or a custom provider. LiteLLM adds per-developer budgets, and Ollama, LM Studio and llama.cpp serve models locally. Remote Control and voice dictation stop working behind any of them.
This page is for the developer who has to make an agent work on a network where code may not leave the building, and the tech lead who then gives eight people the same setup with a spending limit. Security approved “local models only” for one repository; one laptop’s agent forgot its instructions after two turns, and the next used a :cloud tag without noticing. You need a setup that is verifiably local, budgeted per person, and loud when misconfigured.
What this gateway and local-model setup gives you
Section titled “What this gateway and local-model setup gives you”- The one rule every Claude Code gateway follows, and where it silently goes wrong.
- A decision table from your situation to a tool.
- A worked example: Claude Code on a local Ollama model, plus Codex and Cursor.
- A team workflow: Ollama behind LiteLLM on an air-gapped network, with a budget per developer and checks that prove nothing leaves.
- Three prompts, traps, and failure modes with fixes.
Which model to serve, and how to measure the quality gap against Claude Opus 5.5 or GPT-6 Astra, is on open-weight and self-hosted models. Model versions and prices live on the models hub. An organisation-wide gateway rollout for Claude models is on the LLM gateway page, and the hosting decision itself on where the model runs.
How does Claude Code talk to a gateway or local model?
Section titled “How does Claude Code talk to a gateway or local model?”Claude Code speaks the Anthropic Messages API. Any server that implements POST /v1/messages can stand in for Anthropic: point ANTHROPIC_BASE_URL at it and give Claude Code a credential. Ollama, LM Studio and llama.cpp’s llama-server implement that endpoint themselves, so a proxy is optional for one machine. LiteLLM and Claude Code Router sit in front of several backends.
The credential variable is the part people miss. Claude Code’s gateway documentation (checked 2026-09-26) states three rules:
| Variable | Sent as | Use it when |
|---|---|---|
ANTHROPIC_AUTH_TOKEN | Authorization: Bearer header | The gateway expects a bearer token; the default choice if nobody told you |
ANTHROPIC_API_KEY | x-api-key header | The gateway expects an API key; interactive sessions ask you once to approve it |
apiKeyHelper (settings) | Both headers | The credential rotates or comes from a vault |
- Setting only
ANTHROPIC_BASE_URLdoes not switch you over. A gateway credential variable takes precedence over a saved claude.ai login. Without one, the login stays the active credential, and its limits and billing apply. - A settings-file
envblock beats a shell export of the same variable. Never put a credential in a project’s.claude/settings.json, which is committed; use~/.claude/settings.jsonor managed settings. /statusis the proof. Its Status tab shows anAnthropic base URLline with your gateway address and anAuth tokenorAPI keyline naming the variable you set. ALogin methodline naming a claude.ai account means the gateway credential never arrived.
Ollama’s own Claude Code guide adds ANTHROPIC_API_KEY="" next to ANTHROPIC_AUTH_TOKEN, so a stale API key elsewhere in your environment cannot win.
Which gateway or local model server should you use?
Section titled “Which gateway or local model server should you use?”Start from who uses it and where the model runs.
| Your situation | Pick | Why |
|---|---|---|
| One developer, try a local model today | Ollama (ollama launch claude) | One command wires Claude Code or Codex to a local model; Codex also has --oss |
| One developer, a desktop app and a GUI model picker | LM Studio | lms CLI, a local /v1/messages endpoint, and Codex’s --local-provider lmstudio |
| Custom quantisation, no daemon, unusual hardware | llama.cpp (llama-server) | Maximum control; tool use requires --jinja |
| One developer, switch Claude Code between several hosted providers | CC Switch or Claude Code Router | CC Switch is a desktop app that rewrites provider config; CCR is a local gateway with routing rules and fallbacks |
| One developer, many hosted models behind one key | OpenRouter | Hosted router with an Anthropic-compatible endpoint (secondary source, see below) |
| A team with keys, budgets and logs, local or hosted models | LiteLLM proxy | Virtual keys per developer, max_budget per key or team, one place to swap providers |
| Claude models through your organisation’s gateway or cloud contract | Not this page | LLM gateway and where the model runs |
Popularity as of 2026-09-26 (GitHub stars, read through the GitHub API): Ollama 181,740; CC Switch 136,915; llama.cpp 129,535; LiteLLM 59,639; Claude Code Router 37,429; LM Studio’s open-source lms CLI 5,318 (the app itself is closed source). Stars measure attention, not safety.
Run Claude Code on a local Ollama model
Section titled “Run Claude Code on a local Ollama model”This is Ollama’s own Claude Code integration (docs/integrations/claude-code.mdx in ollama/ollama, checked 2026-09-26). It works because Ollama “supports a subset of the Anthropic Messages API”.
-
Start Ollama with an agent-sized context. Ollama’s default context depends on VRAM: 4k tokens below 24 GiB, 32k at 24 to 48 GiB, 256k at 48 GiB or more. Ollama’s own guidance for coding tools is at least 64,000 tokens, so on a typical laptop the default silently truncates the agent’s system prompt:
Terminal window # Terminal 1OLLAMA_CONTEXT_LENGTH=64000 ollama serve -
Launch Claude Code wired to a local model.
qwen3.5is the tag Ollama’s guide uses; replace it with a local tag you have evaluated:Terminal window # Terminal 2, interactive: pick a model, Ollama starts Claude Code for youollama launch claude# Headless, for scripts and CI: no selectors, pulls the model if neededollama launch claude --model qwen3.5 --yes -- -p "How does this repository work? Name the entry point and the test command."--yesskips the selectors and requires--model. Everything after--goes to Claude Code, so-pruns it in print mode. -
Confirm the model really runs locally and with the context you set:
Terminal window ollama ps # CONTEXT should read 64000; PROCESSOR should read 100% GPU
You should see Claude Code’s answer naming files that exist in your repository. If the answer is generic or ignores the files, the model did not call tools: check the context in ollama ps first, then try a model with tool calling. Ollama’s guide says to “use tools with compatible models”, and its compatibility page says tool-choice controls, deferred tools and hosted web search “are not fully supported”.
The manual setup from the same guide, without ollama launch:
export ANTHROPIC_AUTH_TOKEN=ollamaexport ANTHROPIC_API_KEY=""export ANTHROPIC_BASE_URL=http://localhost:11434claude --model qwen3.5Run the same local model in Codex or Cursor
Section titled “Run the same local model in Codex or Cursor”Covered above: ollama launch claude, or ANTHROPIC_BASE_URL plus ANTHROPIC_AUTH_TOKEN and ANTHROPIC_API_KEY="". For LM Studio the pattern is the same after lms server start: ANTHROPIC_BASE_URL=http://localhost:1234, with the loaded model’s ID as --model. The endpoint is documented in LM Studio’s docs; the Claude Code wiring comes from a search extract of lmstudio.ai (secondary), so run the smoke test before relying on it. For llama-server, start it with --jinja (tool use needs it) and -c 65536, then point ANTHROPIC_BASE_URL at its port; llama.cpp does not document the Claude Code combination, so the same smoke test applies.
Codex has a built-in switch for local providers (--oss and --local-provider, codex-cli 0.157.1 --help):
codex --oss --local-provider ollama -m qwen3.5codex --oss --local-provider lmstudio -m LOADED_MODEL_ID # after `lms server start`ollama launch codex # Ollama writes a dedicated Codex profilecodex --profile ollama-launch -c 'web_search="disabled"' # that profile, with Ollama's hosted web search offollama launch codex --restore removes the profile again.
Ollama’s Codex guide also asks for at least 64k context. A custom provider in config.toml must speak the Responses API (wire_api = "responses"); the full block is on open-weight and self-hosted models.
Whether Cursor’s agent can use a model served on your machine or network is not verified as of 2026-09-26. Do not plan a local-only or air-gapped setup on Cursor until your Cursor account team confirms the data path in writing. Both claude and codex run in Cursor’s integrated terminal, so a developer can keep Cursor as the editor and use a terminal agent for the restricted repository.
Switch providers for one developer: CC Switch, Claude Code Router and OpenRouter
Section titled “Switch providers for one developer: CC Switch, Claude Code Router and OpenRouter”These three are individual tools. None of them adds team budgets, and two of them send code to hosted providers.
CC Switch (farion1231/cc-switch, MIT) is a Tauri desktop app that switches providers, MCP servers, skills and prompts for Claude Code, Codex, Gemini CLI, OpenCode and six other tools named in its README. Install it with brew install --cask cc-switch on macOS, or the MSI, portable or Linux builds from its GitHub releases. Impostor sites exist and the app holds your provider keys, so download it only from GitHub releases or ccswitch.io, the one site its repository calls official.
Claude Code Router (npm @musistudio/claude-code-router 3.1.1, MIT) calls itself “a local model gateway and control plane for coding agents”. Version 3 is mainly a desktop app, with a CLI alternative that needs Node 22 or later:
npm install -g @musistudio/claude-code-routerccr ui # management UI at http://127.0.0.1:3458; the gateway listens on http://127.0.0.1:3456In the UI: Providers > Add Provider, then Server > Start, then Agent Config > choose Claude Code (or Codex, OpenCode and others) > apply.
OpenRouter is a hosted router with an Anthropic-compatible endpoint. Its cookbook sets ANTHROPIC_BASE_URL=https://openrouter.ai/api and puts the OpenRouter key in ANTHROPIC_AUTH_TOKEN, and it warns that tool use and extended thinking may not work on non-Anthropic models. Both details are secondary (a search extract of openrouter.ai, not the page itself, as of 2026-09-26); check the cookbook before you configure it.
Air-gapped team setup with a budget per developer
Section titled “Air-gapped team setup with a budget per developer”This is the workflow the rest of the page builds towards: local model server → LiteLLM gateway with keys and budgets → managed agent settings → smoke tests → evals → review as usual. The model server never listens beyond its own host, every developer has a key with a spending limit, and because the network blocks egress to hosted endpoints, a laptop without its gateway key fails instead of silently using a claude.ai login.
-
Serve the model on one host, on localhost only. Run Ollama (or vLLM, covered on the open-weight page) on the GPU server with
OLLAMA_CONTEXT_LENGTH=64000, pull only local tags, and leave it onlocalhost:11434. LiteLLM runs on the same host and is the only thing developers reach, so the model server needs no authentication of its own. At the network edge, block developer laptops’ egress to hosted model endpoints and allow onlyllm-gw.internal; without that block, a laptop with no gateway key falls back to a saved claude.ai login. -
Install LiteLLM pinned, from your internal mirror. PyPI versions 1.82.7 and 1.82.8, published on 24 March 2026, were credential stealers; 1.82.8 also added a
.pthfile that runs on any Python start. Both are gone from PyPI (the PyPI JSON on 2026-09-26 lists 1.82.6 and then 1.83.0), and the incident is tracked in BerriAI/litellm#24518. A mirror or cache can still hold them, so pin the version and verify hashes:Terminal window uv tool install 'litellm[proxy]==1.102.1'# For a server image: pin with hashes and install with pip --require-hashesecho 'litellm[proxy]==1.102.1' > requirements.inuv pip compile --generate-hashes requirements.in -o requirements.txt -
Describe the model and its internal price. Virtual keys and budgets need a Postgres database, which the proxy reads from
DATABASE_URL. A local model costs nothing per token by default, so a budget would never trigger. Give it an internal price per token (your GPU cost divided by measured throughput) withinput_cost_per_tokenandoutput_cost_per_token:litellm-config.yaml model_list:- model_name: local-coderlitellm_params:model: ollama_chat/qwen3.5api_base: http://localhost:11434input_cost_per_token: 0.0000004 # internal price, not a vendor priceoutput_cost_per_token: 0.0000016general_settings:master_key: os.environ/LITELLM_MASTER_KEYdatabase_url: os.environ/DATABASE_URLTerminal window litellm --config litellm-config.yaml # listens on port 4000 -
Issue one key per developer with a budget.
modelslimits what the key can call,max_budgetcaps spend in your internal currency, andbudget_durationresets it (for example30d). The master key comes from your secret store and never appears on a developer machine:Terminal window curl -sS https://llm-gw.internal:4000/key/generate \-H "Authorization: Bearer $LITELLM_MASTER_KEY" \-H "Content-Type: application/json" \-d '{"key_alias": "dev-anna", "models": ["local-coder"], "max_budget": 50, "budget_duration": "30d"}'The gateway URL is
https: developer keys and the master key cross the internal network, so terminate TLS at a reverse proxy in front of LiteLLM (or in LiteLLM) with a certificate from your internal CA. If Claude Code then reports certificate errors, seeNODE_EXTRA_CA_CERTSunder what breaks behind a gateway.For a team-level cap,
POST /team/newtakesmax_budgettoo, and/key/infoshows a key’s spend. -
Point every agent at the gateway through managed settings. Map every Claude Code model alias to the served model, because Claude Code also calls the Haiku alias for background work such as session titles, and your gateway does not serve it. Turn off nonessential traffic, which otherwise still goes to Anthropic and GitHub for version checks, telemetry and release notes:
Managed settings: env block {"env": {"ANTHROPIC_BASE_URL": "https://llm-gw.internal:4000","ANTHROPIC_MODEL": "local-coder","ANTHROPIC_DEFAULT_OPUS_MODEL": "local-coder","ANTHROPIC_DEFAULT_SONNET_MODEL": "local-coder","ANTHROPIC_DEFAULT_HAIKU_MODEL": "local-coder","CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC": "1"},"permissions": {"deny": ["WebSearch", "WebFetch"]},"skipWebFetchPreflight": true}Each developer supplies their own key as
ANTHROPIC_AUTH_TOKEN(a shell export,~/.claude/settings.json, or anapiKeyHelperthat reads your vault).CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFICalso disables auto-updates, so ship Claude Code through your package manager or an internal mirror.skipWebFetchPreflightstops the WebFetch domain check, which callsapi.anthropic.comeven with the variable set. Thedenyrule removes web search and web fetch altogether, because on an air-gapped network either one is a path out (or, through Ollama, a call to its hosted search). Where managed settings live and how they are enforced is on managed policy.For Codex, add a provider in
config.tomlwithbase_url = "https://llm-gw.internal:4000/v1",env_keynaming the variable that holds the developer’s key, andwire_api = "responses". LiteLLM documents a Responses-API endpoint; confirm it in your version, and let the smoke test show whether Codex’s tool calls survive the round trip. -
Verify the setup before anyone works on it. Run the checks in the next section on one laptop, then roll out.
-
Review agent output the same way as before. A local model changes where tokens are generated, not how you prove the change is right: tests, CI, a review bot and a human approval still gate the merge.
How do you prove the gateway setup works and stays local?
Section titled “How do you prove the gateway setup works and stays local?”Run these checks on one developer machine before rollout and after every gateway, model or Claude Code upgrade. The tech lead (or the platform owner for the gateway) signs off on the checks; the pull request reviewer still signs off on each change.
| Check | How | Pass condition |
|---|---|---|
| Gateway answers | curl -sS -w '\n%{http_code}\n' -X POST "$ANTHROPIC_BASE_URL/v1/messages" -H "Authorization: Bearer $ANTHROPIC_AUTH_TOKEN" -H "anthropic-version: 2023-06-01" -H "content-type: application/json" -d '{"model": "local-coder", "max_tokens": 1, "messages": [{"role": "user", "content": "."}]}' | HTTP 200 and a JSON body with "type":"message" and a content array; a 401 means the key or the header is wrong |
| Session is on the gateway | /status in Claude Code | Anthropic base URL shows the gateway; Auth token names ANTHROPIC_AUTH_TOKEN; no Login method line |
| Tools work | The smoke-test prompt below | The answer names real files and a real command; no generic prose |
| Context fits | ollama ps on the server; /context in Claude Code | Server context 64,000 or more; system prompt and tools leave room for work |
| Budget bites | Issue a test key with "max_budget": 0.01 and send prompts until it is spent | The gateway rejects further requests; /key/info shows the spend |
| Nothing leaves | Egress logs from the laptop and the server during a session | No connections to api.anthropic.com, ollama.com or other hosted endpoints |
Output quality is a separate question from connectivity. Before the team relies on a local model, run the same 20 to 30 tasks from your own merged pull requests through the gateway and through the default model, as described in the eval protocol on open-weight and self-hosted models. No primary source quantifies the quality drop of a local model in Claude Code, so your own pass rate is the only number that counts. The tooling for repeatable runs is on evaluating coding agents.
Copy-paste prompts for gateways and local models
Section titled “Copy-paste prompts for gateways and local models”What breaks behind a gateway, and how to recover
Section titled “What breaks behind a gateway, and how to recover”- Remote Control and
/voiceare gone. Both need a claude.ai identity and are unavailable whileANTHROPIC_API_KEY,ANTHROPIC_AUTH_TOKENor anapiKeyHelperis active. Since v2.1.196, Remote Control is also disabled wheneverANTHROPIC_BASE_URLpoints at a non-Anthropic host. Recovery: accept it for gateway sessions, or unset the gateway variables and log in with claude.ai for work that may use Anthropic.claude doctornames what blocks Remote Control. For alternatives, see remote and mobile clients and voice input. - Slack and cloud sessions ignore the gateway. They always use Anthropic’s API. Recovery: do not enable those surfaces for users whose traffic must stay on the gateway.
- A startup warning names two credential sources. A gateway credential and a saved login are both present. Recovery:
/logoutto keep only the gateway credential. - Claude Code asks you to log in although the curl test passes. The credential sits in a project settings file, which applies only after the first-run wizard and trust prompt. Recovery: export
ANTHROPIC_AUTH_TOKENin the shell, or put it in~/.claude/settings.jsonor managed settings. 400errors namingcontext_managementorExtra inputs are not permitted. The upstream rejects fields Claude Code sends to Anthropic endpoints. Recovery:CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS=1.- The gateway reports a context limit in its own words. Claude Code does not recognise the error, so it does not compact and retry. Recovery:
/compact, then setCLAUDE_CODE_AUTO_COMPACT_WINDOWto the gateway’s limit. That variable is clamped to at least 100,000 tokens, so for a 64K local model/compactremains the recovery, and a larger served context is the real fix. - Background calls fail with an unknown model. Claude Code asked for the Haiku alias. Recovery: set all four
ANTHROPIC_*MODELvariables to the served model, as in step 5. - The agent forgets instructions or answers without reading files. Truncated context or a model without tool calling. Recovery:
ollama psfor the context, then the smoke-test prompt; switch model if tools still fail. 403with an HTML body, and nothing in the gateway’s logs. A web application firewall in front of the gateway blocked the request body, because prompts contain source code and XML-style tags. Recovery: exempt the/v1/messagespath from body inspection.- Certificate errors while curl works. Claude Code does not trust your internal CA. Recovery: set
NODE_EXTRA_CA_CERTSto the CA bundle. - A budget never triggers. The local model has no price, so spend stays at zero. Recovery: set
input_cost_per_tokenandoutput_cost_per_tokenas in step 3.
Where to go next with gateways and local models
Section titled “Where to go next with gateways and local models”- Open-weight and self-hosted models: which model to serve, vLLM for a shared GPU server, and the eval protocol that measures the quality gap.
- LLM gateway for Claude Code: rolling out a gateway for Claude models across an organisation.
- What did the agents cost?: gateway budgets next to
/usage, ccusage and spend reports. - Where the model runs: SaaS, cloud routes, gateways and self-hosting compared for a CTO.
- Agent sandboxes compared: isolate the agent’s machine as well as its model traffic.
- Agent tools overview: the other shelves, from multiplexers to review bots.
Frequently asked questions
How do I point Claude Code at a gateway or a local model?
Set ANTHROPIC_BASE_URL to a server that implements the Anthropic Messages API (/v1/messages) and set a credential variable, ANTHROPIC_AUTH_TOKEN or ANTHROPIC_API_KEY. Without the credential variable a saved claude.ai login stays active. Run /status to confirm both.
What stops working in Claude Code behind a gateway?
Remote Control and voice dictation, because both need a claude.ai identity; Remote Control is also disabled whenever ANTHROPIC_BASE_URL points at a non-Anthropic host. Claude Code in Slack and cloud sessions always use Anthropic's API.
Is LiteLLM safe to install?
Pin it. PyPI versions 1.82.7 and 1.82.8, published on 24 March 2026, were malicious credential stealers. Both are gone from PyPI, but an unpinned install in a cached or mirrored environment can still resolve to them.
Do Ollama models always run locally?
No. Ollama model tags ending in :cloud (or -cloud) run on Ollama's cloud, so your code leaves the machine. Pin local tags only when data residency is the reason you use Ollama.