Skip to content

Gateways and Local Models: LiteLLM, CC Switch, Claude Code Router, Ollama and LM Studio

A gateway or local model server lets Claude Code and Codex run on models other than the vendor’s own endpoint: Claude Code through ANTHROPIC_BASE_URL plus a credential variable, Codex through --oss or a custom provider. LiteLLM adds per-developer budgets, and Ollama, LM Studio and llama.cpp serve models locally. Remote Control and voice dictation stop working behind any of them.

This page is for the developer who has to make an agent work on a network where code may not leave the building, and the tech lead who then gives eight people the same setup with a spending limit. Security approved “local models only” for one repository; one laptop’s agent forgot its instructions after two turns, and the next used a :cloud tag without noticing. You need a setup that is verifiably local, budgeted per person, and loud when misconfigured.

What this gateway and local-model setup gives you

Section titled “What this gateway and local-model setup gives you”
  • The one rule every Claude Code gateway follows, and where it silently goes wrong.
  • A decision table from your situation to a tool.
  • A worked example: Claude Code on a local Ollama model, plus Codex and Cursor.
  • A team workflow: Ollama behind LiteLLM on an air-gapped network, with a budget per developer and checks that prove nothing leaves.
  • Three prompts, traps, and failure modes with fixes.

Which model to serve, and how to measure the quality gap against Claude Opus 5.5 or GPT-6 Astra, is on open-weight and self-hosted models. Model versions and prices live on the models hub. An organisation-wide gateway rollout for Claude models is on the LLM gateway page, and the hosting decision itself on where the model runs.

How does Claude Code talk to a gateway or local model?

Section titled “How does Claude Code talk to a gateway or local model?”

Claude Code speaks the Anthropic Messages API. Any server that implements POST /v1/messages can stand in for Anthropic: point ANTHROPIC_BASE_URL at it and give Claude Code a credential. Ollama, LM Studio and llama.cpp’s llama-server implement that endpoint themselves, so a proxy is optional for one machine. LiteLLM and Claude Code Router sit in front of several backends.

The credential variable is the part people miss. Claude Code’s gateway documentation (checked 2026-09-26) states three rules:

VariableSent asUse it when
ANTHROPIC_AUTH_TOKENAuthorization: Bearer headerThe gateway expects a bearer token; the default choice if nobody told you
ANTHROPIC_API_KEYx-api-key headerThe gateway expects an API key; interactive sessions ask you once to approve it
apiKeyHelper (settings)Both headersThe credential rotates or comes from a vault
  1. Setting only ANTHROPIC_BASE_URL does not switch you over. A gateway credential variable takes precedence over a saved claude.ai login. Without one, the login stays the active credential, and its limits and billing apply.
  2. A settings-file env block beats a shell export of the same variable. Never put a credential in a project’s .claude/settings.json, which is committed; use ~/.claude/settings.json or managed settings.
  3. /status is the proof. Its Status tab shows an Anthropic base URL line with your gateway address and an Auth token or API key line naming the variable you set. A Login method line naming a claude.ai account means the gateway credential never arrived.

Ollama’s own Claude Code guide adds ANTHROPIC_API_KEY="" next to ANTHROPIC_AUTH_TOKEN, so a stale API key elsewhere in your environment cannot win.

Which gateway or local model server should you use?

Section titled “Which gateway or local model server should you use?”

Start from who uses it and where the model runs.

Your situationPickWhy
One developer, try a local model todayOllama (ollama launch claude)One command wires Claude Code or Codex to a local model; Codex also has --oss
One developer, a desktop app and a GUI model pickerLM Studiolms CLI, a local /v1/messages endpoint, and Codex’s --local-provider lmstudio
Custom quantisation, no daemon, unusual hardwarellama.cpp (llama-server)Maximum control; tool use requires --jinja
One developer, switch Claude Code between several hosted providersCC Switch or Claude Code RouterCC Switch is a desktop app that rewrites provider config; CCR is a local gateway with routing rules and fallbacks
One developer, many hosted models behind one keyOpenRouterHosted router with an Anthropic-compatible endpoint (secondary source, see below)
A team with keys, budgets and logs, local or hosted modelsLiteLLM proxyVirtual keys per developer, max_budget per key or team, one place to swap providers
Claude models through your organisation’s gateway or cloud contractNot this pageLLM gateway and where the model runs

Popularity as of 2026-09-26 (GitHub stars, read through the GitHub API): Ollama 181,740; CC Switch 136,915; llama.cpp 129,535; LiteLLM 59,639; Claude Code Router 37,429; LM Studio’s open-source lms CLI 5,318 (the app itself is closed source). Stars measure attention, not safety.

This is Ollama’s own Claude Code integration (docs/integrations/claude-code.mdx in ollama/ollama, checked 2026-09-26). It works because Ollama “supports a subset of the Anthropic Messages API”.

  1. Start Ollama with an agent-sized context. Ollama’s default context depends on VRAM: 4k tokens below 24 GiB, 32k at 24 to 48 GiB, 256k at 48 GiB or more. Ollama’s own guidance for coding tools is at least 64,000 tokens, so on a typical laptop the default silently truncates the agent’s system prompt:

    Terminal window
    # Terminal 1
    OLLAMA_CONTEXT_LENGTH=64000 ollama serve
  2. Launch Claude Code wired to a local model. qwen3.5 is the tag Ollama’s guide uses; replace it with a local tag you have evaluated:

    Terminal window
    # Terminal 2, interactive: pick a model, Ollama starts Claude Code for you
    ollama launch claude
    # Headless, for scripts and CI: no selectors, pulls the model if needed
    ollama launch claude --model qwen3.5 --yes -- -p "How does this repository work? Name the entry point and the test command."

    --yes skips the selectors and requires --model. Everything after -- goes to Claude Code, so -p runs it in print mode.

  3. Confirm the model really runs locally and with the context you set:

    Terminal window
    ollama ps # CONTEXT should read 64000; PROCESSOR should read 100% GPU

You should see Claude Code’s answer naming files that exist in your repository. If the answer is generic or ignores the files, the model did not call tools: check the context in ollama ps first, then try a model with tool calling. Ollama’s guide says to “use tools with compatible models”, and its compatibility page says tool-choice controls, deferred tools and hosted web search “are not fully supported”.

The manual setup from the same guide, without ollama launch:

Terminal window
export ANTHROPIC_AUTH_TOKEN=ollama
export ANTHROPIC_API_KEY=""
export ANTHROPIC_BASE_URL=http://localhost:11434
claude --model qwen3.5

Run the same local model in Codex or Cursor

Section titled “Run the same local model in Codex or Cursor”

Covered above: ollama launch claude, or ANTHROPIC_BASE_URL plus ANTHROPIC_AUTH_TOKEN and ANTHROPIC_API_KEY="". For LM Studio the pattern is the same after lms server start: ANTHROPIC_BASE_URL=http://localhost:1234, with the loaded model’s ID as --model. The endpoint is documented in LM Studio’s docs; the Claude Code wiring comes from a search extract of lmstudio.ai (secondary), so run the smoke test before relying on it. For llama-server, start it with --jinja (tool use needs it) and -c 65536, then point ANTHROPIC_BASE_URL at its port; llama.cpp does not document the Claude Code combination, so the same smoke test applies.

Switch providers for one developer: CC Switch, Claude Code Router and OpenRouter

Section titled “Switch providers for one developer: CC Switch, Claude Code Router and OpenRouter”

These three are individual tools. None of them adds team budgets, and two of them send code to hosted providers.

CC Switch (farion1231/cc-switch, MIT) is a Tauri desktop app that switches providers, MCP servers, skills and prompts for Claude Code, Codex, Gemini CLI, OpenCode and six other tools named in its README. Install it with brew install --cask cc-switch on macOS, or the MSI, portable or Linux builds from its GitHub releases. Impostor sites exist and the app holds your provider keys, so download it only from GitHub releases or ccswitch.io, the one site its repository calls official.

Claude Code Router (npm @musistudio/claude-code-router 3.1.1, MIT) calls itself “a local model gateway and control plane for coding agents”. Version 3 is mainly a desktop app, with a CLI alternative that needs Node 22 or later:

Terminal window
npm install -g @musistudio/claude-code-router
ccr ui # management UI at http://127.0.0.1:3458; the gateway listens on http://127.0.0.1:3456

In the UI: Providers > Add Provider, then Server > Start, then Agent Config > choose Claude Code (or Codex, OpenCode and others) > apply.

OpenRouter is a hosted router with an Anthropic-compatible endpoint. Its cookbook sets ANTHROPIC_BASE_URL=https://openrouter.ai/api and puts the OpenRouter key in ANTHROPIC_AUTH_TOKEN, and it warns that tool use and extended thinking may not work on non-Anthropic models. Both details are secondary (a search extract of openrouter.ai, not the page itself, as of 2026-09-26); check the cookbook before you configure it.

Air-gapped team setup with a budget per developer

Section titled “Air-gapped team setup with a budget per developer”

This is the workflow the rest of the page builds towards: local model server → LiteLLM gateway with keys and budgets → managed agent settings → smoke tests → evals → review as usual. The model server never listens beyond its own host, every developer has a key with a spending limit, and because the network blocks egress to hosted endpoints, a laptop without its gateway key fails instead of silently using a claude.ai login.

  1. Serve the model on one host, on localhost only. Run Ollama (or vLLM, covered on the open-weight page) on the GPU server with OLLAMA_CONTEXT_LENGTH=64000, pull only local tags, and leave it on localhost:11434. LiteLLM runs on the same host and is the only thing developers reach, so the model server needs no authentication of its own. At the network edge, block developer laptops’ egress to hosted model endpoints and allow only llm-gw.internal; without that block, a laptop with no gateway key falls back to a saved claude.ai login.

  2. Install LiteLLM pinned, from your internal mirror. PyPI versions 1.82.7 and 1.82.8, published on 24 March 2026, were credential stealers; 1.82.8 also added a .pth file that runs on any Python start. Both are gone from PyPI (the PyPI JSON on 2026-09-26 lists 1.82.6 and then 1.83.0), and the incident is tracked in BerriAI/litellm#24518. A mirror or cache can still hold them, so pin the version and verify hashes:

    Terminal window
    uv tool install 'litellm[proxy]==1.102.1'
    # For a server image: pin with hashes and install with pip --require-hashes
    echo 'litellm[proxy]==1.102.1' > requirements.in
    uv pip compile --generate-hashes requirements.in -o requirements.txt
  3. Describe the model and its internal price. Virtual keys and budgets need a Postgres database, which the proxy reads from DATABASE_URL. A local model costs nothing per token by default, so a budget would never trigger. Give it an internal price per token (your GPU cost divided by measured throughput) with input_cost_per_token and output_cost_per_token:

    litellm-config.yaml
    model_list:
    - model_name: local-coder
    litellm_params:
    model: ollama_chat/qwen3.5
    api_base: http://localhost:11434
    input_cost_per_token: 0.0000004 # internal price, not a vendor price
    output_cost_per_token: 0.0000016
    general_settings:
    master_key: os.environ/LITELLM_MASTER_KEY
    database_url: os.environ/DATABASE_URL
    Terminal window
    litellm --config litellm-config.yaml # listens on port 4000
  4. Issue one key per developer with a budget. models limits what the key can call, max_budget caps spend in your internal currency, and budget_duration resets it (for example 30d). The master key comes from your secret store and never appears on a developer machine:

    Terminal window
    curl -sS https://llm-gw.internal:4000/key/generate \
    -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
    -H "Content-Type: application/json" \
    -d '{"key_alias": "dev-anna", "models": ["local-coder"], "max_budget": 50, "budget_duration": "30d"}'

    The gateway URL is https: developer keys and the master key cross the internal network, so terminate TLS at a reverse proxy in front of LiteLLM (or in LiteLLM) with a certificate from your internal CA. If Claude Code then reports certificate errors, see NODE_EXTRA_CA_CERTS under what breaks behind a gateway.

    For a team-level cap, POST /team/new takes max_budget too, and /key/info shows a key’s spend.

  5. Point every agent at the gateway through managed settings. Map every Claude Code model alias to the served model, because Claude Code also calls the Haiku alias for background work such as session titles, and your gateway does not serve it. Turn off nonessential traffic, which otherwise still goes to Anthropic and GitHub for version checks, telemetry and release notes:

    Managed settings: env block
    {
    "env": {
    "ANTHROPIC_BASE_URL": "https://llm-gw.internal:4000",
    "ANTHROPIC_MODEL": "local-coder",
    "ANTHROPIC_DEFAULT_OPUS_MODEL": "local-coder",
    "ANTHROPIC_DEFAULT_SONNET_MODEL": "local-coder",
    "ANTHROPIC_DEFAULT_HAIKU_MODEL": "local-coder",
    "CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC": "1"
    },
    "permissions": {
    "deny": ["WebSearch", "WebFetch"]
    },
    "skipWebFetchPreflight": true
    }

    Each developer supplies their own key as ANTHROPIC_AUTH_TOKEN (a shell export, ~/.claude/settings.json, or an apiKeyHelper that reads your vault). CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC also disables auto-updates, so ship Claude Code through your package manager or an internal mirror. skipWebFetchPreflight stops the WebFetch domain check, which calls api.anthropic.com even with the variable set. The deny rule removes web search and web fetch altogether, because on an air-gapped network either one is a path out (or, through Ollama, a call to its hosted search). Where managed settings live and how they are enforced is on managed policy.

    For Codex, add a provider in config.toml with base_url = "https://llm-gw.internal:4000/v1", env_key naming the variable that holds the developer’s key, and wire_api = "responses". LiteLLM documents a Responses-API endpoint; confirm it in your version, and let the smoke test show whether Codex’s tool calls survive the round trip.

  6. Verify the setup before anyone works on it. Run the checks in the next section on one laptop, then roll out.

  7. Review agent output the same way as before. A local model changes where tokens are generated, not how you prove the change is right: tests, CI, a review bot and a human approval still gate the merge.

How do you prove the gateway setup works and stays local?

Section titled “How do you prove the gateway setup works and stays local?”

Run these checks on one developer machine before rollout and after every gateway, model or Claude Code upgrade. The tech lead (or the platform owner for the gateway) signs off on the checks; the pull request reviewer still signs off on each change.

CheckHowPass condition
Gateway answerscurl -sS -w '\n%{http_code}\n' -X POST "$ANTHROPIC_BASE_URL/v1/messages" -H "Authorization: Bearer $ANTHROPIC_AUTH_TOKEN" -H "anthropic-version: 2023-06-01" -H "content-type: application/json" -d '{"model": "local-coder", "max_tokens": 1, "messages": [{"role": "user", "content": "."}]}'HTTP 200 and a JSON body with "type":"message" and a content array; a 401 means the key or the header is wrong
Session is on the gateway/status in Claude CodeAnthropic base URL shows the gateway; Auth token names ANTHROPIC_AUTH_TOKEN; no Login method line
Tools workThe smoke-test prompt belowThe answer names real files and a real command; no generic prose
Context fitsollama ps on the server; /context in Claude CodeServer context 64,000 or more; system prompt and tools leave room for work
Budget bitesIssue a test key with "max_budget": 0.01 and send prompts until it is spentThe gateway rejects further requests; /key/info shows the spend
Nothing leavesEgress logs from the laptop and the server during a sessionNo connections to api.anthropic.com, ollama.com or other hosted endpoints

Output quality is a separate question from connectivity. Before the team relies on a local model, run the same 20 to 30 tasks from your own merged pull requests through the gateway and through the default model, as described in the eval protocol on open-weight and self-hosted models. No primary source quantifies the quality drop of a local model in Claude Code, so your own pass rate is the only number that counts. The tooling for repeatable runs is on evaluating coding agents.

Copy-paste prompts for gateways and local models

Section titled “Copy-paste prompts for gateways and local models”

What breaks behind a gateway, and how to recover

Section titled “What breaks behind a gateway, and how to recover”
  • Remote Control and /voice are gone. Both need a claude.ai identity and are unavailable while ANTHROPIC_API_KEY, ANTHROPIC_AUTH_TOKEN or an apiKeyHelper is active. Since v2.1.196, Remote Control is also disabled whenever ANTHROPIC_BASE_URL points at a non-Anthropic host. Recovery: accept it for gateway sessions, or unset the gateway variables and log in with claude.ai for work that may use Anthropic. claude doctor names what blocks Remote Control. For alternatives, see remote and mobile clients and voice input.
  • Slack and cloud sessions ignore the gateway. They always use Anthropic’s API. Recovery: do not enable those surfaces for users whose traffic must stay on the gateway.
  • A startup warning names two credential sources. A gateway credential and a saved login are both present. Recovery: /logout to keep only the gateway credential.
  • Claude Code asks you to log in although the curl test passes. The credential sits in a project settings file, which applies only after the first-run wizard and trust prompt. Recovery: export ANTHROPIC_AUTH_TOKEN in the shell, or put it in ~/.claude/settings.json or managed settings.
  • 400 errors naming context_management or Extra inputs are not permitted. The upstream rejects fields Claude Code sends to Anthropic endpoints. Recovery: CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS=1.
  • The gateway reports a context limit in its own words. Claude Code does not recognise the error, so it does not compact and retry. Recovery: /compact, then set CLAUDE_CODE_AUTO_COMPACT_WINDOW to the gateway’s limit. That variable is clamped to at least 100,000 tokens, so for a 64K local model /compact remains the recovery, and a larger served context is the real fix.
  • Background calls fail with an unknown model. Claude Code asked for the Haiku alias. Recovery: set all four ANTHROPIC_*MODEL variables to the served model, as in step 5.
  • The agent forgets instructions or answers without reading files. Truncated context or a model without tool calling. Recovery: ollama ps for the context, then the smoke-test prompt; switch model if tools still fail.
  • 403 with an HTML body, and nothing in the gateway’s logs. A web application firewall in front of the gateway blocked the request body, because prompts contain source code and XML-style tags. Recovery: exempt the /v1/messages path from body inspection.
  • Certificate errors while curl works. Claude Code does not trust your internal CA. Recovery: set NODE_EXTRA_CA_CERTS to the CA bundle.
  • A budget never triggers. The local model has no price, so spend stays at zero. Recovery: set input_cost_per_token and output_cost_per_token as in step 3.

Where to go next with gateways and local models

Section titled “Where to go next with gateways and local models”

Frequently asked questions

How do I point Claude Code at a gateway or a local model?

Set ANTHROPIC_BASE_URL to a server that implements the Anthropic Messages API (/v1/messages) and set a credential variable, ANTHROPIC_AUTH_TOKEN or ANTHROPIC_API_KEY. Without the credential variable a saved claude.ai login stays active. Run /status to confirm both.

What stops working in Claude Code behind a gateway?

Remote Control and voice dictation, because both need a claude.ai identity; Remote Control is also disabled whenever ANTHROPIC_BASE_URL points at a non-Anthropic host. Claude Code in Slack and cloud sessions always use Anthropic's API.

Is LiteLLM safe to install?

Pin it. PyPI versions 1.82.7 and 1.82.8, published on 24 March 2026, were malicious credential stealers. Both are gone from PyPI, but an unpinned install in a cached or mirrored environment can still resolve to them.

Do Ollama models always run locally?

No. Ollama model tags ending in :cloud (or -cloud) run on Ollama's cloud, so your code leaves the machine. Pin local tags only when data residency is the reason you use Ollama.