Open-weight and self-hosted models for coding agents
Open-weight coding models (GLM-5.3, Qwen3.8, Kimi K3, DeepSeek V4, and Mistral’s Devstral 2 and Medium 3.5) can drive Claude Code, Codex and several vendor CLIs from a server you control. They fit air-gapped networks, strict data rules or a cost floor, and a team should route work to them only after its own evals show the quality gap is acceptable.
Your security team rules that one product line’s code may not leave the building. Procurement wants a second model supplier before a seven-figure renewal. Finance asks why a nightly job that labels 40,000 test failures runs on your most expensive model. Someone proposes “just run Qwen on our own GPUs”, and nobody can say how much worse the agent gets, which tools still work, or what the licence allows.
This page is for the CTO deciding whether to fund a self-hosted model and for the developer asked to wire one into the team’s agent and prove it works. It gives you a September 2026 shortlist with licence status, the harness matrix, verified vLLM and Ollama configuration, an eval protocol with a runner and decision rule, a cost formula, a data-rules checklist, and the traps that silently send code off the machine. Model versions, prices and context windows live on the models hub.
Which open-weight coding models matter in September 2026?
Section titled “Which open-weight coding models matter in September 2026?”Five families cover almost every serious evaluation. The table was verified on 2026-09-26 against each vendor’s GitHub repository or package registry; “check the model card” means the weights licence is on Hugging Face, which this site could not read.
| Model | Vendor | Size (total / active) | Weights licence | Vendor’s own agent | Evidence for coding |
|---|---|---|---|---|---|
| GLM-5.3 (also GLM-5.3-Flash) | Z.ai | 744B / 40B (Flash: 320B / 18B) | Check the model card; the repo code is Apache-2.0 | none named; runs in Claude Code | Only open-weights entry on the official Terminal-Bench 4.0 board |
| Qwen3.8 (2.4T-A95B and 27B) | Alibaba | 2.4T / 95B, and a 27B model | Check the model card | Qwen Code | Vendor: “a Qwen-Max-class model to open release”; scores not verified |
| Kimi K3 | Moonshot AI | 2.8T / 104B | Custom “Kimi K3 License”, not OSI open source | Kimi Code | Vendor-reported only (Terminal-Bench 2.1) |
| DeepSeek V4 (Flash, Pro) | DeepSeek | Flash 284B / 13B, Pro 1.6T / 49B (secondary: SitePoint) | Check the model card | DeepSeek Harness dsh (developer preview) | No primary source |
| Devstral 2 · Devstral Small 2 · Mistral Medium 3.5 | Mistral | Devstral 123B and 24B (secondary: VentureBeat) | Medium 3.5: Modified MIT; Devstral: check the model card | Mistral Vibe | Devstral 2 self-reports 72.2% on SWE-bench Verified (secondary: VentureBeat) |
- GLM-5.3 is the one model with independent evidence. On the official Terminal-Bench 4.0 leaderboard (read 2026-09-26), GLM-5.3 at
maxeffort in Claude Code scores 41.8% ± 3.2 and GPT-5.6 Sol in Codex 37.3% ± 3.8; the 95% intervals overlap, so the board does not separate them. Claude Fable 5.1 in Claude Code scores 57.9% ± 3.8. - Vendor numbers do not cross benchmark versions. Kimi K3 reports 88.3 on Terminal-Bench 2.1, an older, near-saturated version with no conversion to 4.0.
- The SWE-bench Verified board is frozen (newest entry 2026-02-26), so any 2026 figure for these models is a self-report.
None of this predicts results on your repository; the eval section below does. See also reading coding-agent benchmarks.
When is an open-weight model the right call?
Section titled “When is an open-weight model the right call?”An open-weight model answers a constraint, not a quality need. The model rule from the models hub still holds: start on the tool’s default model (Claude Opus 5.5 in Claude Code from v2.1.280, the latest channel; GPT-6 Astra in Codex), tune effort before switching model, and switch only when your own evals say so.
| Your situation | Open-weight model? | Why |
|---|---|---|
| Code may not leave your network (defence, some banks, some public sector) | Yes, self-hosted | No third party processes the code. First check whether a cloud route under your contract satisfies the rule (where the model runs) |
| Data rules require a named region or zero retention, but a cloud provider is allowed | Usually no | A hosted frontier model through your cloud contract is simpler |
| High-volume, narrow jobs (labelling failures, generating fixtures, first-pass triage) | Maybe | Pays off at volume only if the eval pass rate on that job is close to the hosted model’s |
| Hard, open-ended feature work on a large codebase | No, as the default | The board’s only open-weights entry sits well below the best Claude entries |
| A second supplier to avoid lock-in | Yes, as a tested fallback | Keep one open-weight setup passing your eval set (avoiding lock-in) |
| A laptop without network on a plane | Maybe, for small tasks | A 24B–27B model runs locally, with a clear quality drop on multi-file work |
Which harnesses accept a self-hosted model?
Section titled “Which harnesses accept a self-hosted model?”The harness matters as much as the model: the board entry is GLM-5.3 in Claude Code. Each harness speaks one API format, and the serving stack must match it.
| Harness | How it reaches your model | API it needs | What you lose |
|---|---|---|---|
| Claude Code 2.1.283 | ANTHROPIC_BASE_URL plus a token variable; --model or ANTHROPIC_MODEL | Anthropic Messages (/v1/messages) | Remote Control (off when the base URL is not api.anthropic.com); /voice (needs a claude.ai login); MCP tool search, off by default (set ENABLE_TOOL_SEARCH=true if your server forwards tool_reference blocks) |
| Codex 0.157.1 | --oss with --local-provider ollama or lmstudio, or a [model_providers.<id>] block in config.toml | Responses API only (wire_api = "responses"); "chat" is rejected | Features tied to a ChatGPT or OpenAI login, such as codex cloud tasks |
| Cursor | Not verified (cursor.com unreachable on 2026-09-26) | — | See the Cursor tab below |
Qwen Code, Kimi Code, Mistral Vibe, DeepSeek dsh | Each vendor’s CLI, built for its own models | Vendor-specific | Portability of your prompts, hooks and skills |
| OpenCode, Cline, Aider, Goose, Crush | Provider settings in each tool | Mostly OpenAI-compatible | Your team’s Claude Code or Codex configuration |
Serving stacks that implement those APIs, verified on 2026-09-26:
- vLLM (PyPI
vllm0.30.0): OpenAI-compatible “plus Anthropic Messages API”, with tool calling on both. The production choice for a shared GPU server. - Ollama: “a subset of the Anthropic Messages API”;
ollama launch claudeandollama launch codexwire the agents for you. - LM Studio:
POST /v1/messageslocally; Codex reaches it through--local-provider lmstudio. - llama.cpp
llama-server: Anthropic Messages compatible; tool use requires--jinja. - LiteLLM or Claude Code Router: a gateway with virtual keys and budgets in front of several backends (gateways and local models).
Popularity as of 2026-09-26: Ollama has 181,740 GitHub stars (GitHub, ollama/ollama, read 2026-09-26); stars measure attention, not fitness for a regulated network.
How do you connect Claude Code or Codex to a self-hosted model?
Section titled “How do you connect Claude Code or Codex to a self-hosted model?”Start with a single-machine Ollama trial to learn the failure modes, then move to vLLM on a shared server.
Trial on one machine with Ollama. Ollama’s default context is set by VRAM (4k under 24 GiB), which silently truncates an agent’s system prompt, so raise it first:
# Terminal 1: start Ollama with an agent-sized contextOLLAMA_CONTEXT_LENGTH=64000 ollama serve
# Terminal 2: let Ollama launch Claude Code wired to a local model (interactive session)ollama launch claude --model LOCAL_MODEL
# Terminal 3, while the session runs: confirm the loaded contextollama ps # CONTEXT should read 64000; PROCESSOR should read 100% GPUTeam server with vLLM. Map every model alias to the served model: Claude Code also calls the haiku alias for background work, and without the mapping those calls fail with “model not found”:
# Developer machine or CI runner. VLLM_API_KEY comes from your secret store, never a literal.export ANTHROPIC_BASE_URL=http://gpu-01.internal:8000export ANTHROPIC_AUTH_TOKEN="$VLLM_API_KEY" # sent as "Authorization: Bearer ..."export ANTHROPIC_API_KEY="" # stop a saved key from overriding the gatewayexport ANTHROPIC_MODEL=SERVED_MODELexport ANTHROPIC_DEFAULT_OPUS_MODEL=SERVED_MODELexport ANTHROPIC_DEFAULT_SONNET_MODEL=SERVED_MODELexport ANTHROPIC_DEFAULT_HAIKU_MODEL=SERVED_MODELexport CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 # no auto-updates, telemetry or error reportingclaudeRun /status to confirm the base URL, and /context to see how much of the window the system prompt and tools take before you type anything (the context cost). For a team, put the variables in the env block of a managed settings file (LLM gateway).
Trial on one machine. Codex has a built-in switch for Ollama and LM Studio:
OLLAMA_CONTEXT_LENGTH=64000 ollama serve # terminal 1codex --oss --local-provider ollama -m LOCAL_MODEL # terminal 2codex --oss --local-provider lmstudio -m LOCAL_MODEL # or LM Studio after `lms server start`Team server with vLLM. Put the provider in a profile file, $CODEX_HOME/selfhosted.config.toml (~/.codex/selfhosted.config.toml by default). -p selfhosted layers it on top of the base config (codex --help, 0.157.1), so your hosted setup stays the default. The keys below are read from the codex-cli 0.157.1 source:
model = "SERVED_MODEL"model_provider = "vllm"
[model_providers.vllm]name = "vLLM (internal)"base_url = "http://gpu-01.internal:8000/v1"env_key = "VLLM_API_KEY" # the variable Codex reads the key from; never put the key herewire_api = "responses" # "chat" fails with "no longer supported" in 0.157.1Then run a smoke test that forces a tool call:
codex -p selfhosted exec "List the three largest files under src/ and say which one has no tests."Whether Cursor’s agent can use a model served inside your network could not be verified (cursor.com was unreachable on 2026-09-26). Do not plan an air-gapped rollout on Cursor until Cursor confirms the data path in writing. Until then, use Cursor for code that may leave the network, and Claude Code or Codex in its integrated terminal for the restricted repositories.
How do you set up a vLLM server and prove the agent works?
Section titled “How do you set up a vLLM server and prove the agent works?”One GPU server, one agent, one smoke task with a pass or fail answer. It assumes a Linux host with enough GPU memory (see the sizing rule below) and Python with uv.
-
Check the licence before you download anything. Record the licence name and use restrictions from the Hugging Face model card in your model register.
-
Serve the model with tool calling on, which a coding agent needs. Use the tool-call parser the model card names: vLLM lists parsers by family (for example
glm47,kimi_k3,deepseek_v4,qwen3_xml/qwen3_coder,mistralin vLLM 0.30.0), and a newer model may need a newer vLLM release.Terminal window # GPU server. VLLM_API_KEY comes from your secret store.uv venv && uv pip install vllm==0.30.0vllm serve MODEL_REPO_OR_PATH \--served-model-name SERVED_MODEL \--enable-auto-tool-choice --tool-call-parser PARSER \--max-model-len 131072 \--api-key "$VLLM_API_KEY" \--port 8000MODEL_REPO_OR_PATHis the Hugging Face repository or, air-gapped, the local directory holding the weights. -
Point the agent at it with the Claude Code or Codex configuration above.
-
Run a smoke task that needs tools, edits and tests. In a throwaway branch of a real repository:
Terminal window claude -p "Add a failing unit test for parseDuration('90m') returning 5400 seconds, run it, then fix src/time.ts until it passes. Report the test command and its final output." \--permission-mode acceptEdits --allowedTools "Bash(npm *)" --output-format json -
Read the result, not the prose. Run the test yourself (
npm test -- time) and checkgit diff --stat.
What you should see: a JSON result that names the test command, a diff touching src/time.ts and one test file, and a test run you reproduce as green. Prose without edits means tool calling is broken; a truncated plan means the context is too small.
Sizing rule. Weight memory ≈ total parameters × bytes per parameter: 744B at 8 bits is about 744 GB, at 4 bits about 372 GB, before the KV cache. Mixture-of-experts helps speed, not memory. The 24B–27B class (Devstral Small 2, Qwen3.8-27B) fits one workstation.
How do you measure the gap on your own evals?
Section titled “How do you measure the gap on your own evals?”Public boards rank model and harness pairs on someone else’s tasks. Your question is whether the model is good enough for your work in your harness, so build a small eval set from your own history. The full method is on evals for coding agents; this is the version for a model switch.
-
Collect 20 to 30 tasks from merged pull requests. Pick the kind of work you plan to route, each with tests the pull request added. Keep the parent commit, the issue text as the prompt, and the pull request’s test files, at their repository paths, as the hidden checker.
-
Freeze the setup. Same harness version,
CLAUDE.mdorAGENTS.md, MCP servers, skills and effort level. Record them next to the results. -
Run each task in a disposable worktree, on the baseline (the tool’s default model) and on the candidate, three times per setup if you can afford it; one run on 25 tasks cannot separate models a few points apart.
-
Score with the hidden tests, not with the agent’s own report. A task passes when the hidden tests pass and the type check and linter are clean.
-
Compare four numbers, then apply the decision rule below.
A minimal runner, for a repository where each task lives in evals/tasks/<id>/ with prompt.md, base_commit and a hidden-tests/ directory:
#!/usr/bin/env bash# evals/run.sh SETUP_NAME (run on an isolated runner with no production credentials)set -euo pipefailsetup="$1"; mkdir -p "evals/results/$setup"for dir in evals/tasks/*/; do id=$(basename "$dir") wt="$(mktemp -d)/$id" git worktree add --detach "$wt" "$(cat "$dir/base_commit")" >/dev/null ( cd "$wt" # A fresh worktree has no node_modules. Offline, point npm at your internal registry mirror. npm ci --silent || { echo "$id SETUP-FAIL"; exit 0; } claude -p "$(cat "$OLDPWD/$dir/prompt.md")" --permission-mode acceptEdits \ --allowedTools "Bash(npm *)" --output-format json > "$OLDPWD/evals/results/$setup/$id.json" || true cp -r "$OLDPWD/$dir/hidden-tests/." . # hidden tests are stored at their repository paths if npm test --silent && npx tsc --noEmit && npm run lint --silent; then echo "$id PASS"; else echo "$id FAIL"; fi ) | tee -a "evals/results/$setup/summary.txt" git worktree remove --force "$wt"doneUse your repository’s own install, test, type-check and lint commands in place of the npm ones. Run it twice: once with the default environment, once after exporting the self-hosted variables from the Claude Code tab. For Codex, replace the two-line claude command with the line below; --json writes JSONL events, so the result file is .jsonl:
codex -a never exec -p selfhosted --sandbox workspace-write --json "$(cat "$OLDPWD/$dir/prompt.md")" > "$OLDPWD/evals/results/$setup/$id.jsonl" || true| Metric | Definition | Why it matters |
|---|---|---|
| Pass rate | Tasks whose hidden tests, type check and lint pass ÷ tasks run | The quality gap, in your terms |
| Cost per solved task | (GPU-hours × hourly cost, or API spend) ÷ tasks passed | A model that fails half the time is not cheap |
| Intervention rate | Runs that asked a question, looped or hit the turn limit ÷ runs | A person has to step in |
| Tool-call error rate | Malformed or rejected tool calls ÷ tool calls (JSON transcripts) | First sign of a wrong parser |
Decision rule. Route a job type to the open-weight model when its pass rate is within an agreed margin of the baseline (for example, five points on 25 tasks × 3 runs), cost per solved task is lower, and the intervention rate is no higher. Otherwise keep the hosted model and re-run on the next release (the new-model playbook).
What is the cost floor of a self-hosted model?
Section titled “What is the cost floor of a self-hosted model?”Self-hosting replaces a per-token price with a fixed cost that you pay whether the GPUs are busy or not. Put two numbers in front of finance instead of a per-token comparison:
- Self-hosted cost per solved task = (GPU amortisation or reserved rental + power + engineer time to run the server) ÷ (monthly tasks × eval pass rate).
- Hosted cost per solved task = API or plan spend for the same job ÷ tasks solved on the baseline. Break-even is the monthly volume at which the two are equal.
Below break-even, a hosted model (the models hub lists cheaper tiers) wins; above it, self-hosting wins only if the pass rate holds. Easy to miss: re-running the eval for every release, and the reviewer time a lower pass rate adds. See also the economics of agentic engineering.
Which data rules apply to open-weight models?
Section titled “Which data rules apply to open-weight models?”“Open weights” answers where the model runs, not what the licence allows. Use this checklist before the model touches production code.
- Licence named in the register. Record the weights licence from the model card, not the GitHub repository’s code licence (GLM’s code is Apache-2.0; its weights licence is on the card).
- Self-hosted, not the vendor’s API. A vendor’s hosted API or coding plan sends your code to that vendor, under its terms and jurisdiction.
- Local tags only. Ollama tags ending in
:cloudrun on Ollama’s cloud, not your machine. - Egress blocked, not trusted. Block outbound traffic from agent runners at the network layer, and set
CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1. - Hosted open-weight models count as hosted. Kimi K3 in GitHub Copilot is not self-hosting.
Residency and zero retention for frontier models: where the model runs.
How do you keep quality provable on a weaker model?
Section titled “How do you keep quality provable on a weaker model?”Keep every gate you have for the hosted model and add the eval set as a gate. Output is verified the same way whichever model wrote it; only the catch rate changes.
| Gate | What it proves | Owner |
|---|---|---|
| Eval set pass rate, re-run on every model or harness upgrade | The model still solves your kind of task | The platform team |
| Tests, type check and lint in CI on every pull request | This change works | CI, required status checks |
| A failing-test-first rule in the agent’s instructions | The test would catch a regression | The developer who dispatched the task |
| A reviewer agent on the hosted baseline, where data rules allow | A second model checks every diff | The tech lead picks the paths |
| Revert and incident rate per model, monthly | Weaker output is not reaching production | The CTO, with the platform team |
Sign-off stays with people: the developer owns each merge; the platform owner routes job types to models on the eval numbers. When the pass rate drops after an upgrade, roll back first and investigate second. Guardrails for the self-hosted box: permissions and sandboxing.
What breaks when an agent runs on an open-weight model?
Section titled “What breaks when an agent runs on an open-weight model?”The agent answers in prose and never edits a file. Tool calling is off or the parser is wrong. Recovery: restart vLLM with --enable-auto-tool-choice and the model card’s parser (on llama.cpp, add --jinja), then re-run the smoke task.
Claude Code still calls Anthropic. A saved login or API key, or a settings file, overrides the shell variables. Recovery: set ANTHROPIC_API_KEY="" and ANTHROPIC_AUTH_TOKEN, run /status, and move the variables into the settings file’s env block.
Long tasks collapse halfway. --max-model-len is too small or the KV cache does not fit. Recovery: check /context, raise --max-model-len if memory allows, and split the task with a plan file.
The eval looks great and production does not. The eval set holds public code the model trained on, or only easy tasks. Recovery: use private code, check that hidden tests fail first, and add harder tasks from recent incidents.
The server rejects requests with “Unexpected value(s) for the anthropic-beta header” or “Extra inputs are not permitted”. It does not implement Anthropic’s beta headers and tool-schema fields. Recovery: set CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS=1, which strips them (environment-variable reference, checked against v2.1.283), and re-run the smoke task.
The licence blocks the use you planned. The expensive failure. Recovery: route the job type back to the baseline and pick a model legal has approved. Prevent it with step 1 of the server setup.
Where to go next with open-weight models
Section titled “Where to go next with open-weight models”Frequently asked questions
Can Claude Code run an open-weight model?
Yes. Any server that implements the Anthropic Messages API (vLLM, Ollama, LM Studio, llama.cpp or a LiteLLM gateway) works through ANTHROPIC_BASE_URL.
Can Codex run an open-weight model?
Yes. codex-cli 0.157.1 has --oss with --local-provider ollama or lmstudio, and custom providers that speak the Responses API (wire_api = "chat" is no longer supported).
Which open-weight model should I evaluate first?
GLM-5.3, the only open-weights entry on the official Terminal-Bench 4.0 board (41.8% ± 3.2 in Claude Code). Check its weights licence on the model card before use.
How big is the quality gap to Claude Opus 5.5 or GPT-6 Astra?
No primary source measures it for your code. Run 20 to 30 tasks from your own merged pull requests through both setups and compare pass rate, cost per solved task and intervention rate.