Skip to content

Open-weight and self-hosted models for coding agents

Open-weight coding models (GLM-5.3, Qwen3.8, Kimi K3, DeepSeek V4, and Mistral’s Devstral 2 and Medium 3.5) can drive Claude Code, Codex and several vendor CLIs from a server you control. They fit air-gapped networks, strict data rules or a cost floor, and a team should route work to them only after its own evals show the quality gap is acceptable.

Your security team rules that one product line’s code may not leave the building. Procurement wants a second model supplier before a seven-figure renewal. Finance asks why a nightly job that labels 40,000 test failures runs on your most expensive model. Someone proposes “just run Qwen on our own GPUs”, and nobody can say how much worse the agent gets, which tools still work, or what the licence allows.

This page is for the CTO deciding whether to fund a self-hosted model and for the developer asked to wire one into the team’s agent and prove it works. It gives you a September 2026 shortlist with licence status, the harness matrix, verified vLLM and Ollama configuration, an eval protocol with a runner and decision rule, a cost formula, a data-rules checklist, and the traps that silently send code off the machine. Model versions, prices and context windows live on the models hub.

Which open-weight coding models matter in September 2026?

Section titled “Which open-weight coding models matter in September 2026?”

Five families cover almost every serious evaluation. The table was verified on 2026-09-26 against each vendor’s GitHub repository or package registry; “check the model card” means the weights licence is on Hugging Face, which this site could not read.

ModelVendorSize (total / active)Weights licenceVendor’s own agentEvidence for coding
GLM-5.3 (also GLM-5.3-Flash)Z.ai744B / 40B (Flash: 320B / 18B)Check the model card; the repo code is Apache-2.0none named; runs in Claude CodeOnly open-weights entry on the official Terminal-Bench 4.0 board
Qwen3.8 (2.4T-A95B and 27B)Alibaba2.4T / 95B, and a 27B modelCheck the model cardQwen CodeVendor: “a Qwen-Max-class model to open release”; scores not verified
Kimi K3Moonshot AI2.8T / 104BCustom “Kimi K3 License”, not OSI open sourceKimi CodeVendor-reported only (Terminal-Bench 2.1)
DeepSeek V4 (Flash, Pro)DeepSeekFlash 284B / 13B, Pro 1.6T / 49B (secondary: SitePoint)Check the model cardDeepSeek Harness dsh (developer preview)No primary source
Devstral 2 · Devstral Small 2 · Mistral Medium 3.5MistralDevstral 123B and 24B (secondary: VentureBeat)Medium 3.5: Modified MIT; Devstral: check the model cardMistral VibeDevstral 2 self-reports 72.2% on SWE-bench Verified (secondary: VentureBeat)
  • GLM-5.3 is the one model with independent evidence. On the official Terminal-Bench 4.0 leaderboard (read 2026-09-26), GLM-5.3 at max effort in Claude Code scores 41.8% ± 3.2 and GPT-5.6 Sol in Codex 37.3% ± 3.8; the 95% intervals overlap, so the board does not separate them. Claude Fable 5.1 in Claude Code scores 57.9% ± 3.8.
  • Vendor numbers do not cross benchmark versions. Kimi K3 reports 88.3 on Terminal-Bench 2.1, an older, near-saturated version with no conversion to 4.0.
  • The SWE-bench Verified board is frozen (newest entry 2026-02-26), so any 2026 figure for these models is a self-report.

None of this predicts results on your repository; the eval section below does. See also reading coding-agent benchmarks.

When is an open-weight model the right call?

Section titled “When is an open-weight model the right call?”

An open-weight model answers a constraint, not a quality need. The model rule from the models hub still holds: start on the tool’s default model (Claude Opus 5.5 in Claude Code from v2.1.280, the latest channel; GPT-6 Astra in Codex), tune effort before switching model, and switch only when your own evals say so.

Your situationOpen-weight model?Why
Code may not leave your network (defence, some banks, some public sector)Yes, self-hostedNo third party processes the code. First check whether a cloud route under your contract satisfies the rule (where the model runs)
Data rules require a named region or zero retention, but a cloud provider is allowedUsually noA hosted frontier model through your cloud contract is simpler
High-volume, narrow jobs (labelling failures, generating fixtures, first-pass triage)MaybePays off at volume only if the eval pass rate on that job is close to the hosted model’s
Hard, open-ended feature work on a large codebaseNo, as the defaultThe board’s only open-weights entry sits well below the best Claude entries
A second supplier to avoid lock-inYes, as a tested fallbackKeep one open-weight setup passing your eval set (avoiding lock-in)
A laptop without network on a planeMaybe, for small tasksA 24B–27B model runs locally, with a clear quality drop on multi-file work

Which harnesses accept a self-hosted model?

Section titled “Which harnesses accept a self-hosted model?”

The harness matters as much as the model: the board entry is GLM-5.3 in Claude Code. Each harness speaks one API format, and the serving stack must match it.

HarnessHow it reaches your modelAPI it needsWhat you lose
Claude Code 2.1.283ANTHROPIC_BASE_URL plus a token variable; --model or ANTHROPIC_MODELAnthropic Messages (/v1/messages)Remote Control (off when the base URL is not api.anthropic.com); /voice (needs a claude.ai login); MCP tool search, off by default (set ENABLE_TOOL_SEARCH=true if your server forwards tool_reference blocks)
Codex 0.157.1--oss with --local-provider ollama or lmstudio, or a [model_providers.<id>] block in config.tomlResponses API only (wire_api = "responses"); "chat" is rejectedFeatures tied to a ChatGPT or OpenAI login, such as codex cloud tasks
CursorNot verified (cursor.com unreachable on 2026-09-26)—See the Cursor tab below
Qwen Code, Kimi Code, Mistral Vibe, DeepSeek dshEach vendor’s CLI, built for its own modelsVendor-specificPortability of your prompts, hooks and skills
OpenCode, Cline, Aider, Goose, CrushProvider settings in each toolMostly OpenAI-compatibleYour team’s Claude Code or Codex configuration

Serving stacks that implement those APIs, verified on 2026-09-26:

  • vLLM (PyPI vllm 0.30.0): OpenAI-compatible “plus Anthropic Messages API”, with tool calling on both. The production choice for a shared GPU server.
  • Ollama: “a subset of the Anthropic Messages API”; ollama launch claude and ollama launch codex wire the agents for you.
  • LM Studio: POST /v1/messages locally; Codex reaches it through --local-provider lmstudio.
  • llama.cpp llama-server: Anthropic Messages compatible; tool use requires --jinja.
  • LiteLLM or Claude Code Router: a gateway with virtual keys and budgets in front of several backends (gateways and local models).

Popularity as of 2026-09-26: Ollama has 181,740 GitHub stars (GitHub, ollama/ollama, read 2026-09-26); stars measure attention, not fitness for a regulated network.

How do you connect Claude Code or Codex to a self-hosted model?

Section titled “How do you connect Claude Code or Codex to a self-hosted model?”

Start with a single-machine Ollama trial to learn the failure modes, then move to vLLM on a shared server.

Trial on one machine with Ollama. Ollama’s default context is set by VRAM (4k under 24 GiB), which silently truncates an agent’s system prompt, so raise it first:

Terminal window
# Terminal 1: start Ollama with an agent-sized context
OLLAMA_CONTEXT_LENGTH=64000 ollama serve
# Terminal 2: let Ollama launch Claude Code wired to a local model (interactive session)
ollama launch claude --model LOCAL_MODEL
# Terminal 3, while the session runs: confirm the loaded context
ollama ps # CONTEXT should read 64000; PROCESSOR should read 100% GPU

Team server with vLLM. Map every model alias to the served model: Claude Code also calls the haiku alias for background work, and without the mapping those calls fail with “model not found”:

Terminal window
# Developer machine or CI runner. VLLM_API_KEY comes from your secret store, never a literal.
export ANTHROPIC_BASE_URL=http://gpu-01.internal:8000
export ANTHROPIC_AUTH_TOKEN="$VLLM_API_KEY" # sent as "Authorization: Bearer ..."
export ANTHROPIC_API_KEY="" # stop a saved key from overriding the gateway
export ANTHROPIC_MODEL=SERVED_MODEL
export ANTHROPIC_DEFAULT_OPUS_MODEL=SERVED_MODEL
export ANTHROPIC_DEFAULT_SONNET_MODEL=SERVED_MODEL
export ANTHROPIC_DEFAULT_HAIKU_MODEL=SERVED_MODEL
export CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 # no auto-updates, telemetry or error reporting
claude

Run /status to confirm the base URL, and /context to see how much of the window the system prompt and tools take before you type anything (the context cost). For a team, put the variables in the env block of a managed settings file (LLM gateway).

How do you set up a vLLM server and prove the agent works?

Section titled “How do you set up a vLLM server and prove the agent works?”

One GPU server, one agent, one smoke task with a pass or fail answer. It assumes a Linux host with enough GPU memory (see the sizing rule below) and Python with uv.

  1. Check the licence before you download anything. Record the licence name and use restrictions from the Hugging Face model card in your model register.

  2. Serve the model with tool calling on, which a coding agent needs. Use the tool-call parser the model card names: vLLM lists parsers by family (for example glm47, kimi_k3, deepseek_v4, qwen3_xml / qwen3_coder, mistral in vLLM 0.30.0), and a newer model may need a newer vLLM release.

    Terminal window
    # GPU server. VLLM_API_KEY comes from your secret store.
    uv venv && uv pip install vllm==0.30.0
    vllm serve MODEL_REPO_OR_PATH \
    --served-model-name SERVED_MODEL \
    --enable-auto-tool-choice --tool-call-parser PARSER \
    --max-model-len 131072 \
    --api-key "$VLLM_API_KEY" \
    --port 8000

    MODEL_REPO_OR_PATH is the Hugging Face repository or, air-gapped, the local directory holding the weights.

  3. Point the agent at it with the Claude Code or Codex configuration above.

  4. Run a smoke task that needs tools, edits and tests. In a throwaway branch of a real repository:

    Terminal window
    claude -p "Add a failing unit test for parseDuration('90m') returning 5400 seconds, run it, then fix src/time.ts until it passes. Report the test command and its final output." \
    --permission-mode acceptEdits --allowedTools "Bash(npm *)" --output-format json
  5. Read the result, not the prose. Run the test yourself (npm test -- time) and check git diff --stat.

What you should see: a JSON result that names the test command, a diff touching src/time.ts and one test file, and a test run you reproduce as green. Prose without edits means tool calling is broken; a truncated plan means the context is too small.

Sizing rule. Weight memory ≈ total parameters × bytes per parameter: 744B at 8 bits is about 744 GB, at 4 bits about 372 GB, before the KV cache. Mixture-of-experts helps speed, not memory. The 24B–27B class (Devstral Small 2, Qwen3.8-27B) fits one workstation.

How do you measure the gap on your own evals?

Section titled “How do you measure the gap on your own evals?”

Public boards rank model and harness pairs on someone else’s tasks. Your question is whether the model is good enough for your work in your harness, so build a small eval set from your own history. The full method is on evals for coding agents; this is the version for a model switch.

  1. Collect 20 to 30 tasks from merged pull requests. Pick the kind of work you plan to route, each with tests the pull request added. Keep the parent commit, the issue text as the prompt, and the pull request’s test files, at their repository paths, as the hidden checker.

  2. Freeze the setup. Same harness version, CLAUDE.md or AGENTS.md, MCP servers, skills and effort level. Record them next to the results.

  3. Run each task in a disposable worktree, on the baseline (the tool’s default model) and on the candidate, three times per setup if you can afford it; one run on 25 tasks cannot separate models a few points apart.

  4. Score with the hidden tests, not with the agent’s own report. A task passes when the hidden tests pass and the type check and linter are clean.

  5. Compare four numbers, then apply the decision rule below.

A minimal runner, for a repository where each task lives in evals/tasks/<id>/ with prompt.md, base_commit and a hidden-tests/ directory:

#!/usr/bin/env bash
# evals/run.sh SETUP_NAME (run on an isolated runner with no production credentials)
set -euo pipefail
setup="$1"; mkdir -p "evals/results/$setup"
for dir in evals/tasks/*/; do
id=$(basename "$dir")
wt="$(mktemp -d)/$id"
git worktree add --detach "$wt" "$(cat "$dir/base_commit")" >/dev/null
( cd "$wt"
# A fresh worktree has no node_modules. Offline, point npm at your internal registry mirror.
npm ci --silent || { echo "$id SETUP-FAIL"; exit 0; }
claude -p "$(cat "$OLDPWD/$dir/prompt.md")" --permission-mode acceptEdits \
--allowedTools "Bash(npm *)" --output-format json > "$OLDPWD/evals/results/$setup/$id.json" || true
cp -r "$OLDPWD/$dir/hidden-tests/." . # hidden tests are stored at their repository paths
if npm test --silent && npx tsc --noEmit && npm run lint --silent; then echo "$id PASS"; else echo "$id FAIL"; fi
) | tee -a "evals/results/$setup/summary.txt"
git worktree remove --force "$wt"
done

Use your repository’s own install, test, type-check and lint commands in place of the npm ones. Run it twice: once with the default environment, once after exporting the self-hosted variables from the Claude Code tab. For Codex, replace the two-line claude command with the line below; --json writes JSONL events, so the result file is .jsonl:

Terminal window
codex -a never exec -p selfhosted --sandbox workspace-write --json "$(cat "$OLDPWD/$dir/prompt.md")" > "$OLDPWD/evals/results/$setup/$id.jsonl" || true
MetricDefinitionWhy it matters
Pass rateTasks whose hidden tests, type check and lint pass ÷ tasks runThe quality gap, in your terms
Cost per solved task(GPU-hours × hourly cost, or API spend) ÷ tasks passedA model that fails half the time is not cheap
Intervention rateRuns that asked a question, looped or hit the turn limit ÷ runsA person has to step in
Tool-call error rateMalformed or rejected tool calls ÷ tool calls (JSON transcripts)First sign of a wrong parser

Decision rule. Route a job type to the open-weight model when its pass rate is within an agreed margin of the baseline (for example, five points on 25 tasks × 3 runs), cost per solved task is lower, and the intervention rate is no higher. Otherwise keep the hosted model and re-run on the next release (the new-model playbook).

What is the cost floor of a self-hosted model?

Section titled “What is the cost floor of a self-hosted model?”

Self-hosting replaces a per-token price with a fixed cost that you pay whether the GPUs are busy or not. Put two numbers in front of finance instead of a per-token comparison:

  • Self-hosted cost per solved task = (GPU amortisation or reserved rental + power + engineer time to run the server) ÷ (monthly tasks × eval pass rate).
  • Hosted cost per solved task = API or plan spend for the same job ÷ tasks solved on the baseline. Break-even is the monthly volume at which the two are equal.

Below break-even, a hosted model (the models hub lists cheaper tiers) wins; above it, self-hosting wins only if the pass rate holds. Easy to miss: re-running the eval for every release, and the reviewer time a lower pass rate adds. See also the economics of agentic engineering.

Which data rules apply to open-weight models?

Section titled “Which data rules apply to open-weight models?”

“Open weights” answers where the model runs, not what the licence allows. Use this checklist before the model touches production code.

  • Licence named in the register. Record the weights licence from the model card, not the GitHub repository’s code licence (GLM’s code is Apache-2.0; its weights licence is on the card).
  • Self-hosted, not the vendor’s API. A vendor’s hosted API or coding plan sends your code to that vendor, under its terms and jurisdiction.
  • Local tags only. Ollama tags ending in :cloud run on Ollama’s cloud, not your machine.
  • Egress blocked, not trusted. Block outbound traffic from agent runners at the network layer, and set CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1.
  • Hosted open-weight models count as hosted. Kimi K3 in GitHub Copilot is not self-hosting.

Residency and zero retention for frontier models: where the model runs.

How do you keep quality provable on a weaker model?

Section titled “How do you keep quality provable on a weaker model?”

Keep every gate you have for the hosted model and add the eval set as a gate. Output is verified the same way whichever model wrote it; only the catch rate changes.

GateWhat it provesOwner
Eval set pass rate, re-run on every model or harness upgradeThe model still solves your kind of taskThe platform team
Tests, type check and lint in CI on every pull requestThis change worksCI, required status checks
A failing-test-first rule in the agent’s instructionsThe test would catch a regressionThe developer who dispatched the task
A reviewer agent on the hosted baseline, where data rules allowA second model checks every diffThe tech lead picks the paths
Revert and incident rate per model, monthlyWeaker output is not reaching productionThe CTO, with the platform team

Sign-off stays with people: the developer owns each merge; the platform owner routes job types to models on the eval numbers. When the pass rate drops after an upgrade, roll back first and investigate second. Guardrails for the self-hosted box: permissions and sandboxing.

What breaks when an agent runs on an open-weight model?

Section titled “What breaks when an agent runs on an open-weight model?”

The agent answers in prose and never edits a file. Tool calling is off or the parser is wrong. Recovery: restart vLLM with --enable-auto-tool-choice and the model card’s parser (on llama.cpp, add --jinja), then re-run the smoke task.

Claude Code still calls Anthropic. A saved login or API key, or a settings file, overrides the shell variables. Recovery: set ANTHROPIC_API_KEY="" and ANTHROPIC_AUTH_TOKEN, run /status, and move the variables into the settings file’s env block.

Long tasks collapse halfway. --max-model-len is too small or the KV cache does not fit. Recovery: check /context, raise --max-model-len if memory allows, and split the task with a plan file.

The eval looks great and production does not. The eval set holds public code the model trained on, or only easy tasks. Recovery: use private code, check that hidden tests fail first, and add harder tasks from recent incidents.

The server rejects requests with “Unexpected value(s) for the anthropic-beta header” or “Extra inputs are not permitted”. It does not implement Anthropic’s beta headers and tool-schema fields. Recovery: set CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS=1, which strips them (environment-variable reference, checked against v2.1.283), and re-run the smoke task.

The licence blocks the use you planned. The expensive failure. Recovery: route the job type back to the baseline and pick a model legal has approved. Prevent it with step 1 of the server setup.

Frequently asked questions

Can Claude Code run an open-weight model?

Yes. Any server that implements the Anthropic Messages API (vLLM, Ollama, LM Studio, llama.cpp or a LiteLLM gateway) works through ANTHROPIC_BASE_URL.

Can Codex run an open-weight model?

Yes. codex-cli 0.157.1 has --oss with --local-provider ollama or lmstudio, and custom providers that speak the Responses API (wire_api = "chat" is no longer supported).

Which open-weight model should I evaluate first?

GLM-5.3, the only open-weights entry on the official Terminal-Bench 4.0 board (41.8% ± 3.2 in Claude Code). Check its weights licence on the model card before use.

How big is the quality gap to Claude Opus 5.5 or GPT-6 Astra?

No primary source measures it for your code. Run 20 to 30 tasks from your own merged pull requests through both setups and compare pass rate, cost per solved task and intervention rate.