Infrastructure MCP: Docker MCP Toolkit, Kubernetes and Terraform
Infrastructure MCP servers give Claude Code, Codex and Cursor scoped access to running systems: the Kubernetes MCP server reads cluster state, the Terraform MCP server reads the Terraform Registry and HCP Terraform runs, and the Docker MCP Toolkit runs servers in containers behind one gateway. Connected read-only, they let an agent diagnose and propose while changes ship as reviewed code.
This page is for developers who own a service end to end and tech leads who decide what agents may touch in a cluster. The checkout pod in staging restarts every 40 seconds. You run kubectl get pods, kubectl describe, kubectl logs --previous, paste three screens of output into the agent, and it asks for the events you did not copy. The next day the same agent writes a Terraform module with an S3 argument that the provider dropped two major versions ago.
New to MCP? Start with the MCP overview, which explains servers, transports and scopes before this page connects them to a cluster.
What you’ll walk away with from infrastructure MCP
Section titled “What you’ll walk away with from infrastructure MCP”- A read-only Kubernetes MCP setup: a dedicated ServiceAccount, a TOML file with
read_only = true, and Secrets denied at the server - A copy-paste prompt that takes a
CrashLoopBackOffto a root cause and a manifest pull request - A Terraform MCP setup that starts on the public registry and adds HCP Terraform without enabling apply
- An infrastructure-as-code loop: registry lookup, module, plan in HCP Terraform, review, human apply
- The measured context cost of HashiCorp’s skills plugin, and the traps in each server’s catalog entry and plugin
Which infrastructure MCP server does which job?
Section titled “Which infrastructure MCP server does which job?”The three servers answer different questions, and HashiCorp’s skills cover a fourth: how to write the code, not what the live system says.
| Item | Question it answers | How it runs | Write control | Popularity (2026-09-26) |
|---|---|---|---|---|
Kubernetes MCP Server (containers GitHub org, community) | What is the cluster doing, and why is this workload failing? | Local stdio via npx, uvx, a binary or an image; talks to the API server directly, no kubectl needed | read_only = true in TOML, plus the RBAC of the kubeconfig it uses | ★2.1k; npm kubernetes-mcp-server 0.0.67 |
| Terraform MCP Server (HashiCorp) | What is the current provider or module API, and what did the HCP Terraform plan do? | Docker image hashicorp/terraform-mcp-server, stdio or Streamable HTTP | No token: only the 9 public registry tools; with a token, filter with --toolsets or --tools. Apply, discard, cancel, force-unlock and workspace/project/team deletes need ENABLE_TF_OPERATIONS=true | ★1.5k; plugin terraform@claude-plugins-official 10,280 installs |
| Docker MCP Toolkit / Gateway (Docker) | How do I run many servers in containers, with secrets in a keychain, for every client? | docker mcp gateway run, one profile shared across clients | Per-profile tool allowlists; secrets blocked from tool traffic by default | docker/mcp-gateway ★1.6k; catalog repo ★558 |
| HashiCorp agent skills (HashiCorp) | How should this Terraform be written and tested? | 16 Terraform and 4 Packer skills; a plugin or single skills | Not applicable: skills add instructions, not access | ★875; plugin terraform@hashicorp 1.0.0 |
Stars are from GitHub and the install count from the claude.com plugin directory, both read on 2026-09-26. Stars measure attention on a repository, not use of the server.
Set up the Kubernetes MCP server read-only
Section titled “Set up the Kubernetes MCP server read-only”The server uses whatever kubeconfig it finds. If that is your personal admin context, the agent has your admin rights, and read_only = true is the only thing standing between a confused prompt and resources_delete. Give it its own identity instead, so the cluster’s RBAC enforces the limit even if the server config is wrong.
-
Create a read-only ServiceAccount. This binds the built-in
viewClusterRole in one namespace.viewreads pods, logs, events and Deployments, and it does not read Secrets.Terminal window # Terminal, with your normal admin contextkubectl create namespace mcpkubectl create serviceaccount mcp-viewer -n mcpkubectl create rolebinding mcp-viewer-staging --clusterrole=view \--serviceaccount=mcp:mcp-viewer -n stagingkubectl auth can-i list pods --as=system:serviceaccount:mcp:mcp-viewer -n staging # yeskubectl auth can-i delete pods --as=system:serviceaccount:mcp:mcp-viewer -n staging # nokubectl auth can-i get secrets --as=system:serviceaccount:mcp:mcp-viewer -n staging # noUse
kubectl create clusterrolebinding … --clusterrole=viewonly if the agent must read every namespace. -
Build a dedicated kubeconfig with a short-lived token.
kubectl create token(Kubernetes 1.24+) mints a token that expires on its own.Terminal window TOKEN="$(kubectl create token mcp-viewer -n mcp --duration=8h)"API_SERVER="$(kubectl config view --minify -o jsonpath='{.clusters[0].cluster.server}')"kubectl config view --minify --raw \-o jsonpath='{.clusters[0].cluster.certificate-authority-data}' | base64 -d > /tmp/mcp-ca.crtKCFG="$HOME/.kube/mcp-viewer.kubeconfig"kubectl config --kubeconfig="$KCFG" set-cluster mcp --server="$API_SERVER" \--certificate-authority=/tmp/mcp-ca.crt --embed-certs=truekubectl config --kubeconfig="$KCFG" set-credentials mcp-viewer --token="$TOKEN"kubectl config --kubeconfig="$KCFG" set-context mcp --cluster=mcp --user=mcp-viewer --namespace=stagingkubectl config --kubeconfig="$KCFG" use-context mcpchmod 600 "$KCFG"; rm /tmp/mcp-ca.crtIf your cluster’s kubeconfig uses a
certificate-authorityfile path instead of embedded data, pass that path directly. -
Write the server’s TOML file. TOML is the recommended place for settings. Version 0.0.67 still accepts runtime flags such as
--read-onlyand--toolsets, but the next release removes them (the change is onmainas of 2026-09-26), so keepread_onlyin the file rather than passing--read-only.~/.config/kubernetes-mcp-server.toml read_only = true # only tools annotated readOnlyHint=true are exposedtoolsets = ["core", "config"] # the default pair; add "helm" only if you need helm_listdisabled_tools = ["configuration_view"][[denied_resources]] # belt and braces on top of RBACgroup = ""version = "v1"kind = "Secret"configuration_viewprints the kubeconfig as YAML. Over stdio, as configured here, it is exposed by default, so disable it. -
Register the server. Pin the version rather than
@latest, so an upstream release does not change your tool set mid-incident.Terminal window # Terminal, repo root. Single quotes keep ${HOME} unexpanded; Claude Code expands it at launchclaude mcp add kubernetes -s project \-e 'KUBECONFIG=${HOME}/.kube/mcp-viewer.kubeconfig' \-- npx -y kubernetes-mcp-server@0.0.67 --config '${HOME}/.config/kubernetes-mcp-server.toml'Terminal window # Writes [mcp_servers.kubernetes] to ~/.codex/config.toml; your shell expands $HOME herecodex mcp add kubernetes --env KUBECONFIG=$HOME/.kube/mcp-viewer.kubeconfig \-- npx -y kubernetes-mcp-server@0.0.67 --config $HOME/.config/kubernetes-mcp-server.toml// ~/.cursor/mcp.json — use absolute paths{ "mcpServers": { "kubernetes": {"command": "npx","args": ["-y", "kubernetes-mcp-server@0.0.67", "--config", "/Users/you/.config/kubernetes-mcp-server.toml"],"env": { "KUBECONFIG": "/Users/you/.kube/mcp-viewer.kubeconfig" } } } } -
Check what the agent can see. Run
/mcpin Claude Code or Codex, or open the MCP settings in Cursor. Withread_only = trueyou should see none of the write tools, such aspods_delete,pods_exec,pods_run,resources_create_or_update,resources_deleteorresources_scale.
Diagnose a CrashLoopBackOff with the agent
Section titled “Diagnose a CrashLoopBackOff with the agent”This is the example to try first. It is read-only, and it ends with evidence you can check in two minutes, not a wall of reasoning.
What you should see: calls to pods_list_in_namespace, pods_get, pods_log with previous: true, events_list with a field selector such as involvedObject.name=<pod>, and resources_get for the Deployment and ReplicaSets. The report names one cause with evidence, for example “exit code 137, reason OOMKilled, memory limit lowered from 512Mi to 256Mi in the last rollout”. The diff is to a file in the repository. If the agent cannot tell the causes apart, it should say which extra read would decide it rather than guess.
The fix then goes through the same path as any manifest change: a pull request, CI, and your normal deploy. When the pod needs a fix now, a person runs the rollback (kubectl rollout undo deployment/checkout -n staging) with their own credentials; the agent’s identity cannot, by design.
Set up the Terraform MCP server, registry first
Section titled “Set up the Terraform MCP server, registry first”With no token, the Terraform MCP server registers only the registry toolset: nine tools such as search_providers, get_provider_details, get_latest_provider_version, search_modules and get_module_details. They read the public Terraform Registry: the cheapest cure for an agent that writes provider arguments from memory. The --toolsets flag defaults to all, so once a token is present and you pass no --toolsets, every toolset registers: 56 tools (61 with ENABLE_TF_OPERATIONS=true), including write tools such as create_run, create_workspace and delete_variable_in_variable_set.
Add HCP Terraform or Terraform Enterprise when you want the agent to read workspaces and plans. That needs TFE_TOKEN (and TFE_ADDRESS for Terraform Enterprise) and the terraform toolset, 48 tools covering workspaces, runs, plans, variables and state versions. Five of them stay unregistered unless you set ENABLE_TF_OPERATIONS=true: action_run (apply, discard, cancel), delete_workspace_safely, force_unlock_workspace, delete_project and delete_team. Without that variable, create_run offers only plan_and_apply with auto-apply off, plan_only, refresh_state and allow_empty_apply; the auto_approve and is_destroy run types appear only with it. Leave it unset.
# Terminal, repo root. Writes .mcp.json with ${TFE_TOKEN} left for Claude Code to expand at launchclaude mcp add terraform -s project \ -e 'TFE_TOKEN=${TFE_TOKEN}' -e 'TFE_ADDRESS=${TFE_ADDRESS:-https://app.terraform.io}' \ -- docker run -i --rm -e TFE_TOKEN -e TFE_ADDRESS \ hashicorp/terraform-mcp-server:1.3.0 --toolsets=registry,terraformFor the registry only, drop both -e pairs and the --toolsets flag.
codex mcp add terraform -- docker run -i --rm -e TFE_TOKEN -e TFE_ADDRESS \ hashicorp/terraform-mcp-server:1.3.0 --toolsets=registry,terraformThen forward the two variables from your shell, instead of writing the token into the file:
# ~/.codex/config.toml: inside the [mcp_servers.terraform] table the command createdenv_vars = ["TFE_TOKEN", "TFE_ADDRESS"]Export TFE_ADDRESS=https://app.terraform.io for HCP Terraform before starting Codex.
// .cursor/mcp.json — registry only, no token{ "mcpServers": { "terraform": { "command": "docker", "args": ["run", "-i", "--rm", "hashicorp/terraform-mcp-server:1.3.0"] } } }For HCP Terraform, add "-e", "TFE_TOKEN", "-e", "TFE_ADDRESS" to args before the image name, and "--toolsets=registry,terraform" after it (or --tools=, as below). A bare -e NAME makes Docker copy the variable from the environment Cursor starts from, so export both there and confirm with whoami that the token arrived. Never paste the token into the file.
Give the token the least access that works. In HCP Terraform, a team token whose team holds the Plan permission on the relevant workspaces can queue plans and read runs but cannot apply. Ask the agent to call whoami and get_token_permissions once after setup; the answer is your audit record of what it can do.
Do not install terraform@claude-plugins-official in the same session. It registers a second terraform server pinned to hashicorp/terraform-mcp-server:0.4.0, while the image is at 1.3.0.
Add HashiCorp’s skills for how the code is written
Section titled “Add HashiCorp’s skills for how the code is written”The MCP server tells the agent what the current API is. HashiCorp’s skills tell it how HashiCorp wants Terraform written: terraform-style-guide (file layout, naming, for_each over count), terraform-test, refactor-module, terraform-search-import, terraform-stacks, terraform-policy, azure-verified-modules and nine provider-development skills.
claude plugin marketplace add hashicorp/agent-skillsclaude plugin install terraform@hashicorp# The repository also ships .agents/plugins/marketplace.json, named "hashicorp"codex plugin marketplace add hashicorp/agent-skillscodex plugin add terraform@hashicorp# One skill at a time with the skills CLI; the path is from the repository's SKILLS.mdnpx skills add hashicorp/agent-skills/plugins/terraform/skills/terraform-style-guidenpx skills add hashicorp/agent-skills/plugins/terraform/skills/terraform-testMeasured with claude plugin details terraform@hashicorp (plugin 1.0.0, 2026-09-26): 16 skills, no MCP server, ~2,153 tokens always-on in every session. Each skill costs more when it fires: about 2.6k tokens for terraform-style-guide, 4.1k for terraform-test, 5.6k for refactor-module and 6.8k for provider-resources. If your team writes modules and never providers, install the two individual skills above instead of the bundle; the seven provider-* skills are most of the always-on cost.
Ship an infrastructure change: registry lookup, module, plan, review
Section titled “Ship an infrastructure change: registry lookup, module, plan, review”A worked example: add an S3 lifecycle rule that moves objects to Infrequent Access after 30 days, in a repository whose HCP Terraform workspace is connected to version control. The agent does the lookups and the writing; HCP Terraform produces the plan; a person approves the apply.
-
Look up the current API before writing. This is where the registry toolset earns its place.
-
Write the change as a module, with tests. With the HashiCorp skills installed, name them so the agent loads them rather than improvising.
-
Plan in HCP Terraform. Push the branch and open a pull request. The VCS-connected workspace runs a speculative plan on the pull request. In a CLI-driven workspace, run
terraform planwith thecloudblock. Either way HCP Terraform produces the plan with its variables, credentials and policy sets. -
Review the plan as data. Ask the agent to read the plan through the MCP server and reduce it to what a reviewer must decide.
-
A person applies. The reviewer checks the summary against the plan page, merges, and confirms the apply in HCP Terraform. With
ENABLE_TF_OPERATIONSunset, the agent has no apply tool, and its Plan-permission token could not apply anyway.
The loop is the same in every agent. What differs is how you hold it to read-only:
Run step 1 in plan mode; approve the plan, then leave plan mode for step 2. Filter on the server: remove the earlier entry, then add it again with --toolsets=registry,terraform replaced by --tools= and twelve read tools. The server then registers nothing else, so no write tool is left to deny. This switch works the same in all three clients.
claude mcp remove terraform -s projectclaude mcp add terraform -s project \ -e 'TFE_TOKEN=${TFE_TOKEN}' -e 'TFE_ADDRESS=${TFE_ADDRESS:-https://app.terraform.io}' \ -- docker run -i --rm -e TFE_TOKEN -e TFE_ADDRESS hashicorp/terraform-mcp-server:1.3.0 \ --tools=search_providers,get_provider_details,get_latest_provider_version,search_modules,get_module_details,list_workspaces,list_runs,get_run_details,get_plan_logs,get_plan_json_output,whoami,get_token_permissionsUse the same --tools= server switch as in the Claude Code tab, or filter on the client with an allowlist; Codex exposes only the listed tools. Add the key inside the existing [mcp_servers.terraform] table, above any [mcp_servers.terraform.env] subtable (a second [mcp_servers.terraform] header is invalid TOML):
enabled_tools = ["search_providers", "get_provider_details", "get_latest_provider_version", "search_modules", "get_module_details", "list_workspaces", "list_runs", "get_run_details", "get_plan_logs", "get_plan_json_output", "whoami", "get_token_permissions"]Filter on the server, as in the Claude Code tab: replace --toolsets=… in args with --tools= and the same twelve tools (the server rejects both flags together). Run step 1 in Plan Mode; approve the plan, then leave Plan Mode for step 2 and review the diff before accepting it.
Run servers behind the Docker MCP Toolkit gateway
Section titled “Run servers behind the Docker MCP Toolkit gateway”The Docker MCP gateway runs each catalog server in its own container and presents them to your clients as one server. It keeps secrets in the OS keychain instead of environment variables, verifies signatures of Docker’s mcp/ images by default, blocks secrets from tool traffic by default (--block-secrets), and logs tool calls (--log-calls). It pays off when several clients share several SaaS servers.
# Terminal. Docker Desktop 4.59+ ships the plugin; with Docker CE run: docker mcp feature enable profilesdocker mcp catalog pull mcp/docker-mcp-catalogdocker mcp profile create --name dev-tools \ --server catalog://mcp/docker-mcp-catalog/github-official \ --server catalog://mcp/docker-mcp-catalog/terraformdocker mcp oauth authorize github # or store a token from stdin:printf '%s' "$GITHUB_PAT" | docker mcp secret set github.personal_access_tokendocker mcp profile tools dev-tools --disable-all github-officialdocker mcp profile tools dev-tools --enable github-official.list_issues --enable github-official.pull_request_readdocker mcp gateway run --profile dev-tools --dry-run # validates the profile without listeningConnect clients with docker mcp client connect <client> --profile dev-tools. The command reference on main lists claude-code, codex and cursor among the supported clients (read 2026-09-26); older releases may lack codex, so check docker mcp client ls. Add --global for a user-wide entry instead of the current repository. Profile allowlists name tools as <server>.<tool>; list what the profile really exposes with docker mcp tools ls before you rely on one.
Three catalog entries are not what their names suggest (checked in docker/mcp-registry, 2026-09-26):
-
githubin the Docker catalog is the archived reference server. Its title is “GitHub (Archived)”. Usegithub-official, which runsghcr.io/github/github-mcp-server, even though the gateway README’s own example still namesgithub. -
kubernetesin the Docker catalog is a different server. It runsmcp/kubernetes, built fromFlux159/mcp-server-kubernetes, and mounts the kubeconfig path you give it. Point it at your default~/.kube/configand it holds your admin context. Keep Kubernetes as the direct, read-only entry above, or give the catalog entry themcp-viewerkubeconfig. -
terraformin the Docker catalog declares no secrets. Through the gateway you get the public registry tools only. For HCP Terraform, keep the direct Docker entry withTFE_TOKEN.
For allowlisting the gateway across an organization, see MCP registries and gateways.
How much context do infrastructure servers cost?
Section titled “How much context do infrastructure servers cost?”Every registered tool’s schema sits in the context of every session. Measure before you standardise:
| Install | What loads | Always-on cost |
|---|---|---|
terraform@hashicorp plugin 1.0.0 | 16 skills | ~2,153 tokens (measured, claude plugin details) |
| Terraform MCP, no token | 9 registry tools | run /context before and after claude mcp add |
Terraform MCP, token, no --toolsets | 56 tools (61 with ENABLE_TF_OPERATIONS=true) | avoid; always pass --toolsets or --tools |
Terraform MCP, --toolsets=registry,terraform | 52 tools (57 with ENABLE_TF_OPERATIONS=true) | 52 tool schemas instead of 9; measure with /context |
Terraform MCP, --tools= with the 12 read tools | 12 tools | the allowlist used in the loop above |
Kubernetes MCP, core and config | 23 tools, fewer with read_only = true | measure with /context |
Tool counts are from the servers’ own tool registries on 2026-09-26. Trim with the server’s switch first (--tools for Terraform, enabled_tools or disabled_tools in the Kubernetes TOML), because it works in every client. More techniques are in reducing MCP token cost.
How do you verify infrastructure work the agent did?
Section titled “How do you verify infrastructure work the agent did?”You check evidence, not the agent’s account of it:
- The identity is the limit.
kubectl auth can-ishows what themcp-vieweraccount can do, andget_token_permissionsshows what the Terraform token can do. Both answers go into the pull request once, when the setup is reviewed. - Every diagnosis cites its reads. The CrashLoopBackOff report quotes the exit reason, log lines and event it used. A reviewer checks three facts against
kubectl describe pod, not the reasoning. - Gates run before the plan.
terraform fmt -check,terraform validateandterraform testpass locally and again in CI; manifests pass your CI’s validation. - The plan is the review artifact. The reviewer reads the create, update, delete and replace counts and every delete or replace, against the plan page in HCP Terraform. Any unexpected delete or replace blocks the merge.
- A person applies and rolls back. Apply is confirmed in HCP Terraform by the workspace owner; Kubernetes changes go out through the deploy pipeline. Rollback is the same path in reverse: revert the commit, or
kubectl rollout undoby a person with write access. - The acceptance check runs after deploy. Zero restarts for 15 minutes and all replicas Ready for the pod fix; the lifecycle rule visible on the bucket for the Terraform change.
What breaks with infrastructure MCP servers?
Section titled “What breaks with infrastructure MCP servers?”The Kubernetes server connected yesterday and fails today. The ServiceAccount token from kubectl create token expired. Mint a new one and rerun the set-credentials line; nothing else changes. Keep the duration short on purpose.
The agent still sees pods_delete. The TOML file was not loaded, or a key in it is misspelled. Check the --config path in the entry, that Claude Code expanded ${HOME} (run claude mcp get kubernetes), and that the key reads exactly read_only = true at the top level of the file, not under a table header.
pods_log returns nothing useful for a crashing pod. The current container has just started. Ask for previous: true, which returns the terminated container’s logs, the ones with the crash in them.
Events are missing. Kubernetes keeps events for a limited time, so an old crash has none left. Rely on the previous container’s logs and the ReplicaSet diff, and note the gap in the report.
Terraform tools that need a token are missing. TFE_TOKEN never reached the container. docker run -e TFE_TOKEN forwards the variable only if the Docker process has it: in Claude Code check the env block in .mcp.json, in Codex check env_vars, and export the variable before you start the agent.
Codex refuses to start after you edited config.toml. The error reads invalid type: sequence, expected a string in mcp_servers.<name>.env.env_vars. You appended env_vars or enabled_tools at the end of the file, and it landed in the [mcp_servers.<name>.env] subtable that codex mcp add --env writes. Move the line up into the [mcp_servers.<name>] table itself, then run codex mcp get <name> --json to confirm (checked on codex-cli 0.157.1).
The agent writes arguments that terraform validate rejects. It skipped the registry lookup or read the latest docs against an older pinned provider. Make step 1 a separate turn and name the pinned version in the prompt.
Two terraform servers, old tool names. The official Claude plugin (image 0.4.0) and your own entry are both registered. Keep one: claude plugin uninstall terraform@claude-plugins-official, or remove your entry.
The agent acts on text inside a log line or a plan. Pod logs, events and plan output can carry attacker-controlled strings, which reach the model as data it may follow. Read-only identities limit the damage; also keep other write-capable servers out of the same session, and read MCP security before connecting a production cluster.
Connection problems in general are covered in MCP connection issues.