Internal MCP servers — build only the missing interface
An internal MCP server is a Model Context Protocol service your organization builds so coding agents can reach an internal system that no maintained integration exposes safely. It earns the maximum CTO Scorecard Q7 answer only when it closes a measured, recurring access gap through narrow read-first tools, with an owner, least-privilege authorization, contract tests, monitoring, and retirement criteria.
This page is for the CTO who decides what the platform team builds and the tech lead who will own the result. The situation: engineers paste the same service-ownership spreadsheet, feature-flag state, and runbook links into agent sessions every day, and someone proposes “an MCP server for everything internal”. Six months later there are four servers, one has a generic SQL tool, two have no owner, and nobody can say whether any of them made a task faster.
What a justified internal MCP service gives you
Section titled “What a justified internal MCP service gives you”- A four-level Q7 scoring table, so you can place the organization and name the evidence that moves it up.
- A decision table that picks the smallest interface for a gap: an existing server, a CLI plus a skill, a generated file, or a custom server.
- A gap brief and a tool contract template you can adopt as-is.
- Verified commands to connect one internal server in Claude Code, Codex, and Cursor.
- Automated gates that prove the server works in every client and every protocol era, without anyone reading each tool call.
- The metrics and retirement rule that keep the server count honest, and the failure modes with recovery steps.
How does CTO Scorecard Q7 score internal MCP servers?
Section titled “How does CTO Scorecard Q7 score internal MCP servers?”The scorecard asks “Do internal MCP services solve measured recurring access gaps safely?” The number of servers never scores. Evidence does.
| Points | Answer | Evidence that earns it |
|---|---|---|
| 0 | No | None, or servers nobody can list. |
| 1 | A proof of concept exists, but the gap, owner, or controls are not established | A repository and a demo. No baseline, no owner on the service catalog, a shared token. |
| 2 | A justified narrow service is in use with basic ownership and monitoring | A gap brief with a baseline, a named owner, a dashboard with calls and errors. |
| 3 | Justified internal MCP services with owners, least privilege, docs, fixtures, monitoring, and retirement criteria | Everything in 2, plus a tool contract per tool, per-user authorization, a contract-test run in CI, a fixture suite, a quarterly review, and a written retirement rule. |
Level 3 is reachable with one server. An organization with one well-run server scores higher than one with six unowned servers.
When is an internal MCP server the right interface?
Section titled “When is an internal MCP server the right interface?”Most access gaps have a cheaper fix than a new service. An MCP server adds an authenticated endpoint, a schema that every client loads into context, protocol upgrades, and on-call duty. Work down this table and stop at the first row that fits.
| The gap looks like… | Smallest interface | Why it wins |
|---|---|---|
| A public or vendor system that already has a maintained MCP server (GitHub, Sentry, a database) | The existing server, configured read-only | Someone else maintains it. See essential MCP servers. |
| An internal system with a good CLI or REST API, used by one team | The CLI plus a shared skill that documents how to call it | No new service. The agent already runs commands. |
| Slow-changing reference data (ownership, architecture decisions, API catalog) | A generated file in the repository, refreshed by CI | Zero runtime surface, reviewable in diffs. |
| Live internal data, needed by several teams, from more than one agent client, with per-user permissions | An internal MCP server | One governed interface that Claude Code, Codex, and Cursor all speak. |
Build the server when all four conditions in the last row hold: the data is live, the demand spans teams, more than one client needs it, and access depends on who is asking. If only one holds, a skill or a generated file is usually enough.
Build an internal MCP server in seven steps
Section titled “Build an internal MCP server in seven steps”The order matters. Each step produces evidence the next one depends on, and skipping the first two is the most common way to score 1.
-
Write the gap brief and measure the baseline. For two weeks, record the task, who does it, how often, how long it takes, and how it fails today. Use the gap brief template below. Without a baseline you cannot prove the server helped, and Q7 asks for “measured”.
-
Try the simpler options first. Walk the decision table above with the gap brief in hand. Record which option you rejected and why, because the quarterly review will ask.
-
Write one tool contract per tool. Design task-level operations, not raw access.
find_service_owner(endpoint)returning an owner, a repository, and a last-verified timestamp is a bounded business operation.query_internal_database(sql)hands every connected agent an open-ended authority surface. Start with read-only tools and cap every result size. -
Implement the smallest server that satisfies the contracts. Wrap the existing internal API rather than the database behind it, so the API’s authorization still applies. Building your own MCP server covers the SDK code, transports, and packaging. Serve it over Streamable HTTP so there is one deployment, not one process per laptop.
-
Add the production controls. Per-user OAuth instead of a shared service token, per-tool authorization, audit logs that name the user and the tool, rate limits, timeouts, and no write tool until the read tools have proven their value. The MCP security model (Q8) defines the allowlist, write-approval rules, and adversarial fixtures that apply to this server like any other.
-
Connect it in every supported client and add it to the allowlist. Use the per-tool commands in the next section, then add the server’s exact URL to your managed allowlist. Managed policy covers how to distribute that configuration.
-
Graduate on evidence. The server leaves pilot only when the contract gates pass in CI and the metrics beat the baseline from step 1. Otherwise narrow it or retire it.
Gap brief and tool contract templates
Section titled “Gap brief and tool contract templates”Keep both files in the server’s repository, next to the code. The gap brief is the “measured” in Q7; the contract is what reviewers approve instead of reading the implementation.
gap: Agents cannot find the owning team and runbook for an internal endpointusers: [payments, checkout, platform] # teams that hit the gapfrequency_per_week: 40 # counted from 2 weeks of logged sessionscurrent_path: Search the ownership sheet, then ask in #platform-helpcurrent_lead_time_minutes_p50: 12 # measured, not estimatedfailure_modes: [stale sheet, wrong team paged, runbook link missing]alternatives_rejected: existing_server: none exposes our service catalog cli_plus_skill: catalog API needs per-user auth the CLI lacks generated_file: ownership changes daily; a file is stale within hourssuccess_metric: p50 lead time under 2 minutes, wrong-owner rate under 2%owner: platform-team (on-call rota "catalog")review_date: 2026-12-15retire_if: fewer than 10 distinct users in a quarter, or success metric missed twicename: find_service_ownerpurpose: Return the owning team, repository, and runbook for one internal endpointnon_purpose: Listing all services, editing ownership, paging anyoneinput: { endpoint: "string, path such as /v1/invoices, max 200 chars" }output: { team: string, repository: string, runbook_url: string, last_verified: "ISO 8601" }data_classification: internalauthorization: caller's own OAuth identity; catalog API enforces team visibilityside_effects: none # read-only; annotated readOnlyHintlimits: { timeout_ms: 5000, max_results: 1, rate_per_user_per_minute: 30 }errors: [NOT_FOUND, FORBIDDEN, UPSTREAM_TIMEOUT] # agent reports these, never guessesversion: 1.2.0owner: platform-teamdeprecation: announce two releases ahead; old version served for 30 daysThe numbers in the gap brief are placeholders for your own measurements. The retire_if line is the retirement criterion Q7 asks for, written before launch so nobody has to argue for it later.
How do you connect an internal MCP server in Claude Code, Codex, and Cursor?
Section titled “How do you connect an internal MCP server in Claude Code, Codex, and Cursor?”The server is the same; the registration differs per client. Commands checked against Claude Code 2.1.283 and Codex 0.157.1 on 2026-09-26. Each user authenticates as themselves, so no token appears in any file.
Run in a terminal at the repository root. Project scope writes .mcp.json, which you commit so every engineer gets the same entry:
claude mcp add --transport http --scope project catalog https://catalog.mcp.internal.example.com/mcpclaude mcp login catalogThe written entry is {"mcpServers": {"catalog": {"type": "http", "url": "https://catalog.mcp.internal.example.com/mcp"}}}. Tools appear as mcp__catalog__find_service_owner, which is the name you use in permission rules and in --allowedTools. Each engineer approves the project server once on first use; until then claude mcp list shows it as pending approval (Claude Code 2.1.283).
Run in a terminal. codex mcp add writes to the user’s ~/.codex/config.toml:
codex mcp add catalog --url https://catalog.mcp.internal.example.com/mcpcodex mcp login catalogLimit the tool list in the same entry, so a new tool on the server does not reach agents until you add it here:
[mcp_servers.catalog]url = "https://catalog.mcp.internal.example.com/mcp"enabled_tools = ["find_service_owner", "get_runbook"]Add the server to .cursor/mcp.json in the repository and commit it:
{ "mcpServers": { "catalog": { "url": "https://catalog.mcp.internal.example.com/mcp" } }}Cursor’s MCP documentation lists OAuth for remote servers; confirm the login flow on one machine before rollout.
How do you prove the internal server works without reading every call?
Section titled “How do you prove the internal server works without reading every call?”Nobody reviews individual tool calls. Four automated gates and one quarterly sign-off carry the proof.
1. Contract gates in CI, on every server change. The MCP Inspector CLI (@modelcontextprotocol/inspector 2.8.0 on 2026-09-26) lists the tools in each protocol era and checks schema portability. Run it against a staging deployment:
URL=https://catalog.staging.mcp.internal.example.com/mcpnpx @modelcontextprotocol/inspector@2.8.0 --cli "$URL" --protocol-era legacy --method tools/list --strictnpx @modelcontextprotocol/inspector@2.8.0 --cli "$URL" --protocol-era modern --method tools/list --strictnpx @modelcontextprotocol/inspector@2.8.0 --cli "$URL" --method tools/call \ --tool-name find_service_owner --tool-arg endpoint=/v1/invoicesA server that does not offer 2026-07-28 fails the modern call with a non-zero exit, and --strict exits 6 on an error-severity schema portability problem. Both stop the pipeline. Add the staging token through the Inspector’s OAuth options (--stored-auth-only for non-interactive runs), never as a literal header.
2. The fixture suite. Encode each contract’s error list as tests: an unknown endpoint returns NOT_FOUND, a user outside the team gets FORBIDDEN, a slow upstream returns UPSTREAM_TIMEOUT within the limit, and an oversized input is rejected. The adversarial cases (injected instructions in returned data, replayed writes, audit correlation) come from the Q8 fixture suite.
3. A golden-task eval. Keep 10 to 20 real questions from the gap brief with known answers, and run them headless with only the internal server loaded. eval/catalog.mcp.json holds the same catalog entry as .mcp.json, and --strict-mcp-config ignores every other configured server:
claude -p "Who owns /v1/invoices and where is its runbook? Answer as JSON with team and runbook_url." \ --strict-mcp-config --mcp-config eval/catalog.mcp.json \ --allowedTools "mcp__catalog__find_service_owner" --output-format jsonCompare each answer with the expected one in a script. In Codex, isolate the run the same way: --ignore-user-config skips $CODEX_HOME/config.toml (and every server configured there), and the -c overrides define catalog as the only MCP server for this run. Run it from a directory without a project .codex/config.toml, or that file’s servers join the run (Codex 0.157.1):
codex exec --sandbox read-only --ignore-user-config \ -c 'mcp_servers.catalog.url="https://catalog.mcp.internal.example.com/mcp"' \ -c 'mcp_servers.catalog.enabled_tools=["find_service_owner"]' \ "Who owns /v1/invoices and where is its runbook? Answer as JSON with team and runbook_url."Cursor has no headless eval path covered on this page, so the Claude Code and Codex evals stand in for it: they exercise the same server and the same tool contracts. A drop in the pass rate after a server release blocks the release.
4. Production signals. The owner’s dashboard shows calls per week, distinct users, error rate, denied calls, p95 latency, and audit-log completeness, and alerts on any of them leaving its range.
Quarterly sign-off. The owner compares the metrics with the gap brief’s success_metric, applies retire_if, and records the decision: keep, narrow, or retire. The CTO or the platform lead countersigns. That record is the Q7 level-3 evidence.
Copy-paste prompts for an internal MCP design review
Section titled “Copy-paste prompts for an internal MCP design review”These work the same way in Claude Code, Codex, and Cursor. Run them with the server’s repository open.
What breaks with internal MCP servers, and how do you recover?
Section titled “What breaks with internal MCP servers, and how do you recover?”A platform without a need. Servers were built to satisfy a maturity checklist and nobody uses them, but they still need auth, protocol upgrades, and on-call. Recovery: write a gap brief for each server retroactively; retire any server that cannot show a baseline and current users.
The generic tool. A run_query or call_api tool made the server useful for everything and safe for nothing. Recovery: read the audit log for the ten most common calls, turn each into a task-level tool, then remove the generic one and announce the date.
Works in Claude Code, fails in Codex. The server implements only the 2026-07-28 era, or only the older one after a client upgrades. Recovery: add both --protocol-era checks to CI and serve both eras until every supported client negotiates the new one.
A shared service token. Every user gets the token’s authority, and the audit log shows one identity. Recovery: switch to per-user OAuth, confirm the log names individual users, then revoke the token. Agent identity, credentials and secrets covers the options.
Tool sprawl eats context. Every tool’s schema reaches the agent’s context, and a server with 40 tools crowds out the task. Recovery: keep the tool list short, restrict it per client with enabled_tools or permission rules, and see reducing MCP token cost.
A schema change breaks clients silently. A renamed field returns undefined, and the agent guesses an answer. Recovery: version each contract, keep the golden-task eval in the release pipeline, and serve the old version through the deprecation window in the contract.
The owner left. A reorganization removed the team, and the server keeps running with nobody on call. Recovery: make the owner field a required entry in the service catalog, alert when it points at a disbanded team, and apply retire_if at the next review.
Where to go next with internal MCP servers
Section titled “Where to go next with internal MCP servers”- MCP security model (Q8): the allowlist, per-tool approval, and adversarial fixtures this server must pass.
- Building your own MCP server: the SDK code, transports, and packaging behind step 4.
- MCP 2026-07-28: migrate a server and test both protocol eras.
- MCP registries and gateways: publish the approved catalog and audit calls through a gateway.
- Shared skills (the lighter alternative): when a CLI plus a skill closes the gap without a service.
- Platform team: who owns internal agent tooling and how it is funded.
- Knowledge sharing (Q20): keep contracts, fixtures, and incidents discoverable.
- CTO answer key: every scorecard question and its canonical page.