Debug Production from the Editor: Sentry, Grafana, Datadog, and PostHog MCP
Observability MCP servers let Claude Code, Codex, and Cursor read production evidence directly: Sentry for errors and traces, Grafana and Datadog for metrics, logs, and alerts, and PostHog for product metrics. Connected read-only, they turn “fix the top production error” into a loop that starts from a real stack trace and ends with a metric proving the fix held.
This page is for developers who take incidents and tech leads who own the on-call process. The alert fires, you open Grafana to find the spike, switch to Sentry for the stack trace, copy it into the agent, wait for a fix, deploy, and then check three dashboards to see whether it worked. The agent wrote five lines of code; you spent the better part of an hour as the clipboard between four systems.
What you’ll walk away with from observability MCP
Section titled “What you’ll walk away with from observability MCP”- Verified, read-only install lines for Sentry, Grafana, Datadog, and PostHog in all three agents
- A copy-paste prompt that takes the most frequent unresolved Sentry error to a failing test and a local fix
- An incident loop: alert in Grafana or Datadog, error in Sentry, fix, deploy through CI, verify the metric in PostHog
- The measured context cost of each vendor plugin, and when the bare server is the better install
- The traps: the wrong Grafana package, the Sentry header that is not
Bearer, and write tools that stay on unless you turn them off
Which observability MCP server answers which question?
Section titled “Which observability MCP server answers which question?”Each server answers a different incident question. Install the ones whose data you already have; none of them replaces another.
| Server | Question it answers | Transport and auth | Read-only control | Popularity (2026-09-26) |
|---|---|---|---|---|
| Sentry MCP (Sentry) | Which error, where in the code, since which release? | Remote https://mcp.sentry.dev/mcp, OAuth or a Sentry-Bearer token | Default grant is inspect and seer; triage (resolve, assign) is off unless you grant it | ★862; plugin 38,810 installs |
| Grafana MCP (Grafana Labs) | What did metrics, logs, and alerts do around the incident? | Local uvx mcp-grafana with a service account token, or Grafana Cloud remote | --disable-write | ★3.5k |
| Datadog MCP (Datadog, managed) | Which monitor fired, and what do logs and APM traces show? | Remote, one URL per Datadog site; OAuth, and the plugin also supports key auth | No switch in the README; scope with the plugin’s /ddtoolsets and the Datadog user’s role | datadog-labs/mcp-server ★45 (a docs repo; the server is managed) |
| PostHog MCP (PostHog) | Did users feel it, and did the fix restore the funnel? | Remote https://mcp.posthog.com/mcp, OAuth or a personal API key | Allowlist read tools with ?tools=; no read-only flag in the README | PostHog/posthog monorepo ★39.9k |
Popularity is GitHub stars (GitHub API) and install counts from the Claude Marketplace directory (claude.com/plugins), both read 2026-09-26. Stars measure attention on a repository, not use of the server.
Install the four servers read-only
Section titled “Install the four servers read-only”The URLs and packages are identical in every agent; only the command or config file differs. Put team-wide entries in the project file (.mcp.json, .cursor/mcp.json) so they go through review, and keep every credential in an environment variable or a file outside the repository.
Sentry: remote server, OAuth first
Section titled “Sentry: remote server, OAuth first”# Terminal, repo root. --transport http is required; without it Claude Code writes a stdio entryclaude mcp add -s project --transport http sentry https://mcp.sentry.dev/mcpclaude mcp login sentry # or run /mcp inside a sessionToken instead of OAuth (CI, headless). Single quotes keep the variable unexpanded in .mcp.json; Claude Code expands it at launch:
claude mcp add -s project --transport http sentry https://mcp.sentry.dev/mcp \ -H 'Authorization: Sentry-Bearer ${SENTRY_ACCESS_TOKEN}'codex mcp add sentry --url https://mcp.sentry.dev/mcpcodex mcp login sentryToken instead of OAuth: --bearer-token-env-var sends Bearer, which Sentry reserves for OAuth tokens. Send the full header value from an environment variable instead:
# ~/.codex/config.toml; export SENTRY_AUTH_HEADER="Sentry-Bearer $SENTRY_ACCESS_TOKEN" first[mcp_servers.sentry]url = "https://mcp.sentry.dev/mcp"env_http_headers = { "Authorization" = "SENTRY_AUTH_HEADER" }In .cursor/mcp.json:
{ "mcpServers": { "sentry": { "url": "https://mcp.sentry.dev/mcp" } } }Complete the OAuth sign-in when Cursor asks for it.
On the OAuth screen, leave Triage Issues and Manage Projects & Teams unchecked. The defaults, inspect and seer, read issues, events, traces, replays, and releases and run Seer analysis; update_issue belongs to triage. With a token, the server grants every skill, so narrow it in the URL: https://mcp.sentry.dev/mcp?skills=inspect,seer.
Sentry also ships two plugins, and you pick one of them. sentry@claude-plugins-official (source getsentry/plugin-claude, also codex plugin add sentry@openai-curated, and getsentry/plugin-cursor in Cursor’s plugin settings) bundles the hosted server with the sentry-for-ai skills, among them sentry-debug-issue and sentry-fix-stack-traces. sentry-mcp@sentry-mcp (claude plugin marketplace add getsentry/sentry-mcp first) adds a sentry-mcp subagent instead, which keeps large event payloads out of your main context. Installing both registers the server twice. The skills come from the getsentry/sentry-for-ai source repository, which Sentry says to install through one of these plugins rather than directly.
Grafana: the official server, with --disable-write
Section titled “Grafana: the official server, with --disable-write”The official server is PyPI mcp-grafana (1.6.0), run with uvx. Create a service account with the Viewer role, save its token to a file outside the repository, and point GRAFANA_SERVICE_ACCOUNT_TOKEN_FILE at it. The server rereads the file on every request, so a rotated token needs no restart.
claude mcp add -s project grafana \ -e GRAFANA_URL=https://yourstack.grafana.net \ -e 'GRAFANA_SERVICE_ACCOUNT_TOKEN_FILE=${HOME}/.config/grafana/mcp-token' \ -- uvx mcp-grafana --disable-writecodex mcp add grafana \ --env GRAFANA_URL=https://yourstack.grafana.net \ --env GRAFANA_SERVICE_ACCOUNT_TOKEN_FILE=$HOME/.config/grafana/mcp-token \ -- uvx mcp-grafana --disable-writeIn ~/.cursor/mcp.json:
{ "mcpServers": { "grafana": { "command": "uvx", "args": ["mcp-grafana", "--disable-write"], "env": { "GRAFANA_URL": "https://yourstack.grafana.net", "GRAFANA_SERVICE_ACCOUNT_TOKEN_FILE": "/Users/you/.config/grafana/mcp-token" } } } }This entry goes in the user-level file, not .cursor/mcp.json, because the token path is per user. Replace /Users/you with your own home directory.
--disable-write removes dashboard updates, incident creation, alert-rule and silence changes, annotations, and snapshots. It also removes the raw-SQL datasource tools, query_sql and query_influxdb, because they run any statement the datasource credentials allow; add --enable-query to keep them when those credentials are read-only. It also removes the two Sift finders, find_error_pattern_logs and find_slow_requests, because they create investigation records. To keep them, add --enable-write-tools=find_error_pattern_logs,find_slow_requests; they need the Editor role. On Grafana Cloud you can skip the local process: https://mcp.grafana.com/mcp is the endpoint in Grafana’s grafana-cloud-mcp plugin (claude mcp add --transport http grafana-cloud https://mcp.grafana.com/mcp). The flags on this page belong to the local server, so list what the hosted endpoint exposes with /mcp before you rely on it being read-only.
Datadog: one URL per site
Section titled “Datadog: one URL per site”# US1. For EU replace the host with mcp.datadoghq.eu; other sites use their own hostclaude mcp add -s project --transport http datadog https://mcp.datadoghq.com/v1/mcpOr claude plugin install datadog@claude-plugins-official (preview) and run /ddsetup, which walks you through choosing your Datadog site. Do not install both.
codex mcp add datadog --url https://mcp.datadoghq.com/v1/mcpIn our test, codex mcp add started the OAuth flow straight away for Datadog. The alternative is codex plugin add datadog@openai-curated.
{ "mcpServers": { "datadog": { "url": "https://mcp.datadoghq.com/v1/mcp" } } }Match the MCP host to the site in your Datadog URL: app.datadoghq.eu means mcp.datadoghq.eu. Datadog’s README documents no read-only switch (checked 2026-09-26), so log in as a Datadog user whose role fits the job, and with the plugin turn off tool groups you do not need with /ddtoolsets. For CI or headless sessions, the plugin also accepts key authentication: set DD_MCP_DOMAIN (a host such as mcp.datadoghq.com, without https://), DD_API_KEY, and DD_APPLICATION_KEY before you start Claude Code. DD_MCP_TOOLSETS pins the enabled toolsets as a comma-separated list and overrides whatever /ddtoolsets set, which makes it the one non-interactive way to narrow Datadog’s tools (plugin README, checked 2026-09-26).
PostHog: the wizard, or a scoped URL
Section titled “PostHog: the wizard, or a scoped URL”npx @posthog/wizard@latest mcp add (wizard 2.78.0) writes the entry for Cursor, Claude Code, Claude, VS Code, and Zed. Codex is not on its list. For an incident loop, pin an allowlist of read tools instead of the full tool set:
claude mcp add -s project --transport http posthog \ 'https://mcp.posthog.com/mcp?mode=tools&tools=query-funnel,query-trends,query-error-tracking-issue,insights-list,insight-get,dashboard-get,read-data-schema,execute-sql'codex mcp add posthog --url 'https://mcp.posthog.com/mcp?mode=tools&tools=query-funnel,query-trends,query-error-tracking-issue,insights-list,insight-get,dashboard-get,read-data-schema,execute-sql'codex mcp login posthog{ "mcpServers": { "posthog": { "url": "https://mcp.posthog.com/mcp?mode=tools&tools=query-funnel,query-trends,query-error-tracking-issue,insights-list,insight-get,dashboard-get,read-data-schema,execute-sql" } } }By default the server runs in cli mode, except for Cursor, which the server detects and switches to tools mode. In cli mode one posthog tool wraps every command, so a client-side rule cannot single out update-feature-flag. Pinning mode=tools makes every client behave the same: it registers one MCP tool per PostHog tool, and tools= exposes only the listed ones, all of them reads (execute-sql runs HogQL queries). PostHog documents no read-only switch, so the allowlist is your read-only control. Use ?features=flags,dashboards style filters when you want whole product areas instead. EU projects use the mcp-eu.posthog.com host so that OAuth routes to the EU instance.
How much context do observability servers cost?
Section titled “How much context do observability servers cost?”Measure before you standardise. The plugin rows come from claude plugin details on Claude Code 2.1.283; the Grafana row from an MCP tools/list call to mcp-grafana 1.6.0. All were read on 2026-09-26:
| Install | Context cost | What loads |
|---|---|---|
sentry@claude-plugins-official 1.4.0 | ~1,064 tokens in every session | 8 skills, plus the Sentry server (schemas not counted) |
posthog@posthog 1.1.66 | ~29,443 tokens in every session | 164 skills, one error-analyzer agent, 2 hooks, plus the PostHog server |
uvx mcp-grafana --disable-write 1.6.0 | 75,983 characters of tool schemas, roughly 19,000 tokens if every schema loads | 65 tools; the largest roster on this page |
| Datadog remote server | Not measured | The server is managed and lists its tools only after a Datadog login, which the writing environment did not have |
The PostHog row was measured from PostHog’s own marketplace (claude plugin marketplace add PostHog/ai-plugin, then claude plugin install posthog@posthog); the posthog@claude-plugins-official listing points at the same PostHog/ai-plugin repository. The PostHog plugin’s README says “30+ task-specific skills”; version 1.1.66 installs 164. For an incident loop that needs eight read tools, the allowlisted URL above costs a fraction of that. Keep the plugin for product-analytics work, or enable it only in those sessions. The Grafana figure is a ceiling, estimated at four characters per token rather than read from /context: Claude Code and Codex load MCP tools through tool search, so a schema enters the context only when the agent pulls it. Server schemas are on top of the plugin numbers; run /context before and after each claude mcp add, and do the same for Datadog. In Codex, enabled_tools and disabled_tools on a server entry filter what the model sees. More techniques are in reducing MCP token cost.
Fix the most frequent unresolved error locally
Section titled “Fix the most frequent unresolved error locally”This is the example to try first. It uses Sentry alone, it stays read-only, and it ends with a test that proves the fix. You should see search_issues, then search_sentry_tools/execute_sentry_tool running get_issue_details, get_event_stacktrace, or get_trace_details, and optionally analyze_issue_with_seer. The server exposes the detail tools through that catalog pair rather than by their own names.
What you should see: a short issue summary, a frame-to-file mapping, a new test that fails with the same exception type and message as the Sentry event, then a diff and a green run. If the agent cannot reproduce the failure, it says so and lists what it is missing, often a tag or a feature flag state. It should not guess a fix.
For an unfamiliar error, add Seer as a second opinion:
Run the incident loop: alert, error, fix, deploy, verify
Section titled “Run the incident loop: alert, error, fix, deploy, verify”The example above starts from Sentry. A real incident starts from an alert, and it ends when a metric says users are fine again. The agent does the lookups at each step; you and CI own the gates.
-
Read the alert. The agent pulls the firing alert and the surrounding metrics from Grafana or Datadog, and pins down the incident window.
With Datadog, the same request reads: “Show the monitor that fired for service checkout, the matching APM traces with errors, and what changed in the 30 minutes before.”
-
Find the error. Hand the window to Sentry: “Find issues in checkout-api whose events spike inside that window, sorted by event count, and tell me which release introduced each.” The alert gives the time and the symptom; Sentry gives the code location and the release.
-
Reproduce and fix locally. Run the top-error prompt from the previous section, pinned to that issue. The failing test is the acceptance criterion.
-
Ship through CI, not from the editor. The agent opens a pull request that links the Sentry issue and the alert, and carries the failing-then-passing test. Review follows how to review an agent’s pull request; the service owner approves, and the normal pipeline deploys. The MCP servers stay read-only throughout.
-
Verify the metric. After the deploy, the agent checks three things against thresholds you set before the fix: the Sentry issue has no new events in the new release, the alert condition is back under its threshold in Grafana or Datadog, and the user-facing metric in PostHog is back to baseline.
-
Close it as a person. When all three pass, the on-call engineer resolves the Sentry issue and writes the incident note. The agent can draft the note from the evidence it gathered; the incident process itself is in AI-powered incident response.
The loop is the same in every agent. What differs is how you run it:
Run steps 1 to 3 in plan mode, so you approve the plan before the agent edits files. With sentry-mcp@sentry-mcp, Claude delegates Sentry lookups to the sentry-mcp subagent, and your main context holds only the summary. Deny write tools you never want called, in .claude/settings.json:
{ "permissions": { "deny": ["mcp__sentry__update_issue", "mcp__posthog__update-feature-flag"] } }Treat this as a backstop. The real read-only control for Sentry is the OAuth grant or ?skills=inspect,seer, because execute_sentry_tool reaches catalog tools by name, and a deny on update_issue alone may not cover a write sent through it.
Run the loop in the CLI or the app with the servers connected. Filter write tools on the client with disabled_tools = ["update_issue"] in the [mcp_servers.sentry] entry, but treat it as a backstop: the real read-only control is the OAuth grant or ?skills=inspect,seer, because execute_sentry_tool reaches catalog tools by name and a filter on update_issue alone may not cover it. codex review --base main can review the fix branch from step 4 as a second model.
Use Plan Mode for steps 1 to 3 and review the diff before you accept it. As in every client, the controls that hold are server-side: Sentry’s OAuth skills or ?skills=inspect,seer, --disable-write, and PostHog’s tools allowlist. Client-side filters are a backstop.
How do you know the fix worked?
Section titled “How do you know the fix worked?”You check evidence, not the agent’s summary:
- A reproduction test exists, and it failed first. The test must fail on the old code with the same exception as the Sentry event. A test that only passes proves nothing about this incident.
- The pull request links its sources. The Sentry issue ID, the alert, and the incident window are in the description, so a reviewer checks the fix against the evidence and does not have to re-derive it.
- CI gates stay required. Tests, type checks, and lint run before merge, as for any change. The observability servers add evidence; they do not replace gates.
- Thresholds are set before the deploy. Write the pass condition (zero events in the new release, error rate under target, funnel within tolerance) into the pull request, and check it in step 5. Setting the bar after seeing the numbers is how a partial fix gets called done.
- A person resolves. The on-call engineer resolves the Sentry issue and signs off the incident. Keeping
triageoff makes this structural, not a convention. - Rollback is ready. If step 5 fails, revert through the normal pipeline, then reopen the loop at step 2.
What breaks with observability MCP, and how do you recover?
Section titled “What breaks with observability MCP, and how do you recover?”claude mcp add sentry https://mcp.sentry.dev/mcp connects to nothing. Without --transport http, Claude Code 2.1.283 accepts the line and writes a stdio server whose command is the URL. Remove it with claude mcp remove sentry -s project and add it again with the flag.
Sentry rejects your token. Check whether you sent Authorization: Bearer. Sentry reserves Bearer for MCP OAuth tokens; direct tokens use Sentry-Bearer. In Codex, this means env_http_headers, not --bearer-token-env-var.
Grafana starts, but tool calls fail. GRAFANA_URL points at the wrong stack, or the service account lacks the role. Viewer covers dashboards, datasource queries, and incidents; the Sift finders need Editor. GRAFANA_API_KEY is deprecated; use the service account token.
find_error_pattern_logs is missing. --disable-write removed it. Add --enable-write-tools=find_error_pattern_logs,find_slow_requests if you accept Sift creating investigation records.
Loki queries time out or return a flood. A broad selector over a long range scans a lot of data. Ask for a narrow stream selector and a short window, and set --loki-guardrail-mode=enforce so the server rejects unbounded queries and tells the agent how to narrow them.
Datadog login fails or queries return nothing. Check the site first: the MCP host must match your Datadog site. Fix the URL, then log in again. With the plugin, /ddsetup walks you through setting the site and /ddconfig checks the site, authentication, and network access.
The session is slow and the agent forgets instructions. Tool schemas and plugin skills fill the context. Check /context; replace the PostHog plugin with the allowlisted URL, and use the Sentry subagent plugin.
The agent acts on text inside an error message. Event messages, breadcrumbs, and log lines contain user input, and they reach the model as data it may treat as instructions. Keep write tools off, keep production sessions free of other write-capable servers, and read MCP security before connecting production. Events can also hold personal data, which then sits in the session transcript. Scrub it in the SDK before it reaches Sentry.
Connection failures in general are covered in MCP connection issues.