Measuring Agentic Engineering: Telemetry, Cost, Review, Security and Evals
Measuring agentic engineering means answering four questions every week: what the agents cost, what they shipped, what review caught, and what escaped to production. Three layers of tooling answer them: first-party telemetry built into Claude Code and Codex, local usage trackers such as ccusage, and eval, trace and gateway platforms. First-party telemetry is the layer to switch on first.
Your team has run Claude Code, Codex or Cursor for a quarter. The invoice grew, merged pull requests grew, and nobody can say whether the extra changes are good ones, because the only data is a billing total and a few Slack threads. This page is for the developer who has to switch the measurement on and the tech lead who runs the weekly review; a CTO can read the “switch on first” order as a rollout plan.
What this measurement setup gives you
Section titled “What this measurement setup gives you”- A map of the three measurement layers and the job each one does, so you stop buying a platform for a question a built-in setting already answers.
- The 15 observability and quality tools that matter in September 2026, with verified install lines, dated popularity and the page that covers each in depth.
- A switch-on order for a team or an organization: what to enable in week one and what to leave for later.
- A telemetry smoke test for Claude Code and Codex that proves data flows before you build a dashboard, and the route for Cursor.
- A weekly agent report (spend, pull requests, review findings, escaped defects) with a template, three copy-paste prompts, and the traps that make the numbers lie.
Why measure agent work at all?
Section titled “Why measure agent work at all?”Agents change the volume of code faster than they change its quality, and measurement is how you tell the two apart. The 2025 DORA report puts it plainly: “AI doesn’t fix a team; it amplifies what’s already there” (Google Cloud, DORA 2025 report announcement, 2025-09-23). A weekly report is the cheapest way to see which of those your team is.
The report does not replace your delivery metrics. The metric definitions this site uses (accepted change rate, escaped defects per 100 merged changes, cost per accepted change and seven more) live on the canonical metrics page. This page covers the tools that produce the raw numbers.
What are the three layers of agent measurement?
Section titled “What are the three layers of agent measurement?”Each layer answers different questions and has a different owner. Most teams need the first layer in full, the second for individuals, and only parts of the third.
| Layer | Answers | Examples | Who runs it | Cost to run |
|---|---|---|---|---|
| 1. First-party telemetry | Sessions, tokens, estimated cost, lines, commits and pull requests per repository | Claude Code OpenTelemetry and analytics dashboard; Codex [otel]; Cursor admin analytics | Platform team or tech lead, set once in managed config | A collector you already run, or a hosted OTel backend |
| 2. Local usage trackers | “What did I spend today, and how close am I to my plan limit?” | ccusage, Claude-Code-Usage-Monitor, tokscale, /usage in Claude Code | Each developer | Free; reads local logs, nothing leaves the machine |
| 3. Eval, trace and gateway platforms | Did a CLAUDE.md change make the agent better? Which team spent what? Why did this session go wrong? | Langfuse, promptfoo, Inspect SWE, Braintrust; LiteLLM and Cloudflare AI Gateway for per-team spend; claude gateway for SSO and per-group model access | Platform team | A platform to host or buy; a gateway sits in the request path |
Review bots and security scanners sit beside the three layers as gates: they do not measure usage, they produce findings. Their output is what the “review findings” line of the weekly report counts.
Which observability and quality tools matter in September 2026?
Section titled “Which observability and quality tools matter in September 2026?”The table ranks tools by relevance to a team running Claude Code, Codex or Cursor, weighted by adoption. Popularity as of 2026-09-26: GitHub stars from the GitHub API and package versions from npm and PyPI, read that day. Stars measure attention on a repository, not use. Download counts are omitted.
| # | Tool | Measures or gates | Install or enable | Popularity (2026-09-26) | In depth |
|---|---|---|---|---|---|
| 1 | Claude Code OpenTelemetry | Cost, tokens, lines, commits, pull requests | CLAUDE_CODE_ENABLE_TELEMETRY=1 plus OTEL_METRICS_EXPORTER and OTEL_LOGS_EXPORTER | ships in 2.1.283 | Telemetry |
| 2 | Claude Code analytics dashboard, /usage, /insights | Adoption, pull-request attribution; one person’s usage | Team and Enterprise plan feature; /usage in a session | vendor feature | Telemetry |
| 3 | ccusage | Cost from about 18 agents’ local logs, 5-hour blocks | npx ccusage@latest | 18.7k stars; npm 20.0.24 | Cost tracking |
| 4 | Codex OpenTelemetry | Logs, metrics and traces | [otel] table in ~/.codex/config.toml | ships in 0.157.1 | Telemetry |
| 5 | Claude Code Review, /code-review, claude-code-action | Pull request findings | @claude review on a PR; /code-review in a session; anthropics/claude-code-action@v1 | action 9.0k stars | Review bots |
| 6 | codex review, openai/codex-action | Diff findings, locally or in CI | codex review --base main; codex review "Focus on auth" | action 1.2k stars | Review bots |
| 7 | Langfuse and its Claude Code plugin | A trace per Claude Code session | claude plugin marketplace add langfuse/Claude-Observability-Plugin → claude plugin install langfuse-observability@langfuse-observability | 35.1k stars (plugin repo: 24) | Evals |
| 8 | promptfoo | Evals and red-teaming; runs Claude Agent SDK or Codex SDK as the system under test | npx promptfoo@latest init --example getting-started | 25.5k stars; npm 0.123.1 | Evals |
| 9 | LiteLLM proxy | Spend and budgets per virtual key | uv tool install 'litellm[proxy]', pinned and hash-checked | 59.6k stars; PyPI 1.102.1 | Cost tracking |
| 10 | Snyk Agent Scan | MCP configs and skills scanned for injection | uvx snyk-agent-scan@latest | 3.1k stars; PyPI 0.6.4 | Security gates |
| 11 | Semgrep Guardian, semgrep mcp | SAST on every file the agent writes | Claude Code /plugin, then Discover, then Semgrep | 16.8k stars; PyPI 1.178.0 | Security gates |
| 12 | Gitleaks | Secrets in commits | brew install gitleaks, then a pre-commit hook | 29.5k stars | Security gates |
| 13 | TruffleHog | Secrets, verified against the issuer | brew install trufflehog | 28.1k stars | Security gates |
| 14 | anthropics/claude-code-security-review | Security review comments on pull requests | GitHub Action | 6.3k stars | Security gates |
| 15 | Entire CLI | Which prompts and session produced each commit | brew install --cask entireio/tap/entire | 5.1k stars; v0.11.3 (2026-09-25) | Telemetry |
Row 6 takes custom instructions only without a target flag: codex-cli 0.157.1 rejects a prompt combined with --base, --uncommitted or --commit.
Also verified and covered on the deeper pages: Braintrust, Opik, Arize Phoenix, Inspect AI with Inspect SWE, DeepEval, garak, Portkey, Cloudflare AI Gateway, Claude-Code-Usage-Monitor, tokscale, SonarQube MCP and Snyk MCP. Do not start a new project on Helicone: it has been in maintenance mode since March 2026 (secondary sources: the Helicone and Mintlify acquisition posts).
What should a CTO switch on first?
Section titled “What should a CTO switch on first?”Switch on the layer that answers the most expensive open question first. For almost every organization that is “what are we paying, and what did it produce?”, so the order below starts with first-party telemetry and leaves evals for later.
-
Week 1: first-party telemetry, through managed configuration. Put the Claude Code OpenTelemetry variables in managed settings (see managed policy) and ship a Codex
[otel]block with your standardconfig.toml. Developers cannot forget a setting they never had to set. The organization-wide design, with collector pipelines, join keys and retention, is on agent observability. -
Week 1: a pull request marker. On Claude Team or Enterprise, install the Claude GitHub app and enable GitHub analytics so merged pull requests get the
claude-code-assistedlabel, and add anagent-assistedcheckbox to the pull request template for Codex and Cursor. Without a marker, every later metric compares nothing with nothing. -
Week 2: deterministic security gates. Gitleaks as a pre-commit hook and a SAST scan in CI. Agents commit often and never tire, so these gates fire more often than they did for human-only teams, and they cost no tokens.
-
Week 3: one review bot, measured. Start with the bot bundled with your main tool (
/code-reviewor Claude Code Review,codex review, or Bugbot), and record which findings lead to a change. Anthropic’s docs put the managed Claude Code Review (research preview on Team and Enterprise; not available under Zero Data Retention) at an average of $15–25 per review (checked 2026-09-26), so measure it before you roll it out to every repository. -
Week 4: cost attribution. ccusage for individuals. A gateway with virtual keys (LiteLLM or Cloudflare AI Gateway) only when finance needs per-team budgets that OpenTelemetry cannot enforce.
claude gatewayis a separate thing: Anthropic’s self-hosted auth and telemetry gateway for SSO and per-group model access. Cost per accepted change counts seat licences as well as usage and tokens, as metric 8 on the metrics frameworks page does. -
Later: evals. Build an eval set once you change shared
CLAUDE.mdorAGENTS.mdrules, skills or models often enough that “it felt better” is no longer an answer. Claude Code 2.1.283 also shipsclaude plugin eval, which runs a plugin’s eval cases against a no-plugin baseline.
Prove telemetry works with a smoke test
Section titled “Prove telemetry works with a smoke test”The quickest proof that telemetry works is a console exporter: you see a datapoint in your own terminal before any collector exists. The OTLP export, the collector stack and the dashboard panels are on OpenTelemetry and analytics for Claude Code, Codex and Cursor; this section only proves the pipe is open.
In the terminal you start Claude Code from:
# Smoke test: print metrics to this terminal every 10 secondsexport CLAUDE_CODE_ENABLE_TELEMETRY=1export OTEL_METRICS_EXPORTER=consoleexport OTEL_METRIC_EXPORT_INTERVAL=10000claudeWithin one export interval, the terminal prints a claude_code.session.count datapoint. For a team, the exporter variables belong in the env block of managed settings: Claude Code ignores them in a repository’s .claude/settings.json, so a repository cannot redirect your telemetry.
Codex has no console exporter, so the smoke test is a local OTLP collector and an [otel] table in ~/.codex/config.toml (field names from the Codex source at 0.157.1):
[otel]log_user_prompt = falsemetrics_exporter = { otlp-http = { endpoint = "http://localhost:4318/v1/metrics", protocol = "binary" } }Set metrics_exporter explicitly. In 0.157.1 an unset metrics_exporter defaults to statsig, a Codex-internal destination, not off. Start Codex once with codex --strict-config, which exits on any config.toml field that version does not recognize, so a misspelled [otel] key cannot fail in silence. Test a headless codex exec run separately: a reported issue (openai/codex#12913, status unverified on 0.157.1) says it emits no OpenTelemetry metrics. The collector setup is on the telemetry page.
Cursor documents no OpenTelemetry export that we could confirm as of 2026-09-26. Two routes give you data anyway:
# Per-commit provenance: which prompts and session produced each commitbrew install --cask entireio/tap/entirecd your-repoentire enable --agent cursor --telemetry=falseentire statusCursor’s admin dashboard and Analytics API on Teams and Enterprise, plus the Enterprise AI Code Tracking API, are the other route. Those facts are secondary (search extracts of Cursor’s docs), so confirm them in your admin console before you build on them.
Run a weekly agent report: spend, PRs, review findings, escaped defects
Section titled “Run a weekly agent report: spend, PRs, review findings, escaped defects”The weekly report is where the layers meet. Once the sources exist, an agent can draft it and the tech lead reviews it. Report per repository and per cohort (agent-assisted against everything else), never per engineer.
-
Spend. Take estimated cost from
claude_code.cost.usagefor Claude Code. For Codex, take estimated cost fromcodex.turn.cost_microusd(micro-USD per turn, codex-cli 0.157.1), or runnpx ccusage@latest weekly --json --by-agentand take the Codex rows (ccusage 20.0.24;ccusage codexhas no weekly view). ccusage reads only the local machine’s logs, so it gives per-developer figures, not team totals. For Cursor, take usage from the admin dashboard or Analytics API (Teams and Enterprise; secondary, confirm in your console). A gateway, if you run one, reports spend for every tool behind it. Treat all of these as estimates and reconcile with the invoice monthly. -
Pull requests. Count merged pull requests with an agent marker, and the share of all merged pull requests they represent. With the GitHub CLI:
Terminal window SINCE=2026-09-19 # first day of the reporting weekgh pr list --state merged --label claude-code-assisted \--search "merged:>=$SINCE" --json number,title,url --limit 200gh pr list --state merged --label agent-assisted \--search "merged:>=$SINCE" --json number,title,url --limit 200gh pr list --state merged \--search "merged:>=$SINCE" --json number --limit 500 --jq lengthThe last command counts all merged pull requests, the denominator for the share.
-
Review findings. Count findings each bot and each gate posted, and how many led to a code change. That acted-on rate is the precision proxy that decides whether a bot keeps its budget. Record findings from deterministic gates (Gitleaks, Semgrep) separately from model reviewers.
-
Escaped defects. Count production bugs opened this week that trace back to a merged change, per 100 merged changes, split by cohort. This is metric 7 on the canonical metrics page. Every escaped defect in the agent cohort gets a one-line note on which gate should have caught it.
-
Decide one thing. Every report ends with one change to the harness: a new test, a rule in
CLAUDE.mdorAGENTS.md, a scanner rule, or a review-bot setting. A report that changes nothing is a vanity dashboard.
Save this template as docs/agent-report/TEMPLATE.md in the repository that holds your team’s runbooks:
# Agent report — week of YYYY-MM-DD (repository: NAME)
| Line | Agent-assisted | Other | Source || --- | --- | --- | --- || Merged PRs | | | gh pr list, label claude-code-assisted / agent-assisted || Estimated spend (USD) | | n/a | claude_code.cost.usage / codex.turn.cost_microusd / ccusage / gateway || Cost per accepted change (USD) | | n/a | (estimated spend + seat cost) ÷ agent-assisted PRs merged 14–21 days ago and not reverted or fixed since || Review findings posted / acted on | | | bot comments, gate logs || Escaped defects per 100 merged changes | | | bug tracker, linked PRs |
Escaped defects this week: PR, bug, which gate should have caught it.One harness change we make next week: ...Owner: tech lead. Sign-off: engineering manager.Copy-paste prompts for the weekly agent report
Section titled “Copy-paste prompts for the weekly agent report”Run these in Claude Code, Codex or Cursor from a checkout of the repository, with the GitHub CLI signed in with a read-only token. The prompts are the same in all three tools.
BUG_NUMBER is the issue number of the escaped bug.
How do you know the measurement itself is right?
Section titled “How do you know the measurement itself is right?”A measurement system is code, and it fails the way code does: quietly. Five checks keep it honest, and the tech lead signs off on each before the numbers reach a leadership slide.
- A known event produces a known datapoint. Start one session and confirm the session counter moves by one in your backend. Repeat after every collector or Claude Code upgrade.
- Spend reconciles. Once a month, compare OpenTelemetry or gateway cost with the invoice. A gap above your tolerance means an unexported team, a second billing path, or a wrong price table.
- Attribution is spot-checked. Pick five merged pull requests at random and check the label against the session history or the Entire checkpoint trailer.
- Cohorts, not people. No panel groups by
user_idor email. Per-person panels turn the report into surveillance and the metric into a target. - Speed never appears without stability. Merged pull requests and spend always sit beside escaped defects or the 14-day follow-up fix rate.
What breaks when you measure agentic engineering?
Section titled “What breaks when you measure agentic engineering?”| Symptom | Cause | Recovery |
|---|---|---|
| Agent-assisted share jumps sharply in one week | A new marker (the GitHub app, a template checkbox) started counting, not a change in behavior | Annotate the dashboard with the date each marker went live; compare only periods with the same markers |
| Review bot findings rise, nobody acts on them | The bot is tuned for recall; people resolve threads unread | Track acted-on rate per bot; tune it (a REVIEW.md for Claude Code Review) or drop it |
| Spend is flat but the invoice rose | A second billing path: API keys in CI, a gateway, or Cursor usage outside OpenTelemetry | Reconcile monthly; route CI keys through the gateway or tag them |
| Engineers start gaming merged-PR counts | The report was shown per person | Remove per-person views; report cohorts per repository only |
What does measurement cost in context and money?
Section titled “What does measurement cost in context and money?”Measurement adds cost in two places: the model’s context window and your bill.
- Context cost. OpenTelemetry export, ccusage and gateways run outside the model and add no tokens to a session. Plugins and MCP servers do: run
claude plugin details <plugin>for a plugin’s projected token cost, and/contextbefore and after adding a server such as Grafana MCP or Sentry MCP. - Money. ccusage, Gitleaks, TruffleHog and the OpenTelemetry export are free apart from the collector you host. Managed review bots cost per review or per seat; the figure to track is cost per accepted change, not cost per review.