Skip to content

Measuring Agentic Engineering: Telemetry, Cost, Review, Security and Evals

Measuring agentic engineering means answering four questions every week: what the agents cost, what they shipped, what review caught, and what escaped to production. Three layers of tooling answer them: first-party telemetry built into Claude Code and Codex, local usage trackers such as ccusage, and eval, trace and gateway platforms. First-party telemetry is the layer to switch on first.

Your team has run Claude Code, Codex or Cursor for a quarter. The invoice grew, merged pull requests grew, and nobody can say whether the extra changes are good ones, because the only data is a billing total and a few Slack threads. This page is for the developer who has to switch the measurement on and the tech lead who runs the weekly review; a CTO can read the “switch on first” order as a rollout plan.

  • A map of the three measurement layers and the job each one does, so you stop buying a platform for a question a built-in setting already answers.
  • The 15 observability and quality tools that matter in September 2026, with verified install lines, dated popularity and the page that covers each in depth.
  • A switch-on order for a team or an organization: what to enable in week one and what to leave for later.
  • A telemetry smoke test for Claude Code and Codex that proves data flows before you build a dashboard, and the route for Cursor.
  • A weekly agent report (spend, pull requests, review findings, escaped defects) with a template, three copy-paste prompts, and the traps that make the numbers lie.

Agents change the volume of code faster than they change its quality, and measurement is how you tell the two apart. The 2025 DORA report puts it plainly: “AI doesn’t fix a team; it amplifies what’s already there” (Google Cloud, DORA 2025 report announcement, 2025-09-23). A weekly report is the cheapest way to see which of those your team is.

The report does not replace your delivery metrics. The metric definitions this site uses (accepted change rate, escaped defects per 100 merged changes, cost per accepted change and seven more) live on the canonical metrics page. This page covers the tools that produce the raw numbers.

What are the three layers of agent measurement?

Section titled “What are the three layers of agent measurement?”

Each layer answers different questions and has a different owner. Most teams need the first layer in full, the second for individuals, and only parts of the third.

LayerAnswersExamplesWho runs itCost to run
1. First-party telemetrySessions, tokens, estimated cost, lines, commits and pull requests per repositoryClaude Code OpenTelemetry and analytics dashboard; Codex [otel]; Cursor admin analyticsPlatform team or tech lead, set once in managed configA collector you already run, or a hosted OTel backend
2. Local usage trackers“What did I spend today, and how close am I to my plan limit?”ccusage, Claude-Code-Usage-Monitor, tokscale, /usage in Claude CodeEach developerFree; reads local logs, nothing leaves the machine
3. Eval, trace and gateway platformsDid a CLAUDE.md change make the agent better? Which team spent what? Why did this session go wrong?Langfuse, promptfoo, Inspect SWE, Braintrust; LiteLLM and Cloudflare AI Gateway for per-team spend; claude gateway for SSO and per-group model accessPlatform teamA platform to host or buy; a gateway sits in the request path

Review bots and security scanners sit beside the three layers as gates: they do not measure usage, they produce findings. Their output is what the “review findings” line of the weekly report counts.

Which observability and quality tools matter in September 2026?

Section titled “Which observability and quality tools matter in September 2026?”

The table ranks tools by relevance to a team running Claude Code, Codex or Cursor, weighted by adoption. Popularity as of 2026-09-26: GitHub stars from the GitHub API and package versions from npm and PyPI, read that day. Stars measure attention on a repository, not use. Download counts are omitted.

#ToolMeasures or gatesInstall or enablePopularity (2026-09-26)In depth
1Claude Code OpenTelemetryCost, tokens, lines, commits, pull requestsCLAUDE_CODE_ENABLE_TELEMETRY=1 plus OTEL_METRICS_EXPORTER and OTEL_LOGS_EXPORTERships in 2.1.283Telemetry
2Claude Code analytics dashboard, /usage, /insightsAdoption, pull-request attribution; one person’s usageTeam and Enterprise plan feature; /usage in a sessionvendor featureTelemetry
3ccusageCost from about 18 agents’ local logs, 5-hour blocksnpx ccusage@latest18.7k stars; npm 20.0.24Cost tracking
4Codex OpenTelemetryLogs, metrics and traces[otel] table in ~/.codex/config.tomlships in 0.157.1Telemetry
5Claude Code Review, /code-review, claude-code-actionPull request findings@claude review on a PR; /code-review in a session; anthropics/claude-code-action@v1action 9.0k starsReview bots
6codex review, openai/codex-actionDiff findings, locally or in CIcodex review --base main; codex review "Focus on auth"action 1.2k starsReview bots
7Langfuse and its Claude Code pluginA trace per Claude Code sessionclaude plugin marketplace add langfuse/Claude-Observability-Plugin → claude plugin install langfuse-observability@langfuse-observability35.1k stars (plugin repo: 24)Evals
8promptfooEvals and red-teaming; runs Claude Agent SDK or Codex SDK as the system under testnpx promptfoo@latest init --example getting-started25.5k stars; npm 0.123.1Evals
9LiteLLM proxySpend and budgets per virtual keyuv tool install 'litellm[proxy]', pinned and hash-checked59.6k stars; PyPI 1.102.1Cost tracking
10Snyk Agent ScanMCP configs and skills scanned for injectionuvx snyk-agent-scan@latest3.1k stars; PyPI 0.6.4Security gates
11Semgrep Guardian, semgrep mcpSAST on every file the agent writesClaude Code /plugin, then Discover, then Semgrep16.8k stars; PyPI 1.178.0Security gates
12GitleaksSecrets in commitsbrew install gitleaks, then a pre-commit hook29.5k starsSecurity gates
13TruffleHogSecrets, verified against the issuerbrew install trufflehog28.1k starsSecurity gates
14anthropics/claude-code-security-reviewSecurity review comments on pull requestsGitHub Action6.3k starsSecurity gates
15Entire CLIWhich prompts and session produced each commitbrew install --cask entireio/tap/entire5.1k stars; v0.11.3 (2026-09-25)Telemetry

Row 6 takes custom instructions only without a target flag: codex-cli 0.157.1 rejects a prompt combined with --base, --uncommitted or --commit.

Also verified and covered on the deeper pages: Braintrust, Opik, Arize Phoenix, Inspect AI with Inspect SWE, DeepEval, garak, Portkey, Cloudflare AI Gateway, Claude-Code-Usage-Monitor, tokscale, SonarQube MCP and Snyk MCP. Do not start a new project on Helicone: it has been in maintenance mode since March 2026 (secondary sources: the Helicone and Mintlify acquisition posts).

Switch on the layer that answers the most expensive open question first. For almost every organization that is “what are we paying, and what did it produce?”, so the order below starts with first-party telemetry and leaves evals for later.

  1. Week 1: first-party telemetry, through managed configuration. Put the Claude Code OpenTelemetry variables in managed settings (see managed policy) and ship a Codex [otel] block with your standard config.toml. Developers cannot forget a setting they never had to set. The organization-wide design, with collector pipelines, join keys and retention, is on agent observability.

  2. Week 1: a pull request marker. On Claude Team or Enterprise, install the Claude GitHub app and enable GitHub analytics so merged pull requests get the claude-code-assisted label, and add an agent-assisted checkbox to the pull request template for Codex and Cursor. Without a marker, every later metric compares nothing with nothing.

  3. Week 2: deterministic security gates. Gitleaks as a pre-commit hook and a SAST scan in CI. Agents commit often and never tire, so these gates fire more often than they did for human-only teams, and they cost no tokens.

  4. Week 3: one review bot, measured. Start with the bot bundled with your main tool (/code-review or Claude Code Review, codex review, or Bugbot), and record which findings lead to a change. Anthropic’s docs put the managed Claude Code Review (research preview on Team and Enterprise; not available under Zero Data Retention) at an average of $15–25 per review (checked 2026-09-26), so measure it before you roll it out to every repository.

  5. Week 4: cost attribution. ccusage for individuals. A gateway with virtual keys (LiteLLM or Cloudflare AI Gateway) only when finance needs per-team budgets that OpenTelemetry cannot enforce. claude gateway is a separate thing: Anthropic’s self-hosted auth and telemetry gateway for SSO and per-group model access. Cost per accepted change counts seat licences as well as usage and tokens, as metric 8 on the metrics frameworks page does.

  6. Later: evals. Build an eval set once you change shared CLAUDE.md or AGENTS.md rules, skills or models often enough that “it felt better” is no longer an answer. Claude Code 2.1.283 also ships claude plugin eval, which runs a plugin’s eval cases against a no-plugin baseline.

The quickest proof that telemetry works is a console exporter: you see a datapoint in your own terminal before any collector exists. The OTLP export, the collector stack and the dashboard panels are on OpenTelemetry and analytics for Claude Code, Codex and Cursor; this section only proves the pipe is open.

In the terminal you start Claude Code from:

Terminal window
# Smoke test: print metrics to this terminal every 10 seconds
export CLAUDE_CODE_ENABLE_TELEMETRY=1
export OTEL_METRICS_EXPORTER=console
export OTEL_METRIC_EXPORT_INTERVAL=10000
claude

Within one export interval, the terminal prints a claude_code.session.count datapoint. For a team, the exporter variables belong in the env block of managed settings: Claude Code ignores them in a repository’s .claude/settings.json, so a repository cannot redirect your telemetry.

Run a weekly agent report: spend, PRs, review findings, escaped defects

Section titled “Run a weekly agent report: spend, PRs, review findings, escaped defects”

The weekly report is where the layers meet. Once the sources exist, an agent can draft it and the tech lead reviews it. Report per repository and per cohort (agent-assisted against everything else), never per engineer.

  1. Spend. Take estimated cost from claude_code.cost.usage for Claude Code. For Codex, take estimated cost from codex.turn.cost_microusd (micro-USD per turn, codex-cli 0.157.1), or run npx ccusage@latest weekly --json --by-agent and take the Codex rows (ccusage 20.0.24; ccusage codex has no weekly view). ccusage reads only the local machine’s logs, so it gives per-developer figures, not team totals. For Cursor, take usage from the admin dashboard or Analytics API (Teams and Enterprise; secondary, confirm in your console). A gateway, if you run one, reports spend for every tool behind it. Treat all of these as estimates and reconcile with the invoice monthly.

  2. Pull requests. Count merged pull requests with an agent marker, and the share of all merged pull requests they represent. With the GitHub CLI:

    Terminal window
    SINCE=2026-09-19 # first day of the reporting week
    gh pr list --state merged --label claude-code-assisted \
    --search "merged:>=$SINCE" --json number,title,url --limit 200
    gh pr list --state merged --label agent-assisted \
    --search "merged:>=$SINCE" --json number,title,url --limit 200
    gh pr list --state merged \
    --search "merged:>=$SINCE" --json number --limit 500 --jq length

    The last command counts all merged pull requests, the denominator for the share.

  3. Review findings. Count findings each bot and each gate posted, and how many led to a code change. That acted-on rate is the precision proxy that decides whether a bot keeps its budget. Record findings from deterministic gates (Gitleaks, Semgrep) separately from model reviewers.

  4. Escaped defects. Count production bugs opened this week that trace back to a merged change, per 100 merged changes, split by cohort. This is metric 7 on the canonical metrics page. Every escaped defect in the agent cohort gets a one-line note on which gate should have caught it.

  5. Decide one thing. Every report ends with one change to the harness: a new test, a rule in CLAUDE.md or AGENTS.md, a scanner rule, or a review-bot setting. A report that changes nothing is a vanity dashboard.

Save this template as docs/agent-report/TEMPLATE.md in the repository that holds your team’s runbooks:

# Agent report — week of YYYY-MM-DD (repository: NAME)
| Line | Agent-assisted | Other | Source |
| --- | --- | --- | --- |
| Merged PRs | | | gh pr list, label claude-code-assisted / agent-assisted |
| Estimated spend (USD) | | n/a | claude_code.cost.usage / codex.turn.cost_microusd / ccusage / gateway |
| Cost per accepted change (USD) | | n/a | (estimated spend + seat cost) ÷ agent-assisted PRs merged 14–21 days ago and not reverted or fixed since |
| Review findings posted / acted on | | | bot comments, gate logs |
| Escaped defects per 100 merged changes | | | bug tracker, linked PRs |
Escaped defects this week: PR, bug, which gate should have caught it.
One harness change we make next week: ...
Owner: tech lead. Sign-off: engineering manager.

Copy-paste prompts for the weekly agent report

Section titled “Copy-paste prompts for the weekly agent report”

Run these in Claude Code, Codex or Cursor from a checkout of the repository, with the GitHub CLI signed in with a read-only token. The prompts are the same in all three tools.

BUG_NUMBER is the issue number of the escaped bug.

How do you know the measurement itself is right?

Section titled “How do you know the measurement itself is right?”

A measurement system is code, and it fails the way code does: quietly. Five checks keep it honest, and the tech lead signs off on each before the numbers reach a leadership slide.

  • A known event produces a known datapoint. Start one session and confirm the session counter moves by one in your backend. Repeat after every collector or Claude Code upgrade.
  • Spend reconciles. Once a month, compare OpenTelemetry or gateway cost with the invoice. A gap above your tolerance means an unexported team, a second billing path, or a wrong price table.
  • Attribution is spot-checked. Pick five merged pull requests at random and check the label against the session history or the Entire checkpoint trailer.
  • Cohorts, not people. No panel groups by user_id or email. Per-person panels turn the report into surveillance and the metric into a target.
  • Speed never appears without stability. Merged pull requests and spend always sit beside escaped defects or the 14-day follow-up fix rate.

What breaks when you measure agentic engineering?

Section titled “What breaks when you measure agentic engineering?”
SymptomCauseRecovery
Agent-assisted share jumps sharply in one weekA new marker (the GitHub app, a template checkbox) started counting, not a change in behaviorAnnotate the dashboard with the date each marker went live; compare only periods with the same markers
Review bot findings rise, nobody acts on themThe bot is tuned for recall; people resolve threads unreadTrack acted-on rate per bot; tune it (a REVIEW.md for Claude Code Review) or drop it
Spend is flat but the invoice roseA second billing path: API keys in CI, a gateway, or Cursor usage outside OpenTelemetryReconcile monthly; route CI keys through the gateway or tag them
Engineers start gaming merged-PR countsThe report was shown per personRemove per-person views; report cohorts per repository only

What does measurement cost in context and money?

Section titled “What does measurement cost in context and money?”

Measurement adds cost in two places: the model’s context window and your bill.

  • Context cost. OpenTelemetry export, ccusage and gateways run outside the model and add no tokens to a session. Plugins and MCP servers do: run claude plugin details <plugin> for a plugin’s projected token cost, and /context before and after adding a server such as Grafana MCP or Sentry MCP.
  • Money. ccusage, Gitleaks, TruffleHog and the OpenTelemetry export are free apart from the collector you host. Managed review bots cost per review or per seat; the figure to track is cost per accepted change, not cost per review.