Skip to content

Observing the agents: traces, tool calls, cost and audit logs

Agent observability means exporting each coding agent’s telemetry (sessions, model requests, tool calls, permission decisions, and cost) through OpenTelemetry, stamping every run with IDs that join it to its pull request and to later incidents, and keeping audit events longer than analytics. Claude Code and Codex export OpenTelemetry natively; for Cursor, hooks and commit provenance fill the gap.

This page is for the developer who runs agent loops in CI, the tech lead who owns the loop’s autonomy, and the CTO who signs the budget and the audit policy. The situation it solves: an agent-written change caused a production incident on Tuesday, and on Wednesday nobody can say which session wrote it, what the agent ran, who approved the tool calls, or what that loop costs per change that survives. You have logs for your services and none for the things writing them.

What you get from agent observability done this way

Section titled “What you get from agent observability done this way”
  • A collector with a pseudonymized analytics pipeline and an identified audit pipeline.
  • Export settings for Claude Code, Codex, and the verified path for Cursor.
  • Four join keys that lead from a commit to its pull request, session, and tool calls.
  • An eight-panel dashboard with its SQL, a retention template, and three copy-paste prompts.

Why do agent runs need their own telemetry?

Section titled “Why do agent runs need their own telemetry?”

Your service telemetry answers “what did the system do?”. Agent telemetry answers a different question: “what did the thing that changed the system do, on whose authority, and at what cost?”. Faros AI’s AI Engineering Report 2026: The Acceleration Whiplash (April 2026, two years of telemetry from 22,000 developers on Faros’s own platform) reports median time in review up 441.5% and incidents per pull request up 242.7%. When reviewers cannot read everything, evidence about the run has to come from telemetry.

Three things make agent telemetry different:

  • The unit is a run, not a request. One run is a session of many model calls and tool calls, sometimes with subagents. Success is decided days later, when the pull request merges and survives.
  • Identity is a developer’s. Claude Code attributes every tool call to the developer who started the session; it does not act under a separate service account (Claude Code monitoring docs, checked 2026-09-26). The same data is both an audit trail and a potential surveillance tool.
  • Content is sensitive by default. Prompts, tool arguments, and model responses can contain secrets and customer data. Both tools keep content out of telemetry unless you opt in.

Configure each tool separately and normalize afterwards.

ConceptClaude Code 2.1.283Codex CLI 0.157.1Cursor
Export mechanismOTel metrics and log events; traces in betaOTel logs, traces, and metrics from the [otel] tableNo OTel export verified on 2026-09-26; use hooks and provenance tools
Session or run IDsession.id (OTel); session_id in claude -p --output-format jsonconversation.id (OTel); thread_id in codex exec --jsonCloud Agents API run IDs and statusChange webhooks (verified 2026-08-28)
Costclaude_code.cost.usage (USD); total_cost_usd per headless runcodex.turn.cost_microusd metric; codex.turn_cost event with usage.estimated_usdNot verified
Tokensclaude_code.token.usage by type, model, effortcodex.turn.token_usage; usage on each turn.completedNot verified
Tool callsclaude_code.tool_result (tool_name, success, duration_ms)codex.tool_result; codex.tool.call metricHooks (JSON over stdio); event names not verified
Permission decisionsclaude_code.tool_decision, claude_code.permission_mode_changedcodex.tool_decision, codex.sandbox_outcomeNot verified
Link to a commitvcs.ref.head.revision on the tool_result of a successful git commit (needs OTEL_LOG_TOOL_DETAILS=1)None; record the thread_id in the pull requestEntire-Checkpoint commit trailer from Entire CLI
Vendor dashboardclaude.ai/analytics/claude-code (Team, Enterprise); platform.claude.com/claude-code (Console)Enterprise analytics dashboard (secondary: OpenAI help pages seen only in search extracts)Not verified on 2026-09-26 (cursor.com unreachable)

Sources: Claude Code monitoring and analytics docs; Codex source at tag rust-v0.157.1 (codex-rs/otel, codex-rs/config/src/types.rs, codex-rs/exec/src/exec_events.rs), all read 2026-09-26. cursor.com was unreachable on 2026-09-26, so Cursor’s analytics and audit endpoints are not named here.

  1. Write the measurement charter before you collect anything. One paragraph, signed by the CTO: which questions the data answers (run success, cost per accepted change, audit), who sees identities (the security team, through the SIEM), who sees pseudonyms (everyone else), and that no metric is used to rate an individual. The metrics frameworks page explains why PR counts and token counts become gamed targets, and privacy and data handling covers vendor data terms and retention. Telemetry that engineers believe is a surveillance tool gets switched off in their shells.

  2. Stand up one OpenTelemetry Collector with split pipelines. Every tool sends OTLP to the same collector. For the analytics backends the collector replaces the email with a keyed hash, deletes the other user identifiers, and drops tool arguments; the SIEM gets the unmodified audit stream:

    # otel-collector.yaml (OpenTelemetry Collector contrib distribution)
    receivers:
    otlp:
    protocols:
    grpc: { endpoint: 0.0.0.0:4317 }
    http: { endpoint: 0.0.0.0:4318 }
    processors:
    batch: {}
    # Keyed hash: PSEUDONYM_SALT is a secret only the collector holds.
    transform/pseudonymize:
    error_mode: ignore
    log_statements:
    - context: log
    statements:
    - set(attributes["user.email"], SHA256(Concat([attributes["user.email"], "${env:PSEUDONYM_SALT}"], ""))) where attributes["user.email"] != nil
    metric_statements:
    - context: datapoint
    statements:
    - set(attributes["user.email"], SHA256(Concat([attributes["user.email"], "${env:PSEUDONYM_SALT}"], ""))) where attributes["user.email"] != nil
    attributes/drop-ids:
    actions:
    - key: user.account_uuid
    action: delete
    - key: user.account_id
    action: delete
    - key: user.id
    action: delete
    - key: enduser.id
    action: delete
    attributes/strip-content:
    actions:
    - key: tool_parameters
    action: delete
    - key: tool_input
    action: delete
    - key: error
    action: delete
    exporters:
    prometheus:
    endpoint: 0.0.0.0:8889
    otlphttp/analytics:
    endpoint: https://logs.internal.example.com/otlp
    otlphttp/siem:
    endpoint: https://siem.internal.example.com:4318
    service:
    pipelines:
    metrics:
    receivers: [otlp]
    processors: [transform/pseudonymize, attributes/drop-ids, batch]
    exporters: [prometheus]
    logs/analytics:
    receivers: [otlp]
    processors: [transform/pseudonymize, attributes/drop-ids, attributes/strip-content, batch]
    exporters: [otlphttp/analytics]
    logs/audit:
    receivers: [otlp]
    processors: [batch]
    exporters: [otlphttp/siem]

    Claude Code attaches user.email, user.account_uuid, user.account_id (the admin-API ID), and user.id (a persistent per-install ID) to every metric and event; hashing only the email leaves three ways back to each person. The attributes processor’s hash action is unsalted, so a list of staff emails reverses it. logs.internal.example.com and siem.internal.example.com stand for your log store and your SIEM’s OTLP receiver.

  3. Turn on export in each tool. Developer machines get it through managed configuration; CI jobs set it in the job.

    Put the exporter settings in managed settings. Claude Code ignores OpenTelemetry exporter variables in a repository’s .claude/settings.json and .claude/settings.local.json, so a repository cannot turn telemetry on, redirect it, or capture content. When managed settings set OTEL_EXPORTER_OTLP_ENDPOINT, Claude Code removes developer-set per-signal endpoints at startup (v2.1.217 and later).

    {
    "env": {
    "CLAUDE_CODE_ENABLE_TELEMETRY": "1",
    "OTEL_METRICS_EXPORTER": "otlp",
    "OTEL_LOGS_EXPORTER": "otlp",
    "OTEL_EXPORTER_OTLP_PROTOCOL": "grpc",
    "OTEL_EXPORTER_OTLP_ENDPOINT": "https://otel-collector.internal.example.com:4317",
    "OTEL_METRICS_INCLUDE_REPOSITORY": "true",
    "OTEL_LOG_TOOL_DETAILS": "1"
    }
    }

    OTEL_LOG_TOOL_DETAILS=1 adds Bash commands, MCP server and tool names, and tool input to events. It is what makes MCP activity auditable and what puts the commit SHA on git commit tool results; it is also where secrets typed on a command line end up, which is why the collector in step 2 strips those attributes from everything except the SIEM stream, and why the endpoint must be TLS (https://) or reachable only over a private network. Prompt text stays off unless you set OTEL_LOG_USER_PROMPTS=1.

    For a headless CI run, set the same variables on the job, and give the run its own session ID so the pull request can carry it:

    .github/workflows/agent-issue-to-pr.yml
    name: agent-issue-to-pr
    on:
    workflow_dispatch:
    inputs:
    issue:
    description: Issue number to implement
    required: true
    type: number
    permissions: {}
    env:
    ISSUE: ${{ inputs.issue }}
    BRANCH: agent/issue-${{ inputs.issue }}
    jobs:
    issue-to-pr:
    # Runs the agent. Read-only token: nothing here can push.
    runs-on: ubuntu-latest
    timeout-minutes: 30
    permissions:
    contents: read
    issues: read
    outputs:
    session_id: ${{ steps.sid.outputs.id }}
    env:
    CLAUDE_CODE_ENABLE_TELEMETRY: "1"
    OTEL_METRICS_EXPORTER: otlp
    OTEL_LOGS_EXPORTER: otlp
    OTEL_EXPORTER_OTLP_PROTOCOL: http/protobuf
    OTEL_EXPORTER_OTLP_ENDPOINT: ${{ secrets.OTEL_ENDPOINT }}
    OTEL_EXPORTER_OTLP_HEADERS: ${{ secrets.OTEL_HEADERS }}
    OTEL_RESOURCE_ATTRIBUTES: ci.run_id=${{ github.run_id }},loop.name=issue-to-pr
    OTEL_METRICS_INCLUDE_RESOURCE_ATTRIBUTES: "false"
    steps:
    - uses: actions/checkout@v7
    with:
    persist-credentials: false
    - uses: actions/setup-node@v7
    with:
    node-version: 22
    - name: Install Claude Code
    run: npm install -g @anthropic-ai/claude-code
    - name: Prepare the branch and the prompt
    env:
    GH_TOKEN: ${{ github.token }}
    run: |
    git config user.name "issue-to-pr agent"
    git config user.email "issue-to-pr-agent@users.noreply.github.com"
    git switch -c "$BRANCH"
    { cat .github/prompts/issue-to-pr.md; echo
    gh issue view "$ISSUE" --json number,title,body \
    --jq '"Issue #\(.number): \(.title)\n\n\(.body)"'
    } > "$RUNNER_TEMP/prompt.md"
    - name: Pick a session ID
    id: sid
    run: echo "id=$(cat /proc/sys/kernel/random/uuid)" >> "$GITHUB_OUTPUT"
    - name: Run the agent
    env:
    ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
    SID: ${{ steps.sid.outputs.id }}
    run: |
    status=0
    claude -p "$(cat "$RUNNER_TEMP/prompt.md")" \
    --session-id "$SID" --output-format json \
    --max-budget-usd 10 --permission-mode dontAsk \
    --allowedTools "Read,Grep,Glob,Edit,Write,Bash(npm test *),Bash(git commit *)" \
    > "$RUNNER_TEMP/run.json" || status=$?
    jq -c '{session_id, is_error, subtype, num_turns, total_cost_usd, duration_ms}' \
    "$RUNNER_TEMP/run.json" >> "$RUNNER_TEMP/runs.jsonl"
    exit "$status"
    - name: Keep the run record
    if: always()
    uses: actions/upload-artifact@v7
    with:
    name: agent-run-${{ github.run_id }}
    path: ${{ runner.temp }}/runs.jsonl
    if-no-files-found: ignore
    - name: Bundle the agent's commits
    env:
    BASE: ${{ github.sha }}
    run: git bundle create "$RUNNER_TEMP/agent.bundle" "$BRANCH" "^$BASE"
    - uses: actions/upload-artifact@v7
    with:
    name: agent-commits-${{ github.run_id }}
    path: ${{ runner.temp }}/agent.bundle
    if-no-files-found: error
    open-pr:
    # Holds the write token. Runs no code from the agent's commits.
    needs: issue-to-pr
    runs-on: ubuntu-latest
    timeout-minutes: 5
    permissions:
    contents: write
    pull-requests: write
    steps:
    - uses: actions/checkout@v7
    with:
    persist-credentials: false
    - uses: actions/download-artifact@v7
    with:
    name: agent-commits-${{ github.run_id }}
    path: ${{ runner.temp }}
    - name: Push and open the pull request
    env:
    GH_TOKEN: ${{ github.token }}
    AGENT_SESSION_ID: ${{ needs.issue-to-pr.outputs.session_id }}
    run: |
    git fetch "$RUNNER_TEMP/agent.bundle" "$BRANCH:$BRANCH"
    gh auth setup-git
    git push origin "refs/heads/$BRANCH"
    gh pr create --head "$BRANCH" --title "Fix #$ISSUE (agent)" \
    --body "Closes #$ISSUE. Agent session: $AGENT_SESSION_ID"

    OTEL_METRICS_INCLUDE_RESOURCE_ATTRIBUTES=false keeps ci.run_id on events but off metric labels, where it would create one time series per CI run. The subtype field tells you why a run stopped: success, error_max_turns, error_max_budget_usd, error_during_execution, or error_max_structured_output_retries. The run is started by hand (workflow_dispatch needs write access to the repository) and checks out the ref chosen at dispatch, which is the default branch unless the person starting it picks another in “Use workflow from” or passes gh workflow run --ref. Steps inside the agent job are not isolated from each other: Edit, Write and Bash(npm test *) let the agent rewrite the test script and run any code it likes, the issue text it works from is untrusted, and that code can read ANTHROPIC_API_KEY and the OTLP headers and write to $GITHUB_ENV or $GITHUB_PATH, which every later step of the same job inherits. That is why the agent job holds only a read-only GITHUB_TOKEN, and the push lives in a separate open-pr job: it receives the commits as a git bundle artifact, gets the session ID from a step that ran before the agent, and runs no code from those commits. If the agent made no commits, git bundle create refuses an empty bundle and the run stops before open-pr. Read the issue before you dispatch the run; for workflows triggered by untrusted contributors, follow the CI security rules first. gh pr create with GITHUB_TOKEN works only when “Allow GitHub Actions to create and approve pull requests” is enabled in the repository’s or organization’s Actions settings, and a pull request opened with that token does not trigger pull_request workflows, so the agent’s PR gets no CI run. If CI must run on it, open the PR with a GitHub App token or a fine-grained personal access token instead. The run record is written even when the agent stops with an error, so failed runs reach runs.jsonl too.

    Traces are beta: add CLAUDE_CODE_ENHANCED_TELEMETRY_BETA=1 and OTEL_TRACES_EXPORTER=otlp. In claude -p and Agent SDK sessions, Claude Code reads TRACEPARENT from its environment, so its claude_code.interaction spans nest under your CI job’s trace.

  4. Stamp every run with four join keys, and write them into the pull request. Telemetry you cannot join to outcomes only tells you what you spent. Each run carries:

    KeyClaude CodeCodexCursorWhere it lands
    Session--session-id UUIDthread_idEntire checkpoint ID or cloud run IDEvidence bundle provenance.session
    CI runOTEL_RESOURCE_ATTRIBUTES ci.run_id[otel.span_attributes]runs.jsonlruns.jsonl, events
    Looploop.name resource attributeloop.name span attributeruns.jsonlDashboard grouping
    TaskIssue URL in the prompt and the bundleSameSameEvidence bundle provenance.task

    The pull request’s evidence bundle already has a provenance.session field; the CI job fills it from AGENT_SESSION_ID or the thread_id. Join on the session, not on the commit SHA: a squash merge creates a new SHA, so the SHA recorded at git commit time never appears on the default branch. On Team and Enterprise plans with the GitHub app, Claude Code’s analytics also labels merged pull requests claude-code-assisted, a cross-check that is unavailable under Zero Data Retention.

  5. Load outcomes and compute the dashboard. Compute the outcome panels in your warehouse, since metrics backends cannot join to GitHub, from two tables: agent_runs (from runs.jsonl, with cost from telemetry grouped by session) and pull_requests (from the GitHub API, with the session parsed from the bundle and a 14-day follow-up flag as defined on the metrics page).

    -- Weekly outcome panels per loop. A change is accepted when it merged and was
    -- not reverted or fixed within 14 days, so the last two weeks are provisional.
    WITH pr AS ( -- one row per session, so a run with several PRs is counted once
    SELECT repo, agent_session,
    count(*) AS prs,
    count(*) FILTER (WHERE merged_at IS NOT NULL
    AND NOT followup_within_14d) AS accepted
    FROM pull_requests
    WHERE agent_session IS NOT NULL
    GROUP BY repo, agent_session
    )
    SELECT
    date_trunc('week', r.started_at) AS week,
    r.loop_name,
    count(*) AS runs,
    avg(CASE WHEN r.completed_ok THEN 1.0 ELSE 0 END) AS run_completion_rate,
    sum(CASE WHEN r.completed_ok AND p.prs > 0 THEN 1 ELSE 0 END)::numeric
    / NULLIF(sum(CASE WHEN r.completed_ok THEN 1 ELSE 0 END), 0) AS run_to_pr_rate,
    sum(coalesce(p.accepted, 0)) AS accepted_changes,
    sum(r.cost_usd) / NULLIF(sum(coalesce(p.accepted, 0)), 0) AS run_cost_per_accepted_change
    FROM agent_runs r
    LEFT JOIN pr p
    ON p.repo = r.repo AND p.agent_session = r.session_id
    GROUP BY 1, 2
    ORDER BY 1 DESC, 2;

    completed_ok is NOT is_error for Claude Code and NOT failed for Codex. This query gives run cost only; the organization’s full cost per accepted change adds seats and platform cost, per the metrics definitions and the economics page.

  6. Wire incidents to runs. When an incident is traced to a deploy, the chain is: deploy, merge commit, pull request, provenance.session, then every event with that session.id or conversation.id, ordered by event.timestamp (with event.sequence breaking ties) and grouped by prompt.id in Claude Code. The tool_decision events show whether a risky call was allowed by configuration, a hook, or a person. Add the session link to the incident record in your agent incident process, and tag the cause with the failure taxonomy.

  7. Set retention and access per stream, then write it down. See the template in the next section; configure it in the backends and in Claude Code’s cleanupPeriodDays for local transcripts.

Eight panels answer a tech lead’s weekly and a CTO’s monthly questions. Group by loop, repository, and tool, never by person.

PanelDefinitionSourceWatch for
Run completion rateRuns that ended without error ÷ runs startedruns.jsonlA drop after a model, prompt, or harness change
Run-to-PR rateRuns that opened a pull request ÷ runs that completedruns.jsonl + GitHubRuns that finish but produce nothing reviewable
Accepted change rateMerged and not followed up within 14 days ÷ pull requests openedGitHubThe real throughput; the definition is canonical
Run cost per accepted changeRun cost ÷ accepted changes (the SQL above)WarehouseCost rising while acceptance stays flat
Spend by model and effortclaude_code.cost.usage by model, effort, query_source; codex.turn.cost_microusd by reasoning_effortMetricsSubagent or auxiliary spend growing unnoticed
Tool error ratetool_result with success=false ÷ all, by tool_nameEventsA broken MCP server or test command burning turns
Permission frictiontool_decision rejects by source; permission_mode_changed eventsEventsLoops blocked on approvals, or silent mode escalations
Budget and retry stopserror_max_budget_usd and error_max_turns subtypes; claude_code.api_error eventsruns.jsonl, eventsBudgets set too tight, or provider outages

Set three alerts from day one: a single session above your cost ceiling, any permission_mode_changed event whose to_mode is bypassPermissions outside a sandboxed runner, and retry exhaustion above its weekly baseline. For a sense of scale, Anthropic’s cost documentation (checked 2026-09-26) puts typical enterprise Claude Code spend at “around $13 per developer per active day and $150-250 per developer per month”; your own baseline after four weeks is the number to alert on.

Retention follows what each stream contains. Treat this as a starting policy for your security and legal owners to adjust; obligations under the EU AI Act and sector rules in regulated industries can require longer.

StreamContainsStoreStarting retentionWho can read
MetricsCost, tokens, sessions, lines, and pseudonymized usersMetrics backend13 months, for year-on-year comparisonEngineering
Operational eventstool_result, api_request, api_error, pseudonymizedLog store90 daysEngineering
Audit eventstool_decision, permission_mode_changed, hook_execution_complete, auth, mcp_server_connection, plugin_installed, managed_settings_resolved; codex.tool_decision, codex.sandbox_outcomeSIEMYour security log policy (often one year or more)Security
ContentPrompts, responses, tool output, and raw API bodiesOff by default; if on, a restricted storeShortest workable, for example 30 daysNamed incident responders
Change evidenceEvidence bundle, session link, and checkpoint trailersGit and the code hostLife of the repositoryEveryone with repository access
Local transcriptsClaude Code session files on developer machinesDiskcleanupPeriodDays (default 30), pinned in managed settingsThe developer

Claude Code emits a claude_code.retention_sweep event each time it runs that cleanup; a nonzero files_past_cutoff means files outlived the configured period, which is itself an audit finding.

Copy-paste prompts for agent observability

Section titled “Copy-paste prompts for agent observability”

A dashboard that reads zero or double is worse than none. Verify the pipeline itself:

  • Arrival. After any configuration change, check for claude_code.session.count (metrics) or claude_code.user_prompt (events). If nothing arrives, start Claude Code with claude --debug-file /tmp/claude-otel.log and look for [3P telemetry] errors. For Codex, run one codex exec and look for codex.conversation_starts.
  • Join coverage. Each week, count merged agent pull requests whose provenance.session resolves to at least one telemetry event. Below 95%, a workflow is dropping its session ID; the instrumentation audit prompt finds it.
  • Cost reconciliation. Claude Code’s docs call cost metrics approximations. Once a month, compare summed claude_code.cost.usage and Codex’s estimated cost with the invoice or the Console, per workspace. A gap above a few percent usually means unreported CI runs or a second billing path.
  • Canary run. A scheduled CI job runs a fixed, cheap prompt once a day and fails if its runs.jsonl record and its session’s events are not in the warehouse within an hour.
WhoOwnsSigns
Platform or developer-experience teamCollector, managed settings, config.toml templates, and the canaryPipeline changes
Tech leadLoop names, the dashboard review in the weekly retroAutonomy changes based on the panels
SecurityAudit pipeline, SIEM rules, and content-logging approvalsRetention and access
CTOThe measurement charter and the monthly cost viewBudget ceilings and the charter

What breaks when you instrument coding agents?

Section titled “What breaks when you instrument coding agents?”

Telemetry is enabled but nothing arrives. Claude Code has no default OTLP protocol, so every otlp exporter needs OTEL_EXPORTER_OTLP_PROTOCOL or its per-signal variant. Recovery: set the protocol, then check the debug file for [3P telemetry] errors. Lines prefixed [Anthropic telemetry] are Anthropic’s own operational telemetry, not your pipeline.

Metric storage costs explode. Per-run IDs in OTEL_RESOURCE_ATTRIBUTES and the default session.id on metrics create a new series per run. Recovery: set OTEL_METRICS_INCLUDE_RESOURCE_ATTRIBUTES=false and, in large organizations, OTEL_METRICS_INCLUDE_SESSION_ID=false; keep run and session IDs on events, where cardinality is cheap.

A secret shows up in the log store. OTEL_LOG_TOOL_DETAILS=1 records full Bash commands, and a token passed on a command line goes with it. Recovery: rotate the secret, purge the records, route tool details only to the audit pipeline, and add a collector rule or SIEM alert for credential patterns. Never enable OTEL_LOG_RAW_API_BODIES outside a time-boxed investigation; it exports the whole conversation.

The cost panel disagrees with the invoice. Telemetry cost is a client-side estimate, Claude Code before v2.1.214 double-counted cost and tokens on streams that emit multiple cumulative message_delta frames, and CI runs without telemetry are invisible. Recovery: upgrade, reconcile monthly, and report the invoice as the authoritative number with telemetry as the breakdown.

Commits do not join to sessions. Squash and rebase merges rewrite SHAs, and Codex CLI 0.157.1 emits no commit SHA in its telemetry. Recovery: join on provenance.session in the evidence bundle, and fail the pull request check when the field is empty.

The dashboard becomes a leaderboard. Someone sorts cost by developer. Recovery: the charter forbids it, the analytics pipeline carries only a keyed hash of the email with the account and install IDs deleted, and the per-user view exists only in the SIEM for security investigations.

Codex metrics go somewhere you did not choose. Leaving metrics_exporter unset in 0.157.1 means the statsig default. Recovery: set it explicitly to your collector or to "none" in the config.toml template you distribute.

Telemetry tells you what the runs did. The next pages use it to classify failures, handle incidents, and keep the codebase healthy.