Skip to content

ISO/IEC 42001, NIST AI RMF and SOC 2 evidence from the agent pipeline

An agent pipeline built on the evidence bundle already produces most of what ISO/IEC 42001, the NIST AI Risk Management Framework and a SOC 2 examination ask for. Specs, acceptance results, evals, approvals, telemetry and provenance map to named clauses, subcategories and criteria. The work is mapping them once, archiving them and testing the controls.

Your SOC 2 Type II period closes in December. A customer questionnaire asks whether you are “aligned with the NIST AI RMF”, sales wants to say “ISO/IEC 42001 in progress”, and your auditor has noticed that a third of this year’s merged pull requests came from Claude Code or Codex. Nobody wants a second compliance program. You need to know which artifacts your pipeline already writes count as evidence, and where the gaps are.

This page is for the CTO who owns that answer. It is an engineering reading, not audit or legal advice: agree the control design with your auditor or certification body. It builds on the separation-of-duties controls and merge-time archive in agentic engineering in regulated industries.

What this page gives a CTO preparing for an AI management audit

Section titled “What this page gives a CTO preparing for an AI management audit”
  • A scope decision: which ISO/IEC 42001 controls apply when you use agents, and which when you ship AI features.
  • A control-to-evidence table mapping each pipeline artifact to ISO/IEC 42001, NIST AI RMF and SOC 2.
  • An evidence map file with the query, owner and frequency for every control.
  • Telemetry settings for Claude Code, Codex and Cursor, an inventory script and three copy-paste prompts.

What do ISO/IEC 42001, NIST AI RMF and SOC 2 each ask for?

Section titled “What do ISO/IEC 42001, NIST AI RMF and SOC 2 each ask for?”

The three frameworks answer different questions, so one evidence set can serve all of them.

FrameworkWhat it isWhat you getWhat the assessor looks at
ISO/IEC 42001:2023A certifiable AI management system: clauses 4 to 10 plus 38 reference controls in Annex A, under nine objectives from A.2 to A.10A certificate from an accredited certification bodyYour AI risk assessment and treatment, the Statement of Applicability, and evidence that the controls run
NIST AI RMF 1.0 (NIST AI 100-1)A voluntary framework with four functions, GOVERN, MAP, MEASURE and MANAGE, broken into categories and subcategoriesNo certificate; a self-declared alignment that customer questionnaires ask aboutWhatever the customer or internal audit samples: usually policies, roles, inventory, tests, monitoring and incidents
SOC 2An AICPA attestation report on your controls against the Trust Services Criteria; the Common Criteria CC1 to CC9 cover securityA Type I report (design at a point in time) or a Type II report (operation over a period)A sample of changes, access grants and incidents, tested against your described controls

Two related documents help with scoping: NIST AI 600-1, the Generative AI Profile (2024-07-26), which lists twelve generative-AI risks, and ISO/IEC 42005:2025 on impact assessment (both SECONDARY, from search extracts of the NIST and ISO pages, 2026-09-26). The mapping below uses AI RMF 1.0 as published.

Are your coding agents in scope for an AI management system?

Section titled “Are your coding agents in scope for an AI management system?”

Decide this first, in writing, because it sets which Annex A controls appear in your Statement of Applicability. Running Claude Code, Codex or Cursor makes you a user of third-party AI systems; shipping a chatbot, recommender or agent feature also makes you a developer. Most software companies are both.

SituationISO/IEC 42001 controls that carry the weightNIST AI RMF focusSOC 2 focus
You use coding agents to build software that contains no AIA.4.4 tooling resources, A.9.2 to A.9.4 responsible and intended use, A.10.3 suppliers, A.6.2.8 event logs, A.3.2 rolesGOVERN 1, 2, 6; MANAGE 3 (third-party resources and pre-trained models)CC8.1 changes, CC6 access for agent identities, CC7.2 monitoring, CC9.2 vendors
You also ship AI features built with those agentsAll of the above, plus the life-cycle controls A.6.2.2 to A.6.2.7, data controls A.7, and impact assessment A.5 for each featureAll four functions per AI system, with MEASURE 2 evaluations per releaseThe same, with processing-integrity criteria if your report includes them

You may exclude the agent pipeline, but every excluded Annex A control needs a justification. “Developers use coding agents under our standard change process” can be tested; “coding agents are out of scope” cannot.

Each row is one control objective, the clause, subcategory and criterion it satisfies, and the artifact that answers it, built elsewhere on this site: the evidence bundle, the risk classes in governance and autonomy, managed policy, telemetry, evals and the incident log.

Control objectiveISO/IEC 42001NIST AI RMFSOC 2Evidence from the agent pipeline
AI policy that is enforced, not only published5.2, A.2.2, A.2.3GOVERN 1.1, 1.2, 1.4CC5.3The AI usage policy plus managed settings files (/etc/claude-code/managed-settings.json, /etc/codex/requirements.toml, Cursor admin settings) versioned in git; Claude Code’s managed_settings_resolved event shows which policy each session ran
Roles, accountability and human oversight5.3, A.3.2GOVERN 2.1, 3.2; MAP 3.5CC1.3, CC1.5The operating model RACI; provenance.human_owner in every bundle; CODEOWNERS for sensitive paths; the rule that an agent may author and review but never approve
Inventory of AI tools, models and extensionsA.4.2, A.4.4GOVERN 1.6; MAP 4.1CC9.2; system descriptionThe quarterly inventory from provenance.agent and provenance.model (script below), plus MCP server, skill and plugin allowlists from managed policy
AI risk assessment and treatment6.1, 8.2, 8.3MAP 1.5, 4.2; MANAGE 1.2, 1.3CC3.2, CC3.4The agent threat model; risk classes computed by CI from the diff; the autonomy register of unattended loops
Impact assessment6.1, 8.4, A.5.2 to A.5.5MAP 5.1—One assessment of the agent pipeline, plus one per shipped AI feature; the EU AI Act page covers when a feature changes your obligations
Requirements before buildA.6.2.2, A.6.2.3MAP 1.1, 3.3CC8.1 (authorize, design)spec.link and spec.delta in the bundle; the spec, plan and tasks of the artifact chain
Verification and validationA.6.2.4MEASURE 2.1, 2.3, 2.5CC8.1 (test)acceptance and checks results; evals against baseline; mutation score from oracle strength; oracle_changes showing no test was loosened without code-owner review
Controlled deployment and change approvalA.6.2.5MANAGE 1.1CC8.1 (approve, implement)The ruleset with no bypass actors, the separation-of-duties check, a production environment that prevents self-review, and progressive delivery records
Operation and monitoringA.6.2.6, 9.1MEASURE 2.4, 3.1; MANAGE 4.1CC7.2Agent observability: OpenTelemetry from every agent, tool_decision denials and sandbox outcomes, joined to pull requests
Event logs that surviveA.6.2.8MEASURE 2.8CC7.2, CC4.1Telemetry retained in your SIEM, and the signed merge-time archive of bundle, reviews and check runs (actions/attest)
Access for agent identities— (access control sits in ISO/IEC 27001)GOVERN 3.2CC6.1, CC6.2, CC6.3Per-agent GitHub Apps, short-lived tokens and revocation records from agent identity and secrets; deny rules in managed settings
Data entering the agent’s contextA.7.4, A.7.5MEASURE 2.10CC6.7Data classes and approved model routes from data privacy and model hosting; deny rules for secret and regulated paths
Suppliers and pre-trained modelsA.10.2, A.10.3GOVERN 6.1, 6.2; MANAGE 3.1, 3.2CC9.2Vendor reports from procurement; pinned tool versions; eval results on every model change via the new-model playbook
Incidents and the ability to stop an agentA.3.3, A.8.4GOVERN 4.3; MANAGE 2.3, 2.4, 4.3CC7.3, CC7.4, CC7.5The agent incident log and postmortems; the freeze procedure (revoke the identity, pause the loops, push a restrictive policy); findings turned into evals or policy changes
Competence and AI literacy7.2, 7.3, A.4.6GOVERN 2.2CC1.4, CC2.2Training records from upskilling
Internal audit and continual improvement9.2, 9.3, 10.2MEASURE 1.2, 1.3, 2.13; MANAGE 4.2CC4.1, CC4.2Quarterly control metrics, control tests that fail on demand, and the management review minutes that act on them

Two rows need judgment. MANAGE 3.2 says pre-trained models used for development “are monitored as part of AI system regular monitoring and maintenance”: your eval suite runs when the vendor ships a new default model, not only when your code changes. MANAGE 2.4 asks for mechanisms to “supersede, disengage, or deactivate AI systems”: that is your freeze procedure, and an auditor will ask when you last rehearsed it.

How do you turn the table into audit-ready evidence?

Section titled “How do you turn the table into audit-ready evidence?”

The table describes design; a SOC 2 Type II auditor tests operation over a period. These five steps make every row answerable with a query instead of a meeting.

  1. Write the scope statement and the Statement of Applicability rows for the pipeline. Name the coding agents as AI systems the organization uses and list the Annex A controls from the scope table with their justification. Keep it beside the controls, so both change in one pull request.

  2. Make the evidence bundle mandatory and archive it at merge. The gate is on the evidence bundle page and the archive job on regulated industries.

  3. Commit an evidence map that names the query, owner and frequency for each control. Auditors read it first. Adjust owners and repositories:

    controls/ai-evidence-map.yaml
    version: 3
    scope: "Coding agents used in acme/* repositories (AIMS scope statement v2)"
    controls:
    - id: change-approval
    iso42001: ["A.6.2.5"]
    nist_ai_rmf: ["MANAGE 1.1"]
    soc2: ["CC8.1"]
    evidence:
    - "Signed evidence archives for merged PRs (object-locked bucket, 7-year retention)"
    - "Ruleset regulated-default-branch export, bypass_actors = []"
    query: "Merged PRs in period without a passing 'sod' check on the head commit"
    expected: 0
    owner: "@acme/platform"
    frequency: quarterly
    - id: verification
    iso42001: ["A.6.2.4"]
    nist_ai_rmf: ["MEASURE 2.1", "MEASURE 2.3"]
    soc2: ["CC8.1"]
    evidence:
    - "evidence_bundle.acceptance and .checks in each archive"
    - "Mutation score per service from the nightly oracle-strength job"
    query: "Archives with any acceptance result other than pass, or a 'looser' oracle change without code-owner approval"
    expected: 0
    owner: "@acme/quality"
    frequency: quarterly
    - id: model-change
    iso42001: ["A.10.3"]
    nist_ai_rmf: ["MANAGE 3.2", "GOVERN 6.1"]
    soc2: ["CC9.2"]
    evidence:
    - "Eval run against baseline for every model or tool version adopted"
    query: "Models in the inventory with no eval run recorded before first use"
    expected: 0
    owner: "@acme/platform"
    frequency: on-change
    - id: agent-event-logs
    iso42001: ["A.6.2.8"]
    nist_ai_rmf: ["MEASURE 2.8", "MANAGE 4.1"]
    soc2: ["CC7.2"]
    evidence:
    - "OpenTelemetry logs from Claude Code and Codex in the SIEM, 400-day retention"
    query: "Active agent users in the period with no telemetry in the SIEM"
    expected: 0
    owner: "@acme/security"
    frequency: monthly
  4. Generate the agent and model inventory from the archives, not from a spreadsheet. A.4.4 and GOVERN 1.6 ask for an inventory, and the provenance block already records it per change. This lists a quarter’s merged pull requests with agent, model and risk class:

    Terminal window
    # Terminal, read-only. Needs gh and jq.
    gh pr list --repo acme/payments-api --state merged \
    --search "merged:2026-07-01..2026-09-30" --limit 1000 --json number,body \
    | jq -r '.[] | (.body // "") as $b | [.number,
    (($b | capture("agent:[ \\t]*(?<v>[^\\r\\n]+)") | .v) // "none"),
    (($b | capture("model:[ \\t]*(?<v>[^\\r\\n]+)") | .v) // "none"),
    (($b | capture("class:[ \\t]*(?<v>[a-z]+)") | .v) // "none")] | @csv' \
    > inventory-2026-q3.csv

    Group by agent and model. Every model needs an eval run before its first use (the model-change control), and a none row is a change without a bundle, a finding in itself. Search returns at most 1,000 results, so split busy quarters by month. For attested figures, run the same jq filter over the signed archives.

  5. Test each control the way you test an oracle: with cases that must fail. Open three throwaway pull requests: one with no bundle, one with only the owner’s approval and one with a loosened test. Separately, attempt a push from a revoked agent identity and archive the rejected authentication (the audit log entry). Record the failing runs and the rejection as operating evidence for SOC 2 and internal audit (9.2).

How do Claude Code, Codex and Cursor produce the event logs?

Section titled “How do Claude Code, Codex and Cursor produce the event logs?”

The bundle, rulesets and archive are the same for every tool; the event log that A.6.2.8 and CC7.2 need is not. Keep prompt and tool content out of every export: a log that stores prompts also stores their secrets and personal data.

Put the exporter in managed settings so developers cannot turn it off (on Linux, /etc/claude-code/managed-settings.json):

{
"env": {
"CLAUDE_CODE_ENABLE_TELEMETRY": "1",
"OTEL_LOGS_EXPORTER": "otlp",
"OTEL_METRICS_EXPORTER": "otlp",
"OTEL_EXPORTER_OTLP_PROTOCOL": "grpc",
"OTEL_EXPORTER_OTLP_ENDPOINT": "https://otel-collector.internal:4317",
"OTEL_LOG_MANAGED_SETTINGS": "1"
}
}

Claude Code 2.1.283 emits a claude_code.tool_decision event for every tool call: allowed or denied, and whether configuration, a hook or the user decided. Events carry session.id and, when available, user.email, to join a session to provenance.session in a bundle. Prompt text is <REDACTED> and tool arguments are omitted unless you set OTEL_LOG_USER_PROMPTS=1 or OTEL_LOG_TOOL_DETAILS=1; leave both unset, with OTEL_LOG_TOOL_CONTENT, OTEL_LOG_ASSISTANT_RESPONSES and OTEL_LOG_RAW_API_BODIES.

OTEL_LOG_MANAGED_SETTINGS=1 (from v2.1.274, so on both the stable and latest channels on 2026-09-26) adds the redacted content of the resolved managed settings to the claude_code.managed_settings_resolved event, sent at startup and on each change. Set it as above: the SHA-256 digest (managed_settings.resolved_sha256) is sent only with this flag, so machines without it report their sources but no digest. Confirm with a test event. Machines reporting the same digest ran the same policy: operating evidence for the policy row (A.2.2, CC5.3). A SHA-256 of a short policy can be recovered by hashing guesses, so restrict read access to that field in the SIEM. Record the channel with the version (2026-09-26: latest 2.1.283, stable 2.1.274).

Copy-paste prompts for AI management audit preparation

Section titled “Copy-paste prompts for AI management audit preparation”

These prompts read repositories and settings and change no control. They are identical in Claude Code, Codex and Cursor; run them in plan or read-only mode (Claude Code plan mode, Codex --sandbox read-only, Cursor Ask mode). Have a person check every row before it reaches an auditor.

How do you know the evidence will hold up?

Section titled “How do you know the evidence will hold up?”

Evidence holds up when someone other than its producer can check it unaided. Four practices make that true.

  • Independence of the assessor. MEASURE 1.3 asks for “internal experts who did not serve as front-line developers”; clause 9.2 asks for objective internal audit. Have the security or quality team, not the platform team, run the quarterly sample.
  • Metrics with a denominator. Report the share of merged pull requests with a complete bundle, an independent approval of the head commit and a recorded eval for the model. The targets are 100%, and every exception carries a ticket. The canonical metric definitions are on metrics frameworks.
  • The retained record, not the live one. Compute every metric from the signed archives; a pull request edited after merge is not what you attested.
  • A named sign-off. The CTO, as AI management system owner, signs the management review (clause 9.3) that reads these metrics and decides corrective actions (clause 10.2). An agent can draft the pack; a person signs it.

Do not offer ”% of code written by AI” as control evidence: it measures adoption, and none of the three frameworks asks for it.

What breaks when you map the agent pipeline to AI management standards?

Section titled “What breaks when you map the agent pipeline to AI management standards?”
FailureHow it shows upRecovery
The certificate scope excludes the agents. The AIMS covers AI features only.Asked how coding agents are governed, the Statement of Applicability is silent.Extend the scope to AI systems you use, add the A.4.4, A.9 and A.10.3 rows with evidence, and record the change in management review.
The vendor’s report is presented as your control.The system description cites the vendor’s SOC 2 report for change management.Move it to the CC9.2 supplier row; show CC8.1 with your own archives and approvals.
Evals do not run when the model changes. A new vendor default is adopted the next day.The inventory shows a model with no eval run before its first use.Pin versions in managed settings, adopt models through the new-model playbook, and record a nonconformity with a corrective action (clause 10.2).
The event log stores prompts. Someone enabled prompt logging to debug a session.Secrets or personal data in the SIEM; a data-protection finding.Remove the flags, purge under your retention rules, and add the CI check below the table, which fails when any managed file enables OTEL_LOG_USER_PROMPTS, OTEL_LOG_TOOL_DETAILS, OTEL_LOG_TOOL_CONTENT, OTEL_LOG_ASSISTANT_RESPONSES, OTEL_LOG_RAW_API_BODIES or Codex log_user_prompt = true.
The policy exists only as a document.An auditor asks for the control that enforces a clause and gets a PDF.Mark each clause as enforced (managed settings, hooks) or resting on trust, as the usage policy page does; test the enforced ones.
Codex telemetry is quietly switched off. A developer’s own config.toml overrides /etc/codex/config.toml.The agent-event-logs query finds active Codex users with no events in the SIEM.Move [otel] to /etc/codex/managed_config.toml on managed laptops, as in the Codex tab above. Rely on CI runs and the bundle’s provenance for the rest.
The freeze procedure has never run.MANAGE 2.4 or CC7.4 is tested, and nobody knows how to stop an unattended loop.Rehearse it quarterly from the agent incident playbook and archive the rehearsal.

The CI check for the prompt-logging row fails the job when any file under the policy directory, hidden or gitignored ones included, turns content logging on with 1 or true. It also fails closed: a missing directory or an rg that is absent or errors stops the job instead of passing it. Only exit code 1, meaning no match, lets the job through.

Terminal window
dir=infra/agent-policy
test -d "$dir" || { echo "::error::$dir not found"; exit 1; }
status=0
rg -n -i --hidden --no-ignore \
'OTEL_LOG_(USER_PROMPTS|TOOL_DETAILS|TOOL_CONTENT|ASSISTANT_RESPONSES|RAW_API_BODIES)"\s*:\s*"?(1|true)"?|log_user_prompt\s*=\s*true' \
"$dir" || status=$?
case "$status" in
1) echo "No managed file enables content logging." ;;
0) echo "::error::content logging is enabled in the lines above"; exit 1 ;;
*) echo "::error::rg exited $status (missing or failed)"; exit 1 ;;
esac

Where to go next with AI management standards

Section titled “Where to go next with AI management standards”

Frequently asked questions

Are AI coding agents in scope for ISO/IEC 42001?

They are if your AI management system scope says so, and a certification body will ask why they are not. An organization that uses Claude Code, Codex or Cursor is a user of third-party AI systems, so the Annex A controls on tooling, responsible use, suppliers and event logs apply to the agent pipeline. The full life-cycle controls apply to AI features you ship.

Does a vendor's SOC 2 report cover my use of their coding agent?

No. The vendor's report covers the vendor's controls. Your SOC 2 report has to show your own controls over changes the agent makes: authorization, testing, independent approval and deployment under CC8.1, and access for agent identities under CC6. The vendor report is evidence for your vendor-management control, CC9.2.

Do I need separate evidence for ISO/IEC 42001, NIST AI RMF and SOC 2?

No. The same artifacts answer all three: the evidence bundle and its signed archive, branch rules and deployment approvals, eval results, agent telemetry, the agent inventory and the incident log. Keep one control-to-evidence map and cite the relevant clause, subcategory or criterion in each row.