ISO/IEC 42001, NIST AI RMF and SOC 2 evidence from the agent pipeline
An agent pipeline built on the evidence bundle already produces most of what ISO/IEC 42001, the NIST AI Risk Management Framework and a SOC 2 examination ask for. Specs, acceptance results, evals, approvals, telemetry and provenance map to named clauses, subcategories and criteria. The work is mapping them once, archiving them and testing the controls.
Your SOC 2 Type II period closes in December. A customer questionnaire asks whether you are “aligned with the NIST AI RMF”, sales wants to say “ISO/IEC 42001 in progress”, and your auditor has noticed that a third of this year’s merged pull requests came from Claude Code or Codex. Nobody wants a second compliance program. You need to know which artifacts your pipeline already writes count as evidence, and where the gaps are.
This page is for the CTO who owns that answer. It is an engineering reading, not audit or legal advice: agree the control design with your auditor or certification body. It builds on the separation-of-duties controls and merge-time archive in agentic engineering in regulated industries.
What this page gives a CTO preparing for an AI management audit
Section titled “What this page gives a CTO preparing for an AI management audit”- A scope decision: which ISO/IEC 42001 controls apply when you use agents, and which when you ship AI features.
- A control-to-evidence table mapping each pipeline artifact to ISO/IEC 42001, NIST AI RMF and SOC 2.
- An evidence map file with the query, owner and frequency for every control.
- Telemetry settings for Claude Code, Codex and Cursor, an inventory script and three copy-paste prompts.
What do ISO/IEC 42001, NIST AI RMF and SOC 2 each ask for?
Section titled “What do ISO/IEC 42001, NIST AI RMF and SOC 2 each ask for?”The three frameworks answer different questions, so one evidence set can serve all of them.
| Framework | What it is | What you get | What the assessor looks at |
|---|---|---|---|
| ISO/IEC 42001:2023 | A certifiable AI management system: clauses 4 to 10 plus 38 reference controls in Annex A, under nine objectives from A.2 to A.10 | A certificate from an accredited certification body | Your AI risk assessment and treatment, the Statement of Applicability, and evidence that the controls run |
| NIST AI RMF 1.0 (NIST AI 100-1) | A voluntary framework with four functions, GOVERN, MAP, MEASURE and MANAGE, broken into categories and subcategories | No certificate; a self-declared alignment that customer questionnaires ask about | Whatever the customer or internal audit samples: usually policies, roles, inventory, tests, monitoring and incidents |
| SOC 2 | An AICPA attestation report on your controls against the Trust Services Criteria; the Common Criteria CC1 to CC9 cover security | A Type I report (design at a point in time) or a Type II report (operation over a period) | A sample of changes, access grants and incidents, tested against your described controls |
Two related documents help with scoping: NIST AI 600-1, the Generative AI Profile (2024-07-26), which lists twelve generative-AI risks, and ISO/IEC 42005:2025 on impact assessment (both SECONDARY, from search extracts of the NIST and ISO pages, 2026-09-26). The mapping below uses AI RMF 1.0 as published.
Are your coding agents in scope for an AI management system?
Section titled “Are your coding agents in scope for an AI management system?”Decide this first, in writing, because it sets which Annex A controls appear in your Statement of Applicability. Running Claude Code, Codex or Cursor makes you a user of third-party AI systems; shipping a chatbot, recommender or agent feature also makes you a developer. Most software companies are both.
| Situation | ISO/IEC 42001 controls that carry the weight | NIST AI RMF focus | SOC 2 focus |
|---|---|---|---|
| You use coding agents to build software that contains no AI | A.4.4 tooling resources, A.9.2 to A.9.4 responsible and intended use, A.10.3 suppliers, A.6.2.8 event logs, A.3.2 roles | GOVERN 1, 2, 6; MANAGE 3 (third-party resources and pre-trained models) | CC8.1 changes, CC6 access for agent identities, CC7.2 monitoring, CC9.2 vendors |
| You also ship AI features built with those agents | All of the above, plus the life-cycle controls A.6.2.2 to A.6.2.7, data controls A.7, and impact assessment A.5 for each feature | All four functions per AI system, with MEASURE 2 evaluations per release | The same, with processing-integrity criteria if your report includes them |
You may exclude the agent pipeline, but every excluded Annex A control needs a justification. “Developers use coding agents under our standard change process” can be tested; “coding agents are out of scope” cannot.
The control-to-evidence table
Section titled “The control-to-evidence table”Each row is one control objective, the clause, subcategory and criterion it satisfies, and the artifact that answers it, built elsewhere on this site: the evidence bundle, the risk classes in governance and autonomy, managed policy, telemetry, evals and the incident log.
| Control objective | ISO/IEC 42001 | NIST AI RMF | SOC 2 | Evidence from the agent pipeline |
|---|---|---|---|---|
| AI policy that is enforced, not only published | 5.2, A.2.2, A.2.3 | GOVERN 1.1, 1.2, 1.4 | CC5.3 | The AI usage policy plus managed settings files (/etc/claude-code/managed-settings.json, /etc/codex/requirements.toml, Cursor admin settings) versioned in git; Claude Code’s managed_settings_resolved event shows which policy each session ran |
| Roles, accountability and human oversight | 5.3, A.3.2 | GOVERN 2.1, 3.2; MAP 3.5 | CC1.3, CC1.5 | The operating model RACI; provenance.human_owner in every bundle; CODEOWNERS for sensitive paths; the rule that an agent may author and review but never approve |
| Inventory of AI tools, models and extensions | A.4.2, A.4.4 | GOVERN 1.6; MAP 4.1 | CC9.2; system description | The quarterly inventory from provenance.agent and provenance.model (script below), plus MCP server, skill and plugin allowlists from managed policy |
| AI risk assessment and treatment | 6.1, 8.2, 8.3 | MAP 1.5, 4.2; MANAGE 1.2, 1.3 | CC3.2, CC3.4 | The agent threat model; risk classes computed by CI from the diff; the autonomy register of unattended loops |
| Impact assessment | 6.1, 8.4, A.5.2 to A.5.5 | MAP 5.1 | — | One assessment of the agent pipeline, plus one per shipped AI feature; the EU AI Act page covers when a feature changes your obligations |
| Requirements before build | A.6.2.2, A.6.2.3 | MAP 1.1, 3.3 | CC8.1 (authorize, design) | spec.link and spec.delta in the bundle; the spec, plan and tasks of the artifact chain |
| Verification and validation | A.6.2.4 | MEASURE 2.1, 2.3, 2.5 | CC8.1 (test) | acceptance and checks results; evals against baseline; mutation score from oracle strength; oracle_changes showing no test was loosened without code-owner review |
| Controlled deployment and change approval | A.6.2.5 | MANAGE 1.1 | CC8.1 (approve, implement) | The ruleset with no bypass actors, the separation-of-duties check, a production environment that prevents self-review, and progressive delivery records |
| Operation and monitoring | A.6.2.6, 9.1 | MEASURE 2.4, 3.1; MANAGE 4.1 | CC7.2 | Agent observability: OpenTelemetry from every agent, tool_decision denials and sandbox outcomes, joined to pull requests |
| Event logs that survive | A.6.2.8 | MEASURE 2.8 | CC7.2, CC4.1 | Telemetry retained in your SIEM, and the signed merge-time archive of bundle, reviews and check runs (actions/attest) |
| Access for agent identities | — (access control sits in ISO/IEC 27001) | GOVERN 3.2 | CC6.1, CC6.2, CC6.3 | Per-agent GitHub Apps, short-lived tokens and revocation records from agent identity and secrets; deny rules in managed settings |
| Data entering the agent’s context | A.7.4, A.7.5 | MEASURE 2.10 | CC6.7 | Data classes and approved model routes from data privacy and model hosting; deny rules for secret and regulated paths |
| Suppliers and pre-trained models | A.10.2, A.10.3 | GOVERN 6.1, 6.2; MANAGE 3.1, 3.2 | CC9.2 | Vendor reports from procurement; pinned tool versions; eval results on every model change via the new-model playbook |
| Incidents and the ability to stop an agent | A.3.3, A.8.4 | GOVERN 4.3; MANAGE 2.3, 2.4, 4.3 | CC7.3, CC7.4, CC7.5 | The agent incident log and postmortems; the freeze procedure (revoke the identity, pause the loops, push a restrictive policy); findings turned into evals or policy changes |
| Competence and AI literacy | 7.2, 7.3, A.4.6 | GOVERN 2.2 | CC1.4, CC2.2 | Training records from upskilling |
| Internal audit and continual improvement | 9.2, 9.3, 10.2 | MEASURE 1.2, 1.3, 2.13; MANAGE 4.2 | CC4.1, CC4.2 | Quarterly control metrics, control tests that fail on demand, and the management review minutes that act on them |
Two rows need judgment. MANAGE 3.2 says pre-trained models used for development “are monitored as part of AI system regular monitoring and maintenance”: your eval suite runs when the vendor ships a new default model, not only when your code changes. MANAGE 2.4 asks for mechanisms to “supersede, disengage, or deactivate AI systems”: that is your freeze procedure, and an auditor will ask when you last rehearsed it.
How do you turn the table into audit-ready evidence?
Section titled “How do you turn the table into audit-ready evidence?”The table describes design; a SOC 2 Type II auditor tests operation over a period. These five steps make every row answerable with a query instead of a meeting.
-
Write the scope statement and the Statement of Applicability rows for the pipeline. Name the coding agents as AI systems the organization uses and list the Annex A controls from the scope table with their justification. Keep it beside the controls, so both change in one pull request.
-
Make the evidence bundle mandatory and archive it at merge. The gate is on the evidence bundle page and the archive job on regulated industries.
-
Commit an evidence map that names the query, owner and frequency for each control. Auditors read it first. Adjust owners and repositories:
controls/ai-evidence-map.yaml version: 3scope: "Coding agents used in acme/* repositories (AIMS scope statement v2)"controls:- id: change-approvaliso42001: ["A.6.2.5"]nist_ai_rmf: ["MANAGE 1.1"]soc2: ["CC8.1"]evidence:- "Signed evidence archives for merged PRs (object-locked bucket, 7-year retention)"- "Ruleset regulated-default-branch export, bypass_actors = []"query: "Merged PRs in period without a passing 'sod' check on the head commit"expected: 0owner: "@acme/platform"frequency: quarterly- id: verificationiso42001: ["A.6.2.4"]nist_ai_rmf: ["MEASURE 2.1", "MEASURE 2.3"]soc2: ["CC8.1"]evidence:- "evidence_bundle.acceptance and .checks in each archive"- "Mutation score per service from the nightly oracle-strength job"query: "Archives with any acceptance result other than pass, or a 'looser' oracle change without code-owner approval"expected: 0owner: "@acme/quality"frequency: quarterly- id: model-changeiso42001: ["A.10.3"]nist_ai_rmf: ["MANAGE 3.2", "GOVERN 6.1"]soc2: ["CC9.2"]evidence:- "Eval run against baseline for every model or tool version adopted"query: "Models in the inventory with no eval run recorded before first use"expected: 0owner: "@acme/platform"frequency: on-change- id: agent-event-logsiso42001: ["A.6.2.8"]nist_ai_rmf: ["MEASURE 2.8", "MANAGE 4.1"]soc2: ["CC7.2"]evidence:- "OpenTelemetry logs from Claude Code and Codex in the SIEM, 400-day retention"query: "Active agent users in the period with no telemetry in the SIEM"expected: 0owner: "@acme/security"frequency: monthly -
Generate the agent and model inventory from the archives, not from a spreadsheet. A.4.4 and GOVERN 1.6 ask for an inventory, and the provenance block already records it per change. This lists a quarter’s merged pull requests with agent, model and risk class:
Terminal window # Terminal, read-only. Needs gh and jq.gh pr list --repo acme/payments-api --state merged \--search "merged:2026-07-01..2026-09-30" --limit 1000 --json number,body \| jq -r '.[] | (.body // "") as $b | [.number,(($b | capture("agent:[ \\t]*(?<v>[^\\r\\n]+)") | .v) // "none"),(($b | capture("model:[ \\t]*(?<v>[^\\r\\n]+)") | .v) // "none"),(($b | capture("class:[ \\t]*(?<v>[a-z]+)") | .v) // "none")] | @csv' \> inventory-2026-q3.csvGroup by agent and model. Every model needs an eval run before its first use (the
model-changecontrol), and anonerow is a change without a bundle, a finding in itself. Search returns at most 1,000 results, so split busy quarters by month. For attested figures, run the same jq filter over the signed archives. -
Test each control the way you test an oracle: with cases that must fail. Open three throwaway pull requests: one with no bundle, one with only the owner’s approval and one with a loosened test. Separately, attempt a push from a revoked agent identity and archive the rejected authentication (the audit log entry). Record the failing runs and the rejection as operating evidence for SOC 2 and internal audit (9.2).
How do Claude Code, Codex and Cursor produce the event logs?
Section titled “How do Claude Code, Codex and Cursor produce the event logs?”The bundle, rulesets and archive are the same for every tool; the event log that A.6.2.8 and CC7.2 need is not. Keep prompt and tool content out of every export: a log that stores prompts also stores their secrets and personal data.
Put the exporter in managed settings so developers cannot turn it off (on Linux, /etc/claude-code/managed-settings.json):
{ "env": { "CLAUDE_CODE_ENABLE_TELEMETRY": "1", "OTEL_LOGS_EXPORTER": "otlp", "OTEL_METRICS_EXPORTER": "otlp", "OTEL_EXPORTER_OTLP_PROTOCOL": "grpc", "OTEL_EXPORTER_OTLP_ENDPOINT": "https://otel-collector.internal:4317", "OTEL_LOG_MANAGED_SETTINGS": "1" }}Claude Code 2.1.283 emits a claude_code.tool_decision event for every tool call: allowed or denied, and whether configuration, a hook or the user decided. Events carry session.id and, when available, user.email, to join a session to provenance.session in a bundle. Prompt text is <REDACTED> and tool arguments are omitted unless you set OTEL_LOG_USER_PROMPTS=1 or OTEL_LOG_TOOL_DETAILS=1; leave both unset, with OTEL_LOG_TOOL_CONTENT, OTEL_LOG_ASSISTANT_RESPONSES and OTEL_LOG_RAW_API_BODIES.
OTEL_LOG_MANAGED_SETTINGS=1 (from v2.1.274, so on both the stable and latest channels on 2026-09-26) adds the redacted content of the resolved managed settings to the claude_code.managed_settings_resolved event, sent at startup and on each change. Set it as above: the SHA-256 digest (managed_settings.resolved_sha256) is sent only with this flag, so machines without it report their sources but no digest. Confirm with a test event. Machines reporting the same digest ran the same policy: operating evidence for the policy row (A.2.2, CC5.3). A SHA-256 of a short policy can be recovered by hashing guesses, so restrict read access to that field in the SIEM. Record the channel with the version (2026-09-26: latest 2.1.283, stable 2.1.274).
Put the exporter in the [otel] table of a config.toml:
[otel]environment = "prod"log_user_prompt = false # default; prompts are not exported
[otel.exporter.otlp-http]endpoint = "https://otel-collector.internal:4318/v1/logs"protocol = "binary"Codex CLI 0.157.1 emits codex.conversation_starts (approval and sandbox policy, reasoning effort, MCP servers), codex.tool_decision (approved or denied, and by whom) and codex.sandbox_outcome, among others. The log exporter defaults to none.
Where you put that table decides whether it is a control. In 0.157.1 /etc/codex/config.toml and the cloud-managed configuration rank below the developer’s own ~/.codex/config.toml, and requirements.toml has no otel key to lock the exporter (checked against the 0.157.1 source). So:
- Managed laptops: put
[otel]in the legacy/etc/codex/managed_config.toml(or macOS MDM managed config), which 0.157.1 loads above every other layer, including--config. Confirm with a test event and note the legacy mechanism in the evidence map. - Unmanaged laptops: treat Codex telemetry as best-effort and say so in the evidence map.
- CI: the authoritative record. Point
CODEX_HOMEat a directory your workflow writes, pin the CLI (npm install -g @openai/codex@0.157.1) and keep sandbox modes and MCP allowlists in/etc/codex/requirements.toml.
cursor.com was unreachable on 2026-09-26, so this page does not describe Cursor’s admin exports. Ask your Cursor admin which usage and audit data your plan exports, record the dated answer in the evidence map, and treat Cursor as a manual evidence source until then.
Bugbot review comments land on the pull request, so the merge-time archive captures them in reviews.json. If you enable Cursor’s PR Routing & Approval, which can approve low-risk PRs, make the separation-of-duties check ignore its approvals: an agent approval never counts as the independent approval CC8.1 samples.
Copy-paste prompts for AI management audit preparation
Section titled “Copy-paste prompts for AI management audit preparation”These prompts read repositories and settings and change no control. They are identical in Claude Code, Codex and Cursor; run them in plan or read-only mode (Claude Code plan mode, Codex --sandbox read-only, Cursor Ask mode). Have a person check every row before it reaches an auditor.
How do you know the evidence will hold up?
Section titled “How do you know the evidence will hold up?”Evidence holds up when someone other than its producer can check it unaided. Four practices make that true.
- Independence of the assessor. MEASURE 1.3 asks for “internal experts who did not serve as front-line developers”; clause 9.2 asks for objective internal audit. Have the security or quality team, not the platform team, run the quarterly sample.
- Metrics with a denominator. Report the share of merged pull requests with a complete bundle, an independent approval of the head commit and a recorded eval for the model. The targets are 100%, and every exception carries a ticket. The canonical metric definitions are on metrics frameworks.
- The retained record, not the live one. Compute every metric from the signed archives; a pull request edited after merge is not what you attested.
- A named sign-off. The CTO, as AI management system owner, signs the management review (clause 9.3) that reads these metrics and decides corrective actions (clause 10.2). An agent can draft the pack; a person signs it.
Do not offer ”% of code written by AI” as control evidence: it measures adoption, and none of the three frameworks asks for it.
What breaks when you map the agent pipeline to AI management standards?
Section titled “What breaks when you map the agent pipeline to AI management standards?”| Failure | How it shows up | Recovery |
|---|---|---|
| The certificate scope excludes the agents. The AIMS covers AI features only. | Asked how coding agents are governed, the Statement of Applicability is silent. | Extend the scope to AI systems you use, add the A.4.4, A.9 and A.10.3 rows with evidence, and record the change in management review. |
| The vendor’s report is presented as your control. | The system description cites the vendor’s SOC 2 report for change management. | Move it to the CC9.2 supplier row; show CC8.1 with your own archives and approvals. |
| Evals do not run when the model changes. A new vendor default is adopted the next day. | The inventory shows a model with no eval run before its first use. | Pin versions in managed settings, adopt models through the new-model playbook, and record a nonconformity with a corrective action (clause 10.2). |
| The event log stores prompts. Someone enabled prompt logging to debug a session. | Secrets or personal data in the SIEM; a data-protection finding. | Remove the flags, purge under your retention rules, and add the CI check below the table, which fails when any managed file enables OTEL_LOG_USER_PROMPTS, OTEL_LOG_TOOL_DETAILS, OTEL_LOG_TOOL_CONTENT, OTEL_LOG_ASSISTANT_RESPONSES, OTEL_LOG_RAW_API_BODIES or Codex log_user_prompt = true. |
| The policy exists only as a document. | An auditor asks for the control that enforces a clause and gets a PDF. | Mark each clause as enforced (managed settings, hooks) or resting on trust, as the usage policy page does; test the enforced ones. |
Codex telemetry is quietly switched off. A developer’s own config.toml overrides /etc/codex/config.toml. | The agent-event-logs query finds active Codex users with no events in the SIEM. | Move [otel] to /etc/codex/managed_config.toml on managed laptops, as in the Codex tab above. Rely on CI runs and the bundle’s provenance for the rest. |
| The freeze procedure has never run. | MANAGE 2.4 or CC7.4 is tested, and nobody knows how to stop an unattended loop. | Rehearse it quarterly from the agent incident playbook and archive the rehearsal. |
The CI check for the prompt-logging row fails the job when any file under the policy directory, hidden or gitignored ones included, turns content logging on with 1 or true. It also fails closed: a missing directory or an rg that is absent or errors stops the job instead of passing it. Only exit code 1, meaning no match, lets the job through.
dir=infra/agent-policytest -d "$dir" || { echo "::error::$dir not found"; exit 1; }status=0rg -n -i --hidden --no-ignore \ 'OTEL_LOG_(USER_PROMPTS|TOOL_DETAILS|TOOL_CONTENT|ASSISTANT_RESPONSES|RAW_API_BODIES)"\s*:\s*"?(1|true)"?|log_user_prompt\s*=\s*true' \ "$dir" || status=$?case "$status" in 1) echo "No managed file enables content logging." ;; 0) echo "::error::content logging is enabled in the lines above"; exit 1 ;; *) echo "::error::rg exited $status (missing or failed)"; exit 1 ;;esacWhere to go next with AI management standards
Section titled “Where to go next with AI management standards”Frequently asked questions
Are AI coding agents in scope for ISO/IEC 42001?
They are if your AI management system scope says so, and a certification body will ask why they are not. An organization that uses Claude Code, Codex or Cursor is a user of third-party AI systems, so the Annex A controls on tooling, responsible use, suppliers and event logs apply to the agent pipeline. The full life-cycle controls apply to AI features you ship.
Does a vendor's SOC 2 report cover my use of their coding agent?
No. The vendor's report covers the vendor's controls. Your SOC 2 report has to show your own controls over changes the agent makes: authorization, testing, independent approval and deployment under CC8.1, and access for agent identities under CC6. The vendor report is evidence for your vendor-management control, CC9.2.
Do I need separate evidence for ISO/IEC 42001, NIST AI RMF and SOC 2?
No. The same artifacts answer all three: the evidence bundle and its signed archive, branch rules and deployment approvals, eval results, agent telemetry, the agent inventory and the incident log. Keep one control-to-evidence map and cite the relevant clause, subcategory or criterion in each row.