Skip to content

Agentic engineering in regulated industries

Agentic engineering in regulated industries keeps the controls auditors already test (authorized changes, independent approval, test evidence, a complete audit trail) and changes who produces the evidence. Agents write and pre-review changes; a human who neither owns nor authored the change approves the latest commit; CI records everything in an evidence bundle that is archived at merge.

Your bank’s internal audit team is preparing the annual IT general controls review. Last year each of 25 sampled changes had a ticket, a test run and a second person’s approval. This year half of the pull requests were opened by Claude Code or Codex, some approvals came from a review bot, and on a few the engineer who prompted the agent clicked Approve. The auditor’s first question will be: “Who is the second person here?”

This page is for the CTO who has to answer that question, and the tech lead who has to make the answer true in the repository. It is an engineering reading of the control frameworks, not legal or audit advice: agree the control design with your auditor before you rely on it.

Why an agent approving an agent does not satisfy separation of duties

Section titled “Why an agent approving an agent does not satisfy separation of duties”

Separation of duties exists so that one party cannot make a harmful change alone. The frameworks you are audited against state it in terms of accountable people:

  • NIST SP 800-53 Rev. 5, AC-5, requires you to identify duties to separate and “define system access authorizations to support separation of duties”. CM-5(4) Dual Authorization explains that “any changes to selected system components and information cannot occur unless two qualified individuals approve and implement such changes”, and adds that “the individuals are also accountable for the changes”.
  • PCI DSS v4.0.1 requirement 6.2.3.1 (search extract) asks that manual code reviews be performed by “individuals other than the originating code author” who know code-review techniques, and that changes are approved by management before release.
  • For EU financial entities, the DORA-EU risk-management RTS, Commission Delegated Regulation (EU) 2024/1774, Article 17 (search extract), asks for change-management procedures that keep the functions that approve changes independent from the functions that request and implement them.

An agent is not an accountable individual: a review agent prompted by the same engineer is a second opinion, not a second person. GitHub agrees for its own agent: rulesets now “require an additional approval for unattributed Copilot pull requests” by default (public preview), because “requiring one approval usually means two people are involved in a change”, which stops being true when Copilot opens the pull request under its own app identity (GitHub Docs, checked 2026-09-26).

The consequence is a precise rule rather than a ban: agents may author and review; only a person who is neither the change owner nor the author may approve.

Who may do what in an agent-authored change?

Section titled “Who may do what in an agent-authored change?”

Assign each step to a role and say whether an agent may hold it: this is the core of your control narrative.

StepRoleMay an agent do it?Evidence the pipeline keeps
Request the changeRequester (product owner, ticket author)No: authorization comes from a person or an approved backlogTicket or spec linked in spec.link
Own the changeChange owner: the engineer who starts the agent and answers for the resultNoprovenance.human_owner in the evidence bundle
Write the code and testsAuthorYesAgent identity, tool version, model and session in provenance
Pre-review the changeReviewer agent (Claude Code /code-review, codex exec review, Cursor Bugbot)Yes, as a control that produces findingsReview output archived with the bundle
Approve the changeIndependent approverNo: a person who is neither the change owner nor the pull request authorApproving review on the head commit, recorded by GitHub
Approve sensitive pathsCode owner for auth, money, schema, migrationsNoCODEOWNERS review, required by the ruleset
Deploy to productionPipeline, gated by a release approverThe pipeline deploys; the gate is a person other than the one who triggered itDeployment record with environment approval
Change the gate itselfPlatform ownerNoCODEOWNERS on workflows, policy and checker files

Two edge cases need a written rule before the audit, not during it:

  1. No human owner. A scheduled run, an issue-triggered loop or an agent launched from a shared chat channel has no owner. Treat it as GitHub treats unattributed Copilot pull requests: require two independent human approvals.
  2. The owner edits the agent’s branch. Pushing commits makes the owner an author too. The ruleset below requires approval from someone other than the last pusher, so the push re-opens approval.

What do auditors in finance, health and the public sector expect from the pipeline?

Section titled “What do auditors in finance, health and the public sector expect from the pipeline?”

Auditors rarely ask about the agent. For sampled changes they want authorization, testing, independent approval, the deployment record, and proof that nothing bypassed the process. Map each expectation to a pipeline artifact once, and the audit becomes a query.

Sector and frameworkWhat the auditor testsWhere your agent pipeline answers it
Finance, US-listed: SOX IT general controls over program changeEach sampled change authorized, tested and approved by someone other than its developer, who cannot move it to productionspec.link; acceptance and checks in the bundle; the sod check; the production environment with self-review prevented
Finance, EU: DORA-EU Art. 9(4)(e) and RTS 2024/1774 Art. 17 (see CRA, NIS2 and DORA-EU)Documented, risk-based change management; independence between approving and implementing functionsThe risk class in the bundle; CODEOWNERS routing for high-risk paths; the ruleset with no bypass actors
Card payments: PCI DSS v4.0.1 req. 6.2.3, 6.2.3.1, 6.5.1, 6.5.4Bespoke code reviewed before release by someone other than the author; changes documented with approval, testing and back-out; production roles separatedReviewer-agent findings plus independent human approval; risk.rollback; deployment approval by a person who did not trigger it
Health, US: HIPAA Security Rule audit controls (45 CFR 164.312(b)); FDA 21 CFR Part 11 §11.10(e) for GxP systemsActivity on systems with regulated records is recorded; audit trails are computer-generated, time-stamped and do not obscure earlier entriesAgent telemetry to your SIEM; the signed evidence archive; PHI kept out of agent context (see data privacy)
Public sector, US federal: NIST SP 800-53 Rev. 5 (the basis of FedRAMP baselines) CM-3, CM-4, CM-5(4), AC-5, AU-10, SA-10Changes reviewed with security impact analysis, change records retained, dual authorization where selected, non-repudiation of approvalsBundle risk section; archived approvals tied to named accounts; signed attestation of the archive

IEC 62304 change control, NIS2 measures, ISO/IEC 42001 and SOC 2 read from the same artifacts (spec.link, spec.delta, acceptance, oracle_changes); see CRA, NIS2 and DORA-EU and ISO/IEC 42001, NIST AI RMF and SOC 2 evidence.

How do you enforce separation of duties in the repository?

Section titled “How do you enforce separation of duties in the repository?”

Add four controls in this order; the checks must exist before a ruleset can require them. The steps assume GitHub and the evidence bundle, whose CI job is named evidence.

  1. Give every agent its own identity. An agent on a developer’s personal token makes author and owner one account. Run CI agents as a GitHub App or a dedicated machine account with short-lived tokens, as described in agent identity, credentials and secrets. Leave Allow GitHub Actions to create and approve pull requests off (Settings, Actions, General, Workflow permissions) unless a workflow needs to open pull requests, and never let a workflow token approve one.

  2. Add the separation-of-duties check. It passes only when a person on your approver list who is not the change owner, the pull request author, anyone ever assigned, or anyone who authored or committed a commit on the branch has approved the current head commit, and two such people when the change has no human owner. It runs again whenever a review is submitted or dismissed.

    .github/workflows/sod.yml
    name: separation-of-duties
    on:
    pull_request:
    types: [opened, edited, synchronize, reopened, ready_for_review]
    pull_request_review:
    types: [submitted, dismissed]
    permissions:
    contents: read
    issues: read
    pull-requests: read
    jobs:
    sod:
    runs-on: ubuntu-latest
    steps:
    - name: Require an independent human approval of the head commit
    env:
    GH_TOKEN: ${{ github.token }}
    REPO: ${{ github.repository }}
    PR: ${{ github.event.pull_request.number }}
    run: |
    set -euo pipefail
    gh api "repos/$REPO/pulls/$PR" > pr.json
    HEAD_SHA=$(jq -r .head.sha pr.json)
    BASE_SHA=$(jq -r .base.sha pr.json)
    AUTHOR=$(jq -r .user.login pr.json)
    # GitHub reports a machine account as type "User", so the type proves nothing.
    # Approvers come from a CODEOWNERS-protected list read from the base branch.
    # On a 404, gh api still prints GitHub's JSON error body, so test its exit status.
    if ! gh api -H "Accept: application/vnd.github.raw+json" \
    "repos/$REPO/contents/.github/human-approvers?ref=$BASE_SHA" > approvers.raw; then
    echo "::error::.github/human-approvers is missing on the base branch"; exit 1
    fi
    sed -e 's/#.*//' -e 's/^@//' -e 's/[[:space:]]//g' approvers.raw | grep -v '^$' > humans.txt || true
    if [ ! -s humans.txt ]; then
    echo "::error::.github/human-approvers is empty on the base branch"; exit 1
    fi
    owner=$(jq -r '.body // ""' pr.json \
    | sed -nE 's/^[[:space:]]*human_owner:[[:space:]]*"?@?([A-Za-z0-9-]+)"?.*/\1/p' | head -n1)
    if [ -z "$owner" ]; then
    echo "::error::provenance.human_owner is missing (use none for a run without an owner)"; exit 1
    fi
    required=1
    if [ "$owner" = none ]; then required=2
    elif ! jq -e --arg o "$owner" \
    'any(.assignees[]; (.login | ascii_downcase) == ($o | ascii_downcase))' pr.json >/dev/null; then
    echo "::error::human_owner ($owner) must be an assignee of the pull request"; exit 1
    fi
    # human_owner sits in the PR description, which the author can edit, so
    # every commit author, committer and anyone ever assigned is excluded too.
    gh api "repos/$REPO/pulls/$PR/commits" --paginate \
    --jq '.[] | .author.login?, .committer.login? | select(. != null)' > committers.txt
    gh api "repos/$REPO/issues/$PR/events" --paginate \
    --jq '.[] | select(.event == "assigned") | .assignee.login' >> committers.txt
    printf '%s\n%s\n' "$owner" "$AUTHOR" >> committers.txt
    gh api "repos/$REPO/pulls/$PR/reviews" --paginate > reviews.json
    approvers=$(jq -r --arg sha "$HEAD_SHA" '.[]
    | select(.state == "APPROVED" and .commit_id == $sha and .user.type == "User")
    | .user.login' reviews.json | sort -u)
    independent=$(printf '%s\n' "$approvers" | grep -x -i -F -f humans.txt \
    | grep -v -x -i -F -f committers.txt | grep -v '^$' || true)
    count=$(printf '%s\n' "$independent" | grep -c . || true)
    if [ "$count" -lt "$required" ]; then
    echo "::error::$HEAD_SHA needs $required approval(s) from listed humans other than the owner ($owner), the author (@$AUTHOR) or a committer; found $count"; exit 1
    fi
    echo "Independent approval of $HEAD_SHA by: $(echo $independent)"

    The check ignores approvals of older commits and approvals from any account missing from .github/human-approvers, one login per line, read from the base branch so a pull request cannot add its own approver. The account type alone is not enough: GitHub reports a machine user, an ordinary account an agent signs in to, as type User; only GitHub Apps show as Bot. List people, never service accounts. A missing or empty file fails the check. human_owner is author-controlled: the author, or the owner, can rewrite it, and the edited trigger re-runs the check. So the owner must also be an assignee, and everyone ever assigned is excluded, along with every commit author and committer. Assign the change owner when the agent opens the pull request. The owner is still self-declared: an owner who stays unassigned can name a colleague who is an assignee and then approve. So take the owner from a source the author cannot edit, such as the agent’s session log or the launching actor your agent app records, and have the archive job and your control test cross-check human_owner against it. For a run with no human owner, such as a scheduled or ticket-triggered agent, write human_owner: none: the check then requires two distinct independent approvers of the head commit. The ruleset cannot do this for you, because its approval count applies to every pull request on the branch.

  3. Create the branch ruleset. It requires one approval, code-owner review and approval of the last push by someone other than its pusher, dismisses stale approvals and requires three checks: evidence, sod and agent-review, the reviewer job described under the tool tabs below. Requiring agent-review means a pull request cannot merge before the review of its head commit succeeds, so the archive job finds the reviewer output. The empty bypass list lets you tell an auditor the population is complete. Each check is pinned to integration_id 15368, the GitHub Actions app: an unpinned required check accepts a status of that name from any account that can write commit statuses, including the agent’s own GitHub App or token, which could post sod as passing itself.

    {
    "name": "regulated-default-branch",
    "target": "branch",
    "enforcement": "active",
    "bypass_actors": [],
    "conditions": { "ref_name": { "include": ["~DEFAULT_BRANCH"], "exclude": [] } },
    "rules": [
    { "type": "deletion" },
    { "type": "non_fast_forward" },
    { "type": "pull_request", "parameters": {
    "required_approving_review_count": 1,
    "dismiss_stale_reviews_on_push": true,
    "require_code_owner_review": true,
    "require_last_push_approval": true,
    "required_review_thread_resolution": true } },
    { "type": "required_status_checks", "parameters": {
    "strict_required_status_checks_policy": true,
    "required_status_checks": [
    { "context": "evidence", "integration_id": 15368 },
    { "context": "sod", "integration_id": 15368 },
    { "context": "agent-review", "integration_id": 15368 } ] } }
    ]
    }
    Terminal window
    # Terminal, as a repository admin
    gh api -X POST repos/acme/payments-api/rulesets --input ruleset.json
  4. Gate production deployment on a second person. Create a production environment with the release approvers as required reviewers and self-review prevented. On GitHub Free, Pro and Team, required reviewers work only in public repositories, so private regulated code needs GitHub Enterprise for this step.

    Terminal window
    # Terminal, as a repository admin. 4532992 stands for the numeric id of your
    # release-approvers team: gh api orgs/acme/teams/release-approvers --jq .id
    gh api -X PUT repos/acme/payments-api/environments/production --input - <<'JSON'
    { "prevent_self_review": true,
    "reviewers": [ { "type": "Team", "id": 4532992 } ] }
    JSON

How do you know the separation-of-duties check works?

Section titled “How do you know the separation-of-duties check works?”

Test it like any oracle, with cases that must fail. On a throwaway pull request from an agent identity, every case except the last must leave sod red or the merge blocked.

CaseExpected result
Bundle without human_ownerFails: owner missing
Only the change owner approvesFails: no independent approval
A review bot approvesFails: Bot approvals are ignored
A machine user (type User) that an agent signs in to approvesFails: it is not in .github/human-approvers
human_owner edited to another name, then the real owner approvesFails when the agent assigned the owner at pull request creation: the owner is an assignee and is excluded
The agent’s account posts a success status named sodStill blocked: the ruleset accepts sod only from GitHub Actions
human_owner: none and one independent person approves the head commitFails until a second independent person approves it
An independent person approves, then the agent pushes againFails until that person, or another, approves the new head commit
Merge attempted before the agent-review run on the head commit succeedsBlocked: agent-review is a required check
An independent person approves the head commitPasses

Record the run links in your control-testing file: auditors trust a control you can show failing.

How do you keep the record after the merge?

Section titled “How do you keep the record after the merge?”

Anyone with write access can edit a pull request’s description, and therefore the bundle, after the merge, and Actions artifacts expire. Snapshot the record when the pull request merges, sign it with actions/attest, and copy it to storage your records policy controls.

.github/workflows/archive-evidence.yml
name: archive-evidence
on:
pull_request:
types: [closed]
jobs:
archive:
if: github.event.pull_request.merged == true
runs-on: ubuntu-latest
permissions:
contents: read
pull-requests: read
issues: read
checks: read
actions: read
id-token: write
attestations: write
artifact-metadata: write
steps:
- name: Snapshot the bundle, reviews, comments, check runs and reviewer output
id: snapshot
env:
GH_TOKEN: ${{ github.token }}
REPO: ${{ github.repository }}
PR: ${{ github.event.pull_request.number }}
SHA: ${{ github.event.pull_request.head.sha }}
# The workflow that runs the reviewer agent and uploads evidence/
# as the artifact "agent-review"
REVIEW_WORKFLOW: agent-review.yml
run: |
set -euo pipefail
D="evidence-$PR"; mkdir -p "$D"
gh api "repos/$REPO/pulls/$PR" > "$D/pull-request.json"
gh api "repos/$REPO/pulls/$PR/reviews" --paginate --slurp > "$D/reviews.json"
gh api "repos/$REPO/pulls/$PR/comments" --paginate --slurp > "$D/review-comments.json"
gh api "repos/$REPO/issues/$PR/comments" --paginate --slurp > "$D/issue-comments.json"
gh api "repos/$REPO/commits/$SHA/check-runs" --paginate --slurp > "$D/check-runs.json"
# A missing reviewer output must not cost the rest of the record: archive
# the snapshot, record the gap as an exception, fail the job after upload.
run_id=$(gh run list -R "$REPO" --workflow "$REVIEW_WORKFLOW" --commit "$SHA" \
--status success --limit 1 --json databaseId --jq '.[0].databaseId // empty' || true)
if [ -n "$run_id" ] && gh run download "$run_id" -R "$REPO" -n agent-review -D "$D/agent-review"; then
echo "reviewer_output=present" >> "$GITHUB_OUTPUT"
else
echo "EXCEPTION: no agent-review output from a successful $REVIEW_WORKFLOW run on $SHA at merge" \
> "$D/EXCEPTION-agent-review-missing.txt"
echo "reviewer_output=missing" >> "$GITHUB_OUTPUT"
fi
tar czf "$D.tgz" "$D"
- uses: actions/attest@v4
with:
subject-path: evidence-${{ github.event.pull_request.number }}.tgz
- uses: actions/upload-artifact@v7
with:
name: evidence-${{ github.event.pull_request.number }}
path: evidence-${{ github.event.pull_request.number }}.tgz
- name: Fail when the reviewer output is missing
if: steps.snapshot.outputs.reviewer_output == 'missing'
env:
SHA: ${{ github.event.pull_request.head.sha }}
run: |
echo "::error::no agent-review output for $SHA; the archive records it as a control exception"
exit 1

The snapshot holds the pull request, its reviews, inline review comments, conversation comments, check runs and the reviewer agent’s output from the agent-review artifact. If that output is missing, the job still archives and attests everything else, adds EXCEPTION-agent-review-missing.txt to the archive and then fails, so the gap shows as a red run and a recorded exception rather than a lost record. The attestation binds the archive’s digest to the workflow run that produced it, and gh attestation verify checks it. Add a final step that copies it to write-once storage, such as an object-lock bucket, with your regulator’s retention period. Artifact attestations in private repositories need GitHub Enterprise Cloud.

How do Claude Code, Codex and Cursor fit the control?

Section titled “How do Claude Code, Codex and Cursor fit the control?”

The ruleset, checks and archive are identical for every tool; version pinning, reviewer output and the activity log differ.

Version control of the tool. Claude Code has two release channels: latest (2.1.283 on 2026-09-26) and stable (2.1.274), which is “typically about a week behind” and skips releases with major regressions. Set autoUpdatesChannel: "stable" in managed settings and record the version in provenance.agent. The channel governs workstations only: in CI, install a pinned version (npm install -g @anthropic-ai/claude-code@2.1.283, or the stable 2.1.274) and record claude --version in provenance.agent. The models hub lists the default model.

Reviewer output as evidence. Run the review headless and keep the JSON result next to the bundle:

Terminal window
# CI job on the pull request (Claude Code 2.1.283). Check out with
# fetch-depth: 0, or run git fetch origin main, so origin/main exists.
mkdir -p evidence
claude -p "Review the diff against origin/main for security and correctness. List findings with file:line." \
--bare --setting-sources "" --strict-mcp-config \
--output-format json --permission-mode dontAsk \
--allowedTools "Read,Grep,Glob,Bash(git diff:*)" \
> evidence/claude-review.json

Keep the prompt straight after -p: the variadic --allowedTools would swallow a prompt placed after it. --bare skips hooks and CLAUDE.md discovery, --setting-sources "" loads no settings file at all, so the branch’s project and local settings stay unloaded, and --strict-mcp-config ignores the branch’s MCP servers. Skills still resolve through /skill-name under --bare, so the prompt names none. With --bare the job authenticates only through ANTHROPIC_API_KEY (or a cloud provider’s credentials); an OAuth token is not read. Run it on pull_request with contents: read and a checkout with persist-credentials: false, never on pull_request_target. Name the workflow agent-review.yml and its job agent-review, the check the ruleset requires, and end the job with actions/upload-artifact uploading evidence/ as agent-review, so the archive job collects it at merge.

Activity log. Set CLAUDE_CODE_ENABLE_TELEMETRY=1 and the standard OTEL_* exporter variables in managed settings, so developers cannot turn them off. The Compliance API is an Enterprise-plan feature.

Across all three, keep regulated data out of the agent’s context, or it ends up in session logs and the archive; see where the model runs and data privacy and enterprise policies.

These prompts only read evidence. Run them with read-only permissions.

How do you prove the control keeps working?

Section titled “How do you prove the control keeps working?”

Auditors test operating effectiveness over the period, so track three numbers per quarter:

MetricDefinitionTarget
Bundle completenessMerged pull requests to in-scope repositories with an archived bundle and a passing evidence check on the head commit, divided by all merged pull requests to those repositories, counted from the default branch history or the merged-pull-request list (gh api), not from the archives100%; every exception has a ticket
Independent approval rateMerged pull requests whose archive shows a human approval of the head commit by someone other than the owner and the author, divided by all merged pull requests100%
Emergency changes approved in timeEmergency changes with a retrospective independent approval within the window your policy sets, divided by all emergency changes100%

Take each denominator from the default branch history or gh api repos/acme/payments-api/pulls?state=closed filtered to merged pull requests, and match it against the archive list: a merge whose archive job never ran is missing from the archives, so a ratio computed only from them always reads 100%. Take the numerators from the archives. Change failure rate and lead time are on metrics frameworks.

What breaks in regulated agent pipelines, and how do you recover?

Section titled “What breaks in regulated agent pipelines, and how do you recover?”
FailureHow it shows upRecovery
The owner approves the agent’s work, because the author is a bot.Sampled changes where the approver equals human_owner.Log each as a control exception, get a retrospective independent approval, and make sod a required check.
A review bot counts as the approval. A workflow token, app account or machine user submits an approving review.Approvals from accounts of type Bot or missing from .github/human-approvers; the Allow GitHub Actions to create and approve pull requests setting is on.Turn the setting off at organization level, remove approval rights from agent apps, and re-run the gap prompt.
Approval of an older commit. The agent pushed a fix after approval.The approving review’s commit_id differs from the merged head.Enable dismiss_stale_reviews_on_push and require_last_push_approval; the sod check already compares the commit.
Evidence edited after merge. Someone “tidies” a bundle months later.The live pull request disagrees with the archive.The signed archive is the record; the live page is not. Verify with gh attestation verify and note the edit.
Bypass used for a hotfix, never regularized. An admin merged directly during an incident.Ruleset bypass events, or a production change with no pull request.Write an emergency path: a named break-glass role, a pull request after the fact, and a retrospective independent approval in the policy window. See when an agent causes an incident.
The archive job failed at merge. A GitHub API error or runner outage stopped archive-evidence, or the reviewer output was missing.A merged pull request whose archive-evidence run is red or absent, or an archive holding EXCEPTION-agent-review-missing.txt.Re-run the job at once and log a control exception: a late snapshot reads the live pull request, which anyone with write access can edit, so compare the description with its edit history and note any change made after the merge. For missing reviewer output, re-run the agent-review run on the merged head commit first, then the archive job.
The agent weakened the test that proves the change.oracle_changes with direction looser, or a test edit not declared.The bundle forces the change to high risk and code-owner review. Keep the oracle files owned by people, as described on protecting the oracle.

Where to go next with regulated agent pipelines

Section titled “Where to go next with regulated agent pipelines”

Frequently asked questions

Can one AI agent approve a change another agent wrote?

Not for separation of duties. An agent review is a control that produces evidence, but the frameworks auditors test against name accountable individuals: an approver who is neither the author nor the person who owns the change. Keep the agent review and add a human approval of the latest commit.

Who is the author of an agent-written change for audit purposes?

The human who owns the task and started the agent, recorded as the human owner in the evidence bundle. That person cannot approve the change. When no human owner exists, for example a scheduled agent run, require two human approvals.

What do auditors sample when agents write the code?

The same things as before: authorization for the change, test evidence, independent approval before production, the deployment record, and proof that no change reached production outside the process. The evidence bundle, branch rules and deployment approvals produce each item per pull request.