Codex Security and the agents command center
Codex Security is OpenAI’s application security agent that helps teams find, confirm and fix vulnerabilities; @codex security review asks for a security-focused pull request review, and codex agents (/agents) supervises the fix sessions. A finding closes only when a failing exploit test passes after the fix and a named owner signs off.
This page is for developers who run the fixes, tech leads who own the review flow, and CTOs who have to answer “who checked this?” after an incident. The scenario is familiar: a security pass flags 23 issues on Monday, three engineers each start a Codex session to fix “the important ones”, and by Thursday nobody knows which findings were real, which fixes merged, and which session is still waiting for an approval. The workflow below gives every finding one path from report to merged, tested fix, and puts every session working on it on one screen.
What you get from a Codex security review flow
Section titled “What you get from a Codex security review flow”- A map of the four Codex security surfaces and what each one reports and can block.
- An
AGENTS.mdsecurity section that steers@codex security reviewand is in context for local reviews. - A least-privilege permission table: read-only for scanning, a worktree for fixes, never full access.
- A scripted sweep that writes findings as JSON and fails on confirmed critical or high issues.
- A triage rule that turns a finding into a failing test before any fix is written.
- A supervision routine for fix agents in the command center, including
codex queuefor steering them without opening each one.
Which Codex security surface does what?
Section titled “Which Codex security surface does what?”Codex has four places where security work happens. They differ in who starts them, what they see, and whether they can stop a merge.
| Surface | Where it runs | Started by | What it reports | Can it block a merge? |
|---|---|---|---|---|
| Codex Security | OpenAI’s application security agent, set up by following OpenAI’s Codex Security guide | Typically the application security team | Vulnerabilities to find, confirm and fix, in OpenAI’s own description | Not as described in OpenAI’s docs (checked 2026-08-28); route its findings into your triage queue |
@codex security review | GitHub pull requests | A pull request author or reviewer | Security-focused review comments on the pull request | Only if your merge rules require its comments resolved |
codex review / codex exec review (--base, --uncommitted, --commit) | Your machine or a runner | The author or a script | Review findings on a diff, uncommitted work, or a commit. In 0.157.1 those modes take no prompt, so security instructions can only reach them through AGENTS.md, which local sessions load as project instructions; confirm with a seeded finding that the review applies them. A custom prompt works only without a mode flag. | Yes, when a hook or CI step checks the result |
Agent command center (codex agents, /agents) | Your terminal, against the local app-server daemon | You | Status of every session: running, blocked, needs input | No. It is where you supervise the sessions that fix findings |
The order below matters more than which surface you pick. Cheap, private checks run first, so the pull request review sees fewer defects, and every finding from any surface lands in the same triage queue.
Route every security finding through one review flow
Section titled “Route every security finding through one review flow”-
Connect the repository. Enable Codex code review for the repository as described in connect Codex to GitHub. The same connection carries
@codex security review. -
Write the security rules once. Add security lines to the
## Review guidelinessection ofAGENTS.md(template below). Pull request reviews use that section. Local Codex sessions load AGENTS.md as project instructions, so the rules are in context forcodex reviewtoo; OpenAI documents the section for GitHub reviews only, so confirm that local reviews honour it with the seeded-vulnerability check described under “How do you know Codex Security is catching what matters?” below. -
Name a security owner per risky path. Put the owner in
CODEOWNERSforsrc/auth/,src/billing/, and anything that touches secrets. GitHub then requests that person’s review on any pull request touching those paths, so Codex’s findings there reach an owner, not only the author. -
Run a local sweep before the pull request exists. Use the scripted sweep in the next section so authors fix obvious issues before anyone else sees them.
-
Request
@codex security reviewon risky pull requests. Comment it on any pull request that touches authentication, input parsing, file access, or secrets. Keep on-demand requests until the rules are tuned, then consider automatic reviews as described in code review with Codex. -
Turn on Codex Security by following OpenAI’s Codex Security guide once the flow above works, and send its findings to the same triage queue. Confirm plan eligibility and the setup screens there; this page does not state them.
Here is a security section for AGENTS.md that a reviewer can check from the diff alone. Replace the paths with yours; keep the shape.
## Review guidelines
### Security- P0: a route under src/api/ that reads an ID from the request and loads a record without calling requireOwnership() or requireRole(). Missing authorization.- P0: SQL built with template strings or + concatenation anywhere under src/db/. Only the query builder or parameterised queries are allowed.- P0: a secret, token or private key literal in any file. Secrets come from env only.- P1: a new dependency in package.json that is not on the internal allowlist in docs/security/allowed-deps.md. Check the name exists on the registry.- P1: user-controlled URLs passed to fetch() without the allowlist in src/net/egress.ts.- Do not flag: test fixtures under tests/fixtures/, which contain fake credentials on purpose.- When you report a finding, name the rule above that it violates.The dependency rule is there for a reason. The LLM04 Supply Chain entry of OWASP’s GenAI LLM Top 10 2026 (GenAI Security Project, published 2026-08-04, read 2026-09-26) warns that coding assistants invent plausible package names that attackers then register, a practice it calls “slopsquatting”.
Run a scripted security sweep with structured findings
Section titled “Run a scripted security sweep with structured findings”A review comment is hard to gate on. A JSON file is easy. codex exec takes an --output-schema file and writes its final message to the path you give -o, so the sweep can fail a step on confirmed high-severity findings.
Save the schema as security-findings.schema.json:
{ "type": "object", "additionalProperties": false, "required": ["findings"], "properties": { "findings": { "type": "array", "items": { "type": "object", "additionalProperties": false, "required": ["id", "severity", "status", "file", "line", "cwe", "summary", "evidence", "reproduction"], "properties": { "id": { "type": "string" }, "severity": { "type": "string", "enum": ["critical", "high", "medium", "low"] }, "status": { "type": "string", "enum": ["confirmed", "suspected"] }, "file": { "type": "string" }, "line": { "type": "integer" }, "cwe": { "type": "string" }, "summary": { "type": "string" }, "evidence": { "type": "string" }, "reproduction": { "type": "string" } } } } }}Then run the sweep from the repository root in your terminal. The approval flag goes before exec, because codex exec has no -a of its own in 0.157.1:
# Remove the previous run's output so a failed run cannot pass on stale datarm -f findings.json
codex -a never exec -c 'default_permissions=":read-only"' \ --output-schema security-findings.schema.json \ -o findings.json \ "Security sweep of the changes on this branch against origin/main. Apply the Security rules in AGENTS.md. Mark a finding confirmed only if you can point to the exact input that reaches the vulnerable line; otherwise mark it suspected. Put that input in reproduction. Do not edit files."
# Fail on confirmed critical or high findingsjq -e '[.findings[] | select(.status == "confirmed" and (.severity == "critical" or .severity == "high"))] | length == 0' findings.jsonThe rm -f line matters: if codex exec fails (auth, network, a timeout), no new findings.json is written, and without it jq would pass on the previous run’s file, a silent green. With the file gone, jq fails on the missing file instead.
The built-in :read-only permission profile means the sweep cannot change the code it is judging. -a never means it never stops to ask: a command that needs more access fails instead of waiting for a human who is not there. Permission profiles are OpenAI’s preferred control and are still in beta; if your team stays on the legacy flag, use --sandbox read-only instead, never both. OpenAI says the two systems “do not compose”. In our test with Codex 0.157.1 (2026-09-26), passing both did not error: codex exec ran with the --sandbox value and the permission profile was ignored (the startup header showed sandbox: workspace-write for --sandbox workspace-write plus :read-only), so a mixed setup is not doing what it looks like. Before you trust a sweep, check the sandbox: line in the header that codex exec prints at startup. For the same sweep in CI, follow Codex in CI/CD: never run it on an untrusted ref while secrets are in the environment.
Decide what each security session is allowed to touch
Section titled “Decide what each security session is allowed to touch”Scanning and fixing need different permissions. Give each job the least it needs, and let the worktree, not your main checkout, absorb the fix agent’s writes.
| Job | Permissions in Codex CLI 0.157.1 | Why |
|---|---|---|
| Sweep or review | -c 'default_permissions=":read-only"' (legacy: --sandbox read-only), -a never in scripts | The scanner must not change what it judges |
| Confirm a finding | -c 'default_permissions=":workspace"' (legacy: --sandbox workspace-write) in a --worktree, default on-request approvals | It writes one failing test, nothing else |
| Propose a fix | Same as confirm, or --approve-for-me instead of a profile to route approvals through automatic review in the workspace-write sandbox | Edits stay in an isolated checkout until a pull request |
| Audit an untrusted repository | Choose Open restricted at the “Trust this folder?” prompt | Config, hooks and exec policies from untrusted folders stay disabled |
| Anything | Never --dangerously-bypass-approvals-and-sandbox outside a disposable, network-restricted sandbox | Full access lets the agent edit any file and reach the network without asking |
Two prompts in Codex 0.157.1 are worth knowing before a security team starts. When a conversation collects several cybersecurity-risk flags, Codex warns that “extra safety checks are on” and responses slow down; it points to OpenAI’s Trusted Access for Cyber program for authorised security work. And the full-access dialog carries an extra warning that cyber models “carry a higher risk of dangerous actions”. Neither is a reason to loosen the sandbox; both are reasons to keep offensive testing in a disposable environment.
Turn every finding into a failing test before a fix
Section titled “Turn every finding into a failing test before a fix”A finding from Codex Security or a security review is a claim. It becomes a fact when a test reproduces it. The rule for your team: no fix without a failing test that demonstrates the exploit, and no merge until that test passes.
| Finding state | Next action | Who decides |
|---|---|---|
| Reported, severity critical or high | Start a confirm session within one working day | Security owner for the path |
| Reported, medium or low | Batch weekly; confirm or close as noise | Tech lead |
| Cannot reproduce after one confirm session | Close as “not reproduced”, record why | Security owner |
| Confirmed (failing test exists) | Start a fix session in a worktree | Author or on-call developer |
| Fix pull request green, including the exploit test | Merge after sign-off | Security owner, not the fix author |
| Noise twice for the same pattern | Add a “Do not flag” line to AGENTS.md | Tech lead |
Treat the timings as a starting point to adapt, not a standard. The part not to adapt is the separation: the person who signs off on the fix is not the person, or the agent session, that wrote it.
Supervise fix agents in the command center
Section titled “Supervise fix agents in the command center”Once three or four confirm and fix sessions run at once, switching terminals stops working. The agent command center lists every session on the shared local app-server daemon in one view. Open it from a terminal with codex agents, or from inside a session with /agents.
-
Start each fix in its own worktree and name it after the finding. Run
codex --worktreeand then/renamethe session tosec-142, or press New in the command center. The name is how you find it again, andcodex queueaccepts an exact session name. -
Watch the status column, not the transcripts. In 0.157.1 a session shows as Running, Ready, Blocked, or Needs input. Attention messages say what is wrong: “Waiting for approval.”, “Waiting for your response.” or “Task encountered an error.”
-
Use Filter, Search and Group to keep security sessions apart from feature work. The
sec-name prefix makes those sessions easy to find with a search. -
Steer without switching. Send an instruction to a running session from any terminal:
Terminal window codex queue --thread sec-142 --message "Also add the same ownership check to PATCH /api/invoices/:id" -
Archive when merged, delete rarely. Archive stops any running work in the task and its child agents, and their history can be restored from the resume picker, which is what an incident record needs. Delete permanently deletes that history and cannot be undone, so keep it for sessions that never touched a finding.
To supervise sessions on another machine, such as a build box that runs the long fixes, point the command center at that app server: codex agents --remote wss://HOST:PORT --remote-auth-token-env CODEX_REMOTE_TOKEN, with -C DIR to set the working directory for new tasks there. HOST, PORT and DIR are yours; the token stays in the named environment variable, never on the command line. Multi-agent workflows in Codex covers splitting larger work across sessions.
How do you know Codex Security is catching what matters?
Section titled “How do you know Codex Security is catching what matters?”A quiet security review proves nothing on its own. Check the reviewer before you trust it, and check every fix without reading every line.
- Seeded-vulnerability check. Once a quarter, open a throwaway branch with five planted defects that match your
AGENTS.mdrules: a missing ownership check, string-built SQL, a hard-coded token, an unlisted dependency, an unchecked outbound URL. Run the sweep,@codex security reviewand, once enabled, a Codex Security scan of that branch, and record how many each catches. The catch rate is your baseline; a drop after a rules change is a regression. - Exploit test per fix. The failing-then-passing test in
tests/security/is the proof the fix works, and it stays in the suite so the bug cannot return quietly. - Normal gates still run. Types, lint and the full test suite pass on the fix pull request, like any other change.
- Noise rate. Track findings closed as “not reproduced” per week. Rising noise means the rules are too broad; falling catches on the seeded branch means they are too narrow.
- Human sign-off. The security owner approves every merge on a critical or high finding. Codex’s review is an input to that decision, not the approval.
What goes wrong with Codex Security and the command center?
Section titled “What goes wrong with Codex Security and the command center?”The first scan buries the team. Dozens of medium findings arrive and nobody triages any of them. Recovery: triage only critical and high for two weeks, and turn every repeated false positive into a “Do not flag” line before reading the rest.
A fix agent widens its own scope. Asked to fix one route, it refactors the auth module. Recovery: keep “do not change files outside…” in every fix prompt, reject the pull request, and restart the session from the failing test.
Findings are confirmed by argument, not by test. The session writes a convincing explanation and no test. Recovery: send it back with codex queue --thread sec-142 --message "Write the failing test first; no explanation without it".
An untrusted repository runs its own hooks. Trusting a vendor’s repository to audit it enables its config, hooks and exec policies. Recovery: open untrusted code with Open restricted. A session that ran while the folder was trusted can keep that configuration, so start a new task after changing trust.
The command center shows a stale list. Codex reports “agent list is stale; relaunch to retry”. Recovery: quit and run codex agents again. If sessions are still missing, run codex doctor, which diagnoses the local installation, config, auth and runtime health.
codex queue or codex agents refuses to start. In 0.157.1 --no-daemon cannot be used with either: queuing “must discover the shared server”, and the agents overview “requires a shared server”. Recovery: drop --no-daemon. To reach a different server, replace it with --remote ADDR (the two cannot be combined either).
A daemon update interrupts running fixes. codex app-server daemon update “may interrupt running work”. Recovery: update when the command center shows no Running sessions, then resume the interrupted ones by name.
How do Claude Code and Cursor handle the same job?
Section titled “How do Claude Code and Cursor handle the same job?”This page stays on Codex. Claude Code has /security-review and the Claude Security plugin, covered in security audits with Claude Code, and Agent view as its equivalent of the command center, covered in Agent view in Claude Code. Cursor’s security review features are covered in security in Cursor. Tool-neutral gates such as secret scanners and SAST, which run whichever agent wrote the code, belong in security gates for agent-written code.