Performance optimization with Codex
Performance optimization with Codex is a four-stage loop: a read-only diagnosis that names the bottleneck, a benchmark harness committed before any fix, best-of-N cloud tasks (codex cloud exec --attempts, up to four) that try competing strategies, and a deterministic gate that accepts a change only when the team’s own benchmark and test suite both pass.
This guide is for developers who own a slow service and tech leads who decide when a performance change can merge. Your API’s p95 latency crossed two seconds, the dashboard is red, and the slow request touches three services, a database, and a cache. Your laptop database holds one thousandth of production’s rows, so a local profile points at the wrong code. Commands on this page were checked against Codex CLI 0.157.1 on 2026-09-26.
What this Codex performance workflow gives you
Section titled “What this Codex performance workflow gives you”- A read-only diagnosis prompt that returns ranked bottlenecks with file paths and evidence, and changes nothing.
- A benchmark harness and a committed baseline that Codex is not allowed to edit, so a claimed speed-up is a measured one.
- A best-of-N cloud task that tries up to four optimization strategies, and the commands to compare and apply one attempt.
- A verification sequence that proves the change without reading every line of the diff.
- A weekly regression check where a script decides pass or fail and Codex only explains a failure.
The stage before this one is operating Codex day to day. Prerequisite: a Codex cloud environment whose setup script can seed the database; see Codex cloud environments. For the same loop with other agents, see performance analysis with Claude Code and performance optimization with Cursor.
Why does Codex performance work start with a benchmark, not a fix?
Section titled “Why does Codex performance work start with a benchmark, not a fix?”An agent asked to “make it faster” reports a number, and that number comes from a run you did not see, on data you did not choose. Performance changes also fail quietly: a cache that serves stale data or an index that locks a table in production both pass the unit tests. So the order is fixed:
- Diagnose with a read-only sandbox, so nothing changes while you look.
- Commit a benchmark and a baseline before any optimization exists.
- Let Codex try strategies against that benchmark in isolated cloud containers.
- Re-run the benchmark yourself, outside the agent, and merge only on that result.
Each step below produces an artifact the next step checks: a report, a harness, a set of diffs, and a pass or fail.
Diagnose the bottleneck without changing code
Section titled “Diagnose the bottleneck without changing code”Start the diagnosis in the read-only sandbox, so Codex can read code and run read-only commands but cannot edit files. Diagnosis rewards deeper reasoning, and GPT-6 Astra (the Codex default) runs at low effort unless you raise it. Run in your terminal, at the repository root:
codex --sandbox read-only -c model_reasoning_effort="high"If a read-only database MCP server is already configured for Codex (codex mcp list shows it), Codex can run EXPLAIN on the real query plans instead of guessing them. Which servers to use, and how to keep them read-only, is covered in production performance optimization.
The last line matters more than it looks. A diagnosis that mixes measured plans with guesses sends you to optimize the guess. Treat any bottleneck marked “inferred” as a hypothesis for the benchmark to confirm.
Build a benchmark harness Codex cannot game
Section titled “Build a benchmark harness Codex cannot game”The harness is the oracle for everything that follows, so it lands in its own commit before any optimization. Two rules make it trustworthy:
- Production-scale data. Seed at least 100,000 rows for the tables the diagnosis named, with a realistic distribution (a few heavy users, many light ones). N+1 patterns and missing indexes are invisible on a thousand rows.
- Repeatable numbers. Warm-up requests, several runs, and the median of each run’s p95. A single run on a busy machine measures the machine.
Review this diff yourself; it is the one diff in the loop worth reading line by line, because every later decision trusts it. Then commit it. From here on, any optimization diff that touches perf/ or scripts/seed-perf.ts is rejected without discussion.
Run competing optimizations as best-of-N cloud tasks
Section titled “Run competing optimizations as best-of-N cloud tasks”codex cloud exec --attempts N runs N independent agents on the same prompt, each producing its own diff. The CLI accepts 1 to 4 and rejects any other value (checked on 0.157.1), and codex cloud as a whole is marked experimental. Each attempt is a separate agent run that draws on your plan’s 5-hour and weekly allowance, so N attempts use about N times the allowance of one; see Codex cost management.
Best-of-N pays off in performance work because the strategies trade off differently: a rewritten query changes one file, a cache adds infrastructure and an invalidation problem, and parallel fetches change error handling. Attempts given the same prompt tend to converge, so name the alternatives and make each attempt state which one it took.
Before you submit, the cloud environment must be able to run the harness: its setup script installs dependencies, starts the database, and runs scripts/seed-perf.ts. Codex cloud environments covers setup scripts, caching, and agent internet access. Push the harness commit first, because the container clones your branch from the remote.
Submit it from your terminal, inside the repository, with the prompt saved to a file:
git push origin perf/dashboardcodex cloud exec --env perf-test --branch perf/dashboard --attempts 3 - < prompts/dashboard-perf.mdperf-test is the label or ID of your cloud environment, and - reads the prompt from standard input. The command prints a task URL; its ID is the TASK_ID for the commands below.
Choose an attempt by evidence, not by the best number
Section titled “Choose an attempt by evidence, not by the best number”Compare attempts on metadata and numbers before you open any diff:
# Status, attempt count, and size of each task in this environmentcodex cloud list --env perf-test --json \ | jq '.tasks[] | {id, status, attempt_total, summary}'
# Save every attempt's diff for comparisonfor n in 1 2 3; do codex cloud diff TASK_ID --attempt "$n" > "attempt-$n.diff"; doneThe container is not your production hardware, so compare each attempt’s before and after inside the same container, not absolute milliseconds across attempts. Then weigh the gain against what the attempt adds:
| Question | Prefer the attempt that… | Reject when… |
|---|---|---|
| Did it hit the target? | Clears 500 ms with the before and after from the same container | Quotes no before number, or only one run |
| How large is the change? | Touches the fewest files for the gain | Touches perf/, test assertions, or snapshots |
| What does it add to operate? | Adds no new service | Adds a cache with no freshness test |
| Is the migration safe? | Uses a non-blocking index build | Builds an index that locks a hot table |
If two attempts each win on a different bottleneck, apply one, re-baseline, and run a second task for the next bottleneck. A hand-merged combination of two attempts is a new, unmeasured change.
Verify the optimization without reading every line
Section titled “Verify the optimization without reading every line”The agent’s quoted numbers tell you which attempt to try, not whether to merge it. Verification happens outside the agent, and a named person signs off.
-
Apply the attempt in a clean worktree. Run in your terminal, from the main checkout:
Terminal window git fetch origin perf/dashboardgit worktree add --detach ../verify-perf origin/perf/dashboard && cd ../verify-perfcodex cloud apply TASK_ID --attempt 2git diff --stat -- perf/ scripts/seed-perf.ts # must print nothing -
Re-run the gates yourself. Run the type check, the linter, and the full test suite, then seed and benchmark:
node perf/bench.mjsexits non-zero if the median p95 is more than 10% above the baseline. The number that counts is this one, not the one in the task summary. -
Check the query plan, not the code. If the change adds an index, run
EXPLAINon the new query against the seeded database and confirm that the plan uses it. A composite index in the wrong column order compiles, passes tests, and does nothing. -
Open the pull request with the evidence. Paste the before and after JSON and the plan into the description, and ask for
@codex reviewon the pull request. Route any migration or cache change to the service’s code owner; see code review with Codex. -
Prove it in production and keep a way back. Ship behind a feature flag or to a canary, compare p95 in your application’s traces for the same endpoint over the same window, and roll back when it does not improve. The developer who submitted the task owns the merge and the rollback call.
When the change passes, update perf/baseline.json in its own commit, so the next regression check measures against the new number.
Catch performance regressions every week
Section titled “Catch performance regressions every week”A weekly check keeps the gain. Let a script make the pass or fail decision and use Codex only to explain a failure, because an agent asked “did performance regress?” can talk itself into either answer.
Schedule node perf/bench.mjs in CI (for example, a weekly schedule trigger). When it exits non-zero, run a headless, read-only Codex step that writes its triage to a file:
status=0node perf/bench.mjs > perf/latest.json || status=$? # survives bash -e on CI runnersif [ "$status" -ne 0 ]; then codex exec --sandbox read-only --ephemeral -o perf/triage.md - < prompts/perf-triage.mdfiexit "$status"Check out full history (for example, fetch-depth: 0 in actions/checkout), otherwise the triage step’s git log sees only one commit. The script keeps the benchmark’s exit code, so the job stays red on a regression and scheduled-failure notifications still fire; the triage file is an attachment to that failure, not a verdict. -o (--output-last-message) writes the final message to a file your job can attach to an issue, and --ephemeral keeps the session off disk. Give the job a timeout, because codex exec has no spend cap (checked on 0.157.1). Authentication, secrets, and the workflow file are covered in CI/CD with Codex and running Codex non-interactively. To run the same check from the desktop app instead, see Codex automations.
A human reads the triage and decides the next step: revert, open an optimization task, or accept a deliberate trade-off and re-baseline.
When Codex performance work goes wrong
Section titled “When Codex performance work goes wrong”| Symptom | Cause | Recovery |
|---|---|---|
| The fix is fast locally and slow in production | Benchmark data far smaller than production | Seed at production scale in the cloud environment’s setup script; re-baseline before the next attempt |
| The benchmark improves but users see no change | The attempt edited the harness, the seed, or a test | Reject any diff touching perf/; keep the “benchmark oracle” rule in AGENTS.md and check git diff --stat in step 1 of verification |
| All attempts return the same approach | Same prompt, no named alternatives | Name the strategies and require each attempt to state its choice |
| Numbers swing between runs | Noisy machine, no warm-up, one run | Five runs, warm-up requests, median p95; compare before and after in one container |
| Users see stale data after a cache change | TTL without invalidation on writes | Require a freshness test in the prompt; roll back the flag, then add invalidation |
| The deploy locks a table | Index built with a blocking statement on a hot table | Use a non-blocking build (CREATE INDEX CONCURRENTLY on PostgreSQL) in its own migration; see database work with Codex |
| The triage names no suspect commits | Shallow CI checkout, so git log sees one commit | Fetch full history (fetch-depth: 0) in the scheduled job |
| Best-of-N drains the weekly allowance | Four attempts on a mechanical change | Use attempts only where strategies differ; one attempt for a known fix |
Where to go next with performance
Section titled “Where to go next with performance”- Codex cloud environments: setup scripts, network access, and best-of-N in depth.
- Load, stress, and benchmark testing: turn the harness into a load-test and CI gate for every service.
- Production performance optimization: diagnose p99 regressions, pool exhaustion, and autoscaling with read-only MCP servers.
- Database work with Codex: queries, indexes, and migrations the agent writes safely.
- Codex cost management: what best-of-N and scheduled runs cost against your allowance.
- Security audits with Codex: the next page in this stage.