Skip to content

Performance optimization with Codex

Performance optimization with Codex is a four-stage loop: a read-only diagnosis that names the bottleneck, a benchmark harness committed before any fix, best-of-N cloud tasks (codex cloud exec --attempts, up to four) that try competing strategies, and a deterministic gate that accepts a change only when the team’s own benchmark and test suite both pass.

This guide is for developers who own a slow service and tech leads who decide when a performance change can merge. Your API’s p95 latency crossed two seconds, the dashboard is red, and the slow request touches three services, a database, and a cache. Your laptop database holds one thousandth of production’s rows, so a local profile points at the wrong code. Commands on this page were checked against Codex CLI 0.157.1 on 2026-09-26.

What this Codex performance workflow gives you

Section titled “What this Codex performance workflow gives you”
  • A read-only diagnosis prompt that returns ranked bottlenecks with file paths and evidence, and changes nothing.
  • A benchmark harness and a committed baseline that Codex is not allowed to edit, so a claimed speed-up is a measured one.
  • A best-of-N cloud task that tries up to four optimization strategies, and the commands to compare and apply one attempt.
  • A verification sequence that proves the change without reading every line of the diff.
  • A weekly regression check where a script decides pass or fail and Codex only explains a failure.

The stage before this one is operating Codex day to day. Prerequisite: a Codex cloud environment whose setup script can seed the database; see Codex cloud environments. For the same loop with other agents, see performance analysis with Claude Code and performance optimization with Cursor.

Why does Codex performance work start with a benchmark, not a fix?

Section titled “Why does Codex performance work start with a benchmark, not a fix?”

An agent asked to “make it faster” reports a number, and that number comes from a run you did not see, on data you did not choose. Performance changes also fail quietly: a cache that serves stale data or an index that locks a table in production both pass the unit tests. So the order is fixed:

  1. Diagnose with a read-only sandbox, so nothing changes while you look.
  2. Commit a benchmark and a baseline before any optimization exists.
  3. Let Codex try strategies against that benchmark in isolated cloud containers.
  4. Re-run the benchmark yourself, outside the agent, and merge only on that result.

Each step below produces an artifact the next step checks: a report, a harness, a set of diffs, and a pass or fail.

Diagnose the bottleneck without changing code

Section titled “Diagnose the bottleneck without changing code”

Start the diagnosis in the read-only sandbox, so Codex can read code and run read-only commands but cannot edit files. Diagnosis rewards deeper reasoning, and GPT-6 Astra (the Codex default) runs at low effort unless you raise it. Run in your terminal, at the repository root:

Terminal window
codex --sandbox read-only -c model_reasoning_effort="high"

If a read-only database MCP server is already configured for Codex (codex mcp list shows it), Codex can run EXPLAIN on the real query plans instead of guessing them. Which servers to use, and how to keep them read-only, is covered in production performance optimization.

The last line matters more than it looks. A diagnosis that mixes measured plans with guesses sends you to optimize the guess. Treat any bottleneck marked “inferred” as a hypothesis for the benchmark to confirm.

Build a benchmark harness Codex cannot game

Section titled “Build a benchmark harness Codex cannot game”

The harness is the oracle for everything that follows, so it lands in its own commit before any optimization. Two rules make it trustworthy:

  • Production-scale data. Seed at least 100,000 rows for the tables the diagnosis named, with a realistic distribution (a few heavy users, many light ones). N+1 patterns and missing indexes are invisible on a thousand rows.
  • Repeatable numbers. Warm-up requests, several runs, and the median of each run’s p95. A single run on a busy machine measures the machine.

Review this diff yourself; it is the one diff in the loop worth reading line by line, because every later decision trusts it. Then commit it. From here on, any optimization diff that touches perf/ or scripts/seed-perf.ts is rejected without discussion.

Run competing optimizations as best-of-N cloud tasks

Section titled “Run competing optimizations as best-of-N cloud tasks”

codex cloud exec --attempts N runs N independent agents on the same prompt, each producing its own diff. The CLI accepts 1 to 4 and rejects any other value (checked on 0.157.1), and codex cloud as a whole is marked experimental. Each attempt is a separate agent run that draws on your plan’s 5-hour and weekly allowance, so N attempts use about N times the allowance of one; see Codex cost management.

Best-of-N pays off in performance work because the strategies trade off differently: a rewritten query changes one file, a cache adds infrastructure and an invalidation problem, and parallel fetches change error handling. Attempts given the same prompt tend to converge, so name the alternatives and make each attempt state which one it took.

Before you submit, the cloud environment must be able to run the harness: its setup script installs dependencies, starts the database, and runs scripts/seed-perf.ts. Codex cloud environments covers setup scripts, caching, and agent internet access. Push the harness commit first, because the container clones your branch from the remote.

Submit it from your terminal, inside the repository, with the prompt saved to a file:

Terminal window
git push origin perf/dashboard
codex cloud exec --env perf-test --branch perf/dashboard --attempts 3 - < prompts/dashboard-perf.md

perf-test is the label or ID of your cloud environment, and - reads the prompt from standard input. The command prints a task URL; its ID is the TASK_ID for the commands below.

Choose an attempt by evidence, not by the best number

Section titled “Choose an attempt by evidence, not by the best number”

Compare attempts on metadata and numbers before you open any diff:

Terminal window
# Status, attempt count, and size of each task in this environment
codex cloud list --env perf-test --json \
| jq '.tasks[] | {id, status, attempt_total, summary}'
# Save every attempt's diff for comparison
for n in 1 2 3; do codex cloud diff TASK_ID --attempt "$n" > "attempt-$n.diff"; done

The container is not your production hardware, so compare each attempt’s before and after inside the same container, not absolute milliseconds across attempts. Then weigh the gain against what the attempt adds:

QuestionPrefer the attempt that…Reject when…
Did it hit the target?Clears 500 ms with the before and after from the same containerQuotes no before number, or only one run
How large is the change?Touches the fewest files for the gainTouches perf/, test assertions, or snapshots
What does it add to operate?Adds no new serviceAdds a cache with no freshness test
Is the migration safe?Uses a non-blocking index buildBuilds an index that locks a hot table

If two attempts each win on a different bottleneck, apply one, re-baseline, and run a second task for the next bottleneck. A hand-merged combination of two attempts is a new, unmeasured change.

Verify the optimization without reading every line

Section titled “Verify the optimization without reading every line”

The agent’s quoted numbers tell you which attempt to try, not whether to merge it. Verification happens outside the agent, and a named person signs off.

  1. Apply the attempt in a clean worktree. Run in your terminal, from the main checkout:

    Terminal window
    git fetch origin perf/dashboard
    git worktree add --detach ../verify-perf origin/perf/dashboard && cd ../verify-perf
    codex cloud apply TASK_ID --attempt 2
    git diff --stat -- perf/ scripts/seed-perf.ts # must print nothing
  2. Re-run the gates yourself. Run the type check, the linter, and the full test suite, then seed and benchmark: node perf/bench.mjs exits non-zero if the median p95 is more than 10% above the baseline. The number that counts is this one, not the one in the task summary.

  3. Check the query plan, not the code. If the change adds an index, run EXPLAIN on the new query against the seeded database and confirm that the plan uses it. A composite index in the wrong column order compiles, passes tests, and does nothing.

  4. Open the pull request with the evidence. Paste the before and after JSON and the plan into the description, and ask for @codex review on the pull request. Route any migration or cache change to the service’s code owner; see code review with Codex.

  5. Prove it in production and keep a way back. Ship behind a feature flag or to a canary, compare p95 in your application’s traces for the same endpoint over the same window, and roll back when it does not improve. The developer who submitted the task owns the merge and the rollback call.

When the change passes, update perf/baseline.json in its own commit, so the next regression check measures against the new number.

A weekly check keeps the gain. Let a script make the pass or fail decision and use Codex only to explain a failure, because an agent asked “did performance regress?” can talk itself into either answer.

Schedule node perf/bench.mjs in CI (for example, a weekly schedule trigger). When it exits non-zero, run a headless, read-only Codex step that writes its triage to a file:

Terminal window
status=0
node perf/bench.mjs > perf/latest.json || status=$? # survives bash -e on CI runners
if [ "$status" -ne 0 ]; then
codex exec --sandbox read-only --ephemeral -o perf/triage.md - < prompts/perf-triage.md
fi
exit "$status"

Check out full history (for example, fetch-depth: 0 in actions/checkout), otherwise the triage step’s git log sees only one commit. The script keeps the benchmark’s exit code, so the job stays red on a regression and scheduled-failure notifications still fire; the triage file is an attachment to that failure, not a verdict. -o (--output-last-message) writes the final message to a file your job can attach to an issue, and --ephemeral keeps the session off disk. Give the job a timeout, because codex exec has no spend cap (checked on 0.157.1). Authentication, secrets, and the workflow file are covered in CI/CD with Codex and running Codex non-interactively. To run the same check from the desktop app instead, see Codex automations.

A human reads the triage and decides the next step: revert, open an optimization task, or accept a deliberate trade-off and re-baseline.

SymptomCauseRecovery
The fix is fast locally and slow in productionBenchmark data far smaller than productionSeed at production scale in the cloud environment’s setup script; re-baseline before the next attempt
The benchmark improves but users see no changeThe attempt edited the harness, the seed, or a testReject any diff touching perf/; keep the “benchmark oracle” rule in AGENTS.md and check git diff --stat in step 1 of verification
All attempts return the same approachSame prompt, no named alternativesName the strategies and require each attempt to state its choice
Numbers swing between runsNoisy machine, no warm-up, one runFive runs, warm-up requests, median p95; compare before and after in one container
Users see stale data after a cache changeTTL without invalidation on writesRequire a freshness test in the prompt; roll back the flag, then add invalidation
The deploy locks a tableIndex built with a blocking statement on a hot tableUse a non-blocking build (CREATE INDEX CONCURRENTLY on PostgreSQL) in its own migration; see database work with Codex
The triage names no suspect commitsShallow CI checkout, so git log sees one commitFetch full history (fetch-depth: 0) in the scheduled job
Best-of-N drains the weekly allowanceFour attempts on a mechanical changeUse attempts only where strategies differ; one attempt for a known fix