Skip to content

Choose a primary AI engineering harness

A primary AI engineering harness is the coding agent plus the repository configuration that takes a task from written intent to a verified diff: versioned instructions, a plan gate, bounded permissions, isolation, and a reproducible check command. Claude Code, Codex, and Cursor can all reach the maximum Scorecard Q1 answer; a trial on your own tasks decides which one fits.

You have used two agents for a month. One felt faster, the other asked fewer questions, and a teammate swears by a third. Nobody wrote down which one produced last week’s reverted pull request, and the repository still has no instruction file. This page is for the developer who picks the tool and the tech lead who wants the choice to rest on evidence the team can repeat.

Scorecard Q1 · Plan: How does your primary AI coding setup handle repository work?

Maximum-score answer (3 points): “Primary harness uses versioned instructions, plan and artifact gates, bounded tools, and reproducible verification evidence.”

The lower answers are chat suggestions copied by hand (1 point) and an agent that edits the repository while workflow and verification stay ad hoc (2 points).

  • A definition of “harness” you can check in a repository, not a product preference.
  • A one-to-two-week trial on three of your own tasks, run the same way in each candidate tool.
  • A scoring sheet that ranks candidates by accepted outcome first and cost last.
  • Four copy-paste prompts: readiness audit, plan-only run, bounded execution, and trial comparison.
  • A scope check and a CI gate that verify each run without anyone reading every line.
  • A decision record the tech lead signs, with a date to re-run the trial.

What counts as a harness, and why is the tool not the score?

Section titled “What counts as a harness, and why is the tool not the score?”

Q1 does not score a logo, a model name, or a terminal against an editor. It scores whether the configured setup can do six things, each of which leaves evidence in the repository or in the run log.

CapabilityEvidence you can point atHow you check it
Durable instructionsA reviewed CLAUDE.md, AGENTS.md, or Cursor Rule that names the check commandThe file is in Git and was changed through a pull request
Plan gateA non-editing run produces plan.md with files, risks, and testsThe plan commit exists before the first code commit
Bounded executionPermission mode, sandbox, or approval policy matches the task’s riskThe run log shows the mode; secrets and production credentials are absent
IsolationParallel or risky work runs in its own Git worktree or cloud environmentgit worktree list or the cloud run ID
Reproducible verificationOne command, such as npm run typecheck && npm run lint && npm test, run by the agent and again by CIThe CI result on the pull request
Traceable reviewThe diff maps to intent.md, spec.md, and plan.mdThe scope check below passes

The artifact chain has templates for intent.md, spec.md, and plan.md. The tool map names the feature that implements each row in each tool, checked on Claude Code 2.1.283 and Codex 0.157.1. This page does not repeat it; it tells you how to choose.

Why not choose from a feature matrix or a demo?

Section titled “Why not choose from a feature matrix or a demo?”

Two reasons. First, every row above is available in all three tools, so a feature matrix rarely separates them for your repository. Second, impressions mislead. In METR’s randomised study of experienced open-source developers (published 2025-07-10, early-2025 tooling), developers expected AI to speed them up by 24% and, after the measured slowdown, still believed it had sped them up by 20%. The same study measured them taking 19% longer. Measure on your own tasks.

The tools differ in where you work, not in what they can prove:

If your team mostly…Start the trial withWhy
Works in an IDE and reviews visuallyCursorIDE-first: Agent, Plan Mode, Worktrees, Cloud Agents, and Bugbot on pull requests
Scripts, runs headless in CI, or enforces policy with hooksClaude CodeCLI-first: permission modes, hooks, claude -p with JSON output, subagents
Moves between desktop app, CLI, IDE, and cloud tasksCodexMulti-surface: codex exec, permission profiles, --worktree, codex review, cloud tasks

Treat the table as the order to test in, not the verdict. Always include the tool the team uses today as the baseline.

Plan one to two weeks. The trial compares harnesses, so each candidate gets the same tasks, the same instructions, and the same gate. Model choice is a separate decision: run each tool on its default model, record which one it used, and leave routing to model routing by evidence. Current defaults and prices live on the model comparison guide.

  1. Pick three representative tasks. Each one touches several files, has an observable outcome, and needs no production access. A good set is one bug fix with a failing test, one small feature behind a flag, and one refactor with no behaviour change. Keep a base commit for each.

  2. Write acceptance before any tool runs. For each task, commit intent.md with the owner and what is out of scope, plus a grader: tests the agent does not see, and the check command. The grader decides “accepted”, not the person running the trial.

  3. Make the repository agent-ready once. Commit one instruction file with the exact check command, build steps, and forbidden paths. Keep the content in AGENTS.md and add a one-line CLAUDE.md containing @AGENTS.md, so Claude Code on either release channel reads it; put the same check command in a Cursor Rule. Agent-ready codebase and concise repository context cover the content.

  4. Configure each candidate with the same boundary. Workspace writes only, no production credentials in the environment, network off unless the task needs it. Permissions and sandboxing compares the controls row by row.

  5. Run plan-only first, then execute in a worktree. Commit the accepted plan.md before any code. Execute each task once per candidate, each in its own worktree, with the prompts below. The commands for each tool follow these steps.

  6. Verify with the scope check and CI. Push each run as a draft pull request. CI runs the check command and the hidden tests; the scope check flags any file the plan did not name.

  7. Record the sheet and decide. Fill in the trial record, apply the decision rule, and open a pull request that adds the decision to docs/ai/primary-harness.md.

The plan and execute commands differ by tool. Each block runs in a terminal at the repository root; run the headless command from inside the worktree the plan step created.

Terminal window
# Plan only: plan mode reads and proposes, it does not edit
claude --permission-mode plan -w trial-task-1
# Headless alternative for the execute step: run it inside the plan worktree
cd "$(git worktree list | awk '/trial-task-1/ {print $1}')"
# Cost ceiling and JSON result
claude -p "$(cat trial/task-1/execute-prompt.md)" \
--permission-mode acceptEdits \
--allowedTools "Bash(npm test *)" "Bash(npm run *)" \
--max-budget-usd 5 \
--output-format json > trial/runs/task-1-claude.json

With Claude Code 2.1.283 on the latest channel, interactive terminal sessions start in auto mode, while claude -p starts in Manual, so record which mode each run used. Check the release channel too: the stable channel (2.1.274 on 2026-09-26) starts on an older default model.

Prompts stay the same across tools, so differences in the result come from the harness. Paste them into the agent’s prompt.

How do you verify a trial run without reading every line?

Section titled “How do you verify a trial run without reading every line?”

The grader and CI decide acceptance. A human reads the evidence, not the diff:

  • CI is the authority. The agent runs the check command during the session, but the result that counts is CI’s. That holds only while the agent cannot change the pipeline: on pull_request, GitHub runs the workflow file from the pull request itself, so any diff under .github/ is out of scope for a trial task, and /.github/ belongs under CODEOWNERS with a required code-owner review. Hidden tests live outside the agent’s worktree until CI runs them.
  • The scope check catches silent sprawl. Run it on each draft pull request:
Terminal window
# Terminal, in the run's worktree: flag files the plan did not name
git diff --name-only main...HEAD | while read -r f; do
case "$f" in plan.md|*/plan.md|intent.md|*/intent.md) continue;; esac
grep -qF "$f" plan.md || echo "OUT OF PLAN: $f"
done
  • The evidence report is checked against the plan. Every deviation the agent listed must be either accepted by the task owner or reverted.
  • A reviewer agent reads the diff first. Run the tool’s own review (/code-review in Claude Code, codex review --base main in Codex, Bugbot in Cursor) and a human reads only the findings it marks important. Layered pull-request review sets that up.
  • The task owner signs acceptance. The named owner in intent.md accepts the outcome; the tech lead signs the harness decision.

Record every run on one sheet. The values below are placeholders that show the shape.

trial/record.yaml
date: 2026-09-26
tasks: [task-1-bugfix, task-2-flagged-feature, task-3-refactor]
boundary: workspace writes, no network, no production credentials
candidates:
- harness: claude-code 2.1.283 (latest), default model, CLAUDE.md -> @AGENTS.md
accepted_by_ci_and_hidden_tests: 3/3
out_of_plan_files: 0
human_corrections_per_task: 1
review_minutes_per_task: 12
cost_per_accepted_task_usd: 2.10
- harness: codex 0.157.1, default model, AGENTS.md, profile :workspace
accepted_by_ci_and_hidden_tests: 2/3
out_of_plan_files: 1
human_corrections_per_task: 2
review_minutes_per_task: 18
cost_per_accepted_task_usd: 1.60
decision: primary harness = <tool>; re-run trial on 2026-12-26 or on a major release
signed_off_by: tech lead, 2026-09-26

Decision rule. Rank by accepted tasks first. Among candidates tied on acceptance, prefer fewer out-of-plan files and corrections, then less review time, then lower cost per accepted task. A cheaper run that needs a rerun is not cheaper. Plan and seat sizing is the next scorecard question: size an AI plan from measured workload.

What goes wrong when you pick a primary harness?

Section titled “What goes wrong when you pick a primary harness?”

The team scores a product category. A debate about terminals against editors replaces measurement. Recovery: restate the six capabilities, run the same three tasks in each candidate, and let the sheet decide.

The agent edits before it understands the repository. Code appears before any plan, and the diff touches files nobody expected. Recovery: start every task in plan mode or Plan Mode, commit plan.md, and fail the run when the scope check prints OUT OF PLAN.

The instruction file is not read. Codex ignores the project’s AGENTS.md in an untrusted project (since 0.150.0), Claude Code reads AGENTS.md only when there is no CLAUDE.md and only from v2.1.277 for first-party sessions (v2.1.281 on Bedrock, Google Cloud and Foundry), on the latest channel only, and how Cursor attaches Rules could not be re-verified for this page on 2026-09-26. Recovery: ask each agent, before the first task, “Which instruction files did you load, and what is the check command?” and fix the setup until all three answer the same.

A green test hides a wrong outcome. The agent adjusted a test to match its code. Recovery: keep acceptance tests outside the agent’s worktree until CI runs them, and trace each test to a line in intent.md.

One candidate ran with wider permissions. It finished more tasks because it could reach more. Recovery: discard its runs, reset to the shared boundary, and rerun. Report “cannot finish inside the boundary” as a finding.

The decision never expires. A harness chosen in May is still primary in September after three major releases. Recovery: date the decision record, re-run the three tasks on each major release of the primary tool or every quarter, and keep the instruction file in AGENTS.md so switching costs one import line.

  • One instruction file names the exact check command and is changed only through pull requests.
  • Non-trivial work starts from intent.md and an accepted plan.md committed before code.
  • Runs use a stated permission boundary with no production credentials in the environment.
  • Parallel or risky work runs in its own worktree or cloud environment.
  • CI re-runs the check command and the hidden tests on every agent pull request.
  • A dated decision record names the primary harness, the trial evidence, and the re-run date.

Where to go next from your primary harness

Section titled “Where to go next from your primary harness”

Then open the adapter for the harness you chose: Claude Code, Codex, or Cursor.