What is harness engineering?

Harness engineering is the practice of turning each recurring agent mistake into a permanent control, such as an instruction, a tool, a test or a linter rule, so the mistake cannot recur. Mitchell Hashimoto named it in writing on 5 February 2026, and an OpenAI post put it in its title six days later. Unlike vibe coding, it adds checks that run even when nobody reads the diff.

Origin Mitchell Hashimoto wrote the earliest use of the exact phrase the research found, in “My AI Adoption Journey” (opens in a new tab) on 5 February 2026, and an OpenAI post by Ryan Lopopolo (opens in a new tab) put “harness engineering” in its title on 11 February 2026, after which the term spread.

Last updated
By
Published
In Polish
Inżynieria harnessu
Closest ladder level
Level 5, free guide
Also searched as
agent harness engineering, engineer the harness, Agent = Model + Harness

Then $19.99 a month or $99.99 a year. Cancel from your account page. 30-day refund on a first purchase.

On this page Why it mattersHarness engineering vs vibe codingPrompt, context and harness engineeringThe seven layers of a coding agent’s harnessHow each tool does itQuestionsSources

Why it matters.

When an agent makes the same mistake twice, the reflex is another sentence in the rules file. Harness engineering asks which layer can actually stop it: a sentence is advice the model may ignore, a hook is code that runs on every matching event, and a sandbox is a wall. This site’s rule is to put each control in the most deterministic layer that can express it, and to version, review and evaluate the harness like code. An Anthropic post notes that every component encodes an assumption about what the model cannot do, and such assumptions go stale as models improve.

Harness engineering vs vibe coding

Both reduce how much code a person reads; our reading is that vibe coding stops checking, while harness engineering moves the checking into code that fails loudly.

Harness engineering compared with vibe coding: what it names, who writes the code, who checks it, what stops the loop, where the name came from and where it breaks
QuestionVibe codingHarness engineering
What it names A way of working: accept the model’s output without reading it A discipline: engineer the environment around the agent so a mistake cannot recur
Who writes the code The model The agent; OpenAI describes one internal product built with “0 lines of manually-written code” (vendor-reported, one team)
Who checks it Nobody reads the diff; the human judges behaviour by eye Mechanical checks and agent reviewers first; OpenAI: “Humans may review pull requests, but aren’t required to.”
What stops the loop The human decides it works The checks pass; Hashimoto: give the agent “fast, high quality tools to automatically tell it when it is wrong”
Where the name came from Andrej Karpathy, post on X, 2 February 2025 Mitchell Hashimoto, 5 February 2026, who disclaims inventing it; an OpenAI post by Ryan Lopopolo spread it on 11 February 2026
Where it breaks Anything others maintain; Karpathy called it fine for “throwaway weekend projects” Harness assumptions go stale as models improve (Anthropic), and Dex Horthy argues it is not enough on its own
Horthy’s essay has a section headed “this has nothing to do with vibe coding” and says: “If you love vibe coding, please, go on vibing.” He adds that the rest of it is aimed at “folks solving hard problems in complex codebases”.

Source: Vibe coding vs agentic engineering, 2026-10-10.

Prompt, context and harness engineering

Grigorev’s July 2026 taxonomy gives each layer its own question (it adds loop and graph engineering, not shown here), and sources disagree on whether the harness contains context engineering or the reverse.

The question each of prompt, context and harness engineering answers, and who says so
LayerThe question it answersSaid by
Prompt engineering “what we say when we interact with the agent” Alexey Grigorev, 22 July 2026
Context engineering “what the agent knows before it starts” Alexey Grigorev, 22 July 2026
Harness engineering What constrains and checks the agent This site’s reading
Birgitta Böckeler (2 April 2026) puts the harness inside context engineering: “Engineering a user harness for a coding agent is a specific form of context engineering.” Wikipedia’s “Agent harness” article, read on 10 October 2026, puts it the other way round: the harness “designs the whole operational environment and contains the other two as parts.” Neither nesting is settled, and this page picks neither.

Source: Vibe coding vs agentic engineering, 2026-10-10.

The seven layers of a coding agent’s harness

This site’s harness guide splits the harness into seven layers, each preventing one class of failure and enforcing it differently.

The seven harness layers, what each holds and how firmly each enforces
LayerWhat it holdsEnforcement
1 · Context Project instructions (CLAUDE.md, AGENTS.md, Cursor Rules), memory Advice: the model may ignore it
2 · Tools MCP servers, CLIs, subagents Capability: what the agent can reach
3 · Permissions and sandbox Permission modes, allow/ask/deny rules, OS sandbox, network egress Hard: the client or the OS refuses
4 · Hooks Scripts that run at lifecycle events, such as before a tool call or at stop Deterministic: code runs every time
5 · Skills Procedures loaded on demand (SKILL.md plus scripts) Advice, loaded only when relevant
6 · Plugins Installable bundles of skills, hooks, subagents and MCP servers Distribution: one versioned source
7 · Environments Worktrees, containers, cloud environments, seeded data, port blocks Isolation: separate state per run
The codebase and the oracle, the tests and gates that decide “done”, sit next to the harness rather than inside it. This site’s Setup Pack covers the Harness station: the files an agent reads before it writes anything. The ladder guide places the harness among the six stations of Level 5.

Source: Harness guide, 2026-10-02.

How each tool does it.

Cursor, Claude Code and Codex each name this their own way. Each note carries the date it was checked.

Cursor

Rules, MCP, Subagents, Run modes, Hooks, Agent Skills, Plugins, Worktrees and Cloud Agents cover the layers; file paths and setting names are left out until they are re-verified.

Harness guide, 2026-08-28

Claude Code

The layers live in CLAUDE.md, permissions rules in .claude/settings.json, the hooks block (33 events in Claude Code 2.1.283), .claude/skills/, plugins and --worktree; managed settings outrank project settings.

Harness guide, 2026-10-02

Codex

The layers live in AGENTS.md, .codex/config.toml (loaded only for a trusted repository), permission profiles (beta) and 12 hook events that need persisted trust (Codex CLI 0.157.1); /debug-config shows which layer set each value.

Harness guide, 2026-10-02

Questions about harness engineering.

How is harness engineering different from vibe coding?

Vibe coding accepts the agent’s output without reading it; harness engineering turns each repeated mistake into a test, linter rule, hook or instruction. Our reading: vibe coding stops checking and harness engineering moves the checking into code. Dex Horthy’s essay of 22 July 2026 adds: “If you love vibe coding, please, go on vibing.” He aims the rest of it at people solving hard problems in complex codebases.

Who coined harness engineering?

No coiner is established. Mitchell Hashimoto’s post of 5 February 2026 is the earliest use of the exact phrase the research found, and he wrote: “I don’t need to invent any new terms here; if another one exists, I’ll jump on the bandwagon.” OpenAI’s post by Ryan Lopopolo, 11 February 2026, put it in its title; Birgitta Böckeler noted it “only mentions ‘harness’ once in the text”.

What is the difference between harness engineering and context engineering?

Context engineering decides what the agent sees; harness engineering decides what checks and constrains it. Sources disagree on nesting. Birgitta Böckeler calls the harness “a specific form of context engineering”, while Wikipedia’s “Agent harness” article puts context engineering inside the harness. Neither nesting is settled, so treat them as two questions that overlap.

Is harness engineering enough?

Not according to Dex Horthy of HumanLayer, who wrote on 22 July 2026 that “no amount of harness engineering or loopsmaxxing can solve what is fundamentally a model-training issue.” Anthropic adds that every harness component “encodes an assumption about what the model can’t do on its own”, and those assumptions can go stale as models improve.

Sources.

The primary sources outside this site that this page relies on.

  1. My AI Adoption Journey (opens in a new tab) Mitchell Hashimoto
  2. Harness engineering: leveraging Codex in an agent-first world (opens in a new tab) Ryan Lopopolo, OpenAI
  3. Harness Engineering - first thoughts (opens in a new tab) Birgitta Böckeler, martinfowler.com
  4. Harness engineering for coding agent users (opens in a new tab) Birgitta Böckeler, martinfowler.com
  5. Harness design for long-running application development (opens in a new tab) Prithvi Rajasekaran, Anthropic
  6. Why Software Factories Fail (or: harness engineering is not enough) (opens in a new tab) Dex Horthy, HumanLayer
  7. AI-Native Development: Specifications, Loop and Graph Engineering (opens in a new tab) Alexey Grigorev, Alexey On Data

Read the guides in the same words.

Open every guide with the 7-day free trial. Each term is defined once here, and the guides use it the same way.

Then $19.99 a month or $99.99 a year. Cancel from your account page. 30-day refund on a first purchase.