What is harness engineering?
Harness engineering is the practice of turning each recurring agent mistake into a permanent control, such as an instruction, a tool, a test or a linter rule, so the mistake cannot recur. Mitchell Hashimoto named it in writing on 5 February 2026, and an OpenAI post put it in its title six days later. Unlike vibe coding, it adds checks that run even when nobody reads the diff.
Origin Mitchell Hashimoto wrote the earliest use of the exact phrase the research found, in “My AI Adoption Journey” (opens in a new tab) on 5 February 2026, and an OpenAI post by Ryan Lopopolo (opens in a new tab) put “harness engineering” in its title on 11 February 2026, after which the term spread.
- In Polish
- Inżynieria harnessu
- Closest ladder level
- Level 5, free guide
- Also searched as
- agent harness engineering, engineer the harness, Agent = Model + Harness
Then $19.99 a month or $99.99 a year. Cancel from your account page. 30-day refund on a first purchase.
Why it matters.
When an agent makes the same mistake twice, the reflex is another sentence in the rules file. Harness engineering asks which layer can actually stop it: a sentence is advice the model may ignore, a hook is code that runs on every matching event, and a sandbox is a wall. This site’s rule is to put each control in the most deterministic layer that can express it, and to version, review and evaluate the harness like code. An Anthropic post notes that every component encodes an assumption about what the model cannot do, and such assumptions go stale as models improve.
Harness engineering vs vibe coding
Both reduce how much code a person reads; our reading is that vibe coding stops checking, while harness engineering moves the checking into code that fails loudly.
| Question | Vibe coding | Harness engineering |
|---|---|---|
| What it names | A way of working: accept the model’s output without reading it | A discipline: engineer the environment around the agent so a mistake cannot recur |
| Who writes the code | The model | The agent; OpenAI describes one internal product built with “0 lines of manually-written code” (vendor-reported, one team) |
| Who checks it | Nobody reads the diff; the human judges behaviour by eye | Mechanical checks and agent reviewers first; OpenAI: “Humans may review pull requests, but aren’t required to.” |
| What stops the loop | The human decides it works | The checks pass; Hashimoto: give the agent “fast, high quality tools to automatically tell it when it is wrong” |
| Where the name came from | Andrej Karpathy, post on X, 2 February 2025 | Mitchell Hashimoto, 5 February 2026, who disclaims inventing it; an OpenAI post by Ryan Lopopolo spread it on 11 February 2026 |
| Where it breaks | Anything others maintain; Karpathy called it fine for “throwaway weekend projects” | Harness assumptions go stale as models improve (Anthropic), and Dex Horthy argues it is not enough on its own |
Prompt, context and harness engineering
Grigorev’s July 2026 taxonomy gives each layer its own question (it adds loop and graph engineering, not shown here), and sources disagree on whether the harness contains context engineering or the reverse.
| Layer | The question it answers | Said by |
|---|---|---|
| Prompt engineering | “what we say when we interact with the agent” | Alexey Grigorev, 22 July 2026 |
| Context engineering | “what the agent knows before it starts” | Alexey Grigorev, 22 July 2026 |
| Harness engineering | What constrains and checks the agent | This site’s reading |
The seven layers of a coding agent’s harness
This site’s harness guide splits the harness into seven layers, each preventing one class of failure and enforcing it differently.
| Layer | What it holds | Enforcement |
|---|---|---|
| 1 · Context | Project instructions (CLAUDE.md, AGENTS.md, Cursor Rules), memory | Advice: the model may ignore it |
| 2 · Tools | MCP servers, CLIs, subagents | Capability: what the agent can reach |
| 3 · Permissions and sandbox | Permission modes, allow/ask/deny rules, OS sandbox, network egress | Hard: the client or the OS refuses |
| 4 · Hooks | Scripts that run at lifecycle events, such as before a tool call or at stop | Deterministic: code runs every time |
| 5 · Skills | Procedures loaded on demand (SKILL.md plus scripts) | Advice, loaded only when relevant |
| 6 · Plugins | Installable bundles of skills, hooks, subagents and MCP servers | Distribution: one versioned source |
| 7 · Environments | Worktrees, containers, cloud environments, seeded data, port blocks | Isolation: separate state per run |
How each tool does it.
Cursor, Claude Code and Codex each name this their own way. Each note carries the date it was checked.
Cursor
Rules, MCP, Subagents, Run modes, Hooks, Agent Skills, Plugins, Worktrees and Cloud Agents cover the layers; file paths and setting names are left out until they are re-verified.
Harness guide, 2026-08-28
Claude Code
The layers live in CLAUDE.md, permissions rules in .claude/settings.json, the hooks block (33 events in Claude Code 2.1.283), .claude/skills/, plugins and --worktree; managed settings outrank project settings.
Harness guide, 2026-10-02
Codex
The layers live in AGENTS.md, .codex/config.toml (loaded only for a trusted repository), permission profiles (beta) and 12 hook events that need persisted trust (Codex CLI 0.157.1); /debug-config shows which layer set each value.
Harness guide, 2026-10-02
Questions about harness engineering.
How is harness engineering different from vibe coding?
Vibe coding accepts the agent’s output without reading it; harness engineering turns each repeated mistake into a test, linter rule, hook or instruction. Our reading: vibe coding stops checking and harness engineering moves the checking into code. Dex Horthy’s essay of 22 July 2026 adds: “If you love vibe coding, please, go on vibing.” He aims the rest of it at people solving hard problems in complex codebases.
Who coined harness engineering?
No coiner is established. Mitchell Hashimoto’s post of 5 February 2026 is the earliest use of the exact phrase the research found, and he wrote: “I don’t need to invent any new terms here; if another one exists, I’ll jump on the bandwagon.” OpenAI’s post by Ryan Lopopolo, 11 February 2026, put it in its title; Birgitta Böckeler noted it “only mentions ‘harness’ once in the text”.
What is the difference between harness engineering and context engineering?
Context engineering decides what the agent sees; harness engineering decides what checks and constrains it. Sources disagree on nesting. Birgitta Böckeler calls the harness “a specific form of context engineering”, while Wikipedia’s “Agent harness” article puts context engineering inside the harness. Neither nesting is settled, so treat them as two questions that overlap.
Is harness engineering enough?
Not according to Dex Horthy of HumanLayer, who wrote on 22 July 2026 that “no amount of harness engineering or loopsmaxxing can solve what is fundamentally a model-training issue.” Anthropic adds that every harness component “encodes an assumption about what the model can’t do on its own”, and those assumptions can go stale as models improve.
The vocabulary of AI-driven development
Every name below is defined against vibe coding: who writes the code, who checks it, and what stops the loop.
Read the long-form guide: agentic engineering vs vibe coding
Sources.
The primary sources outside this site that this page relies on.
- My AI Adoption Journey (opens in a new tab) Mitchell Hashimoto
- Harness engineering: leveraging Codex in an agent-first world (opens in a new tab) Ryan Lopopolo, OpenAI
- Harness Engineering - first thoughts (opens in a new tab) Birgitta Böckeler, martinfowler.com
- Harness engineering for coding agent users (opens in a new tab) Birgitta Böckeler, martinfowler.com
- Harness design for long-running application development (opens in a new tab) Prithvi Rajasekaran, Anthropic
- Why Software Factories Fail (or: harness engineering is not enough) (opens in a new tab) Dex Horthy, HumanLayer
- AI-Native Development: Specifications, Loop and Graph Engineering (opens in a new tab) Alexey Grigorev, Alexey On Data
Keep reading.
The guides that go deeper, and the terms and comparisons next to this one.
In the docs
- The full A-Z glossarySubscription
- Vibe coding vs agentic engineering: 21 terms comparedFree
- Level 5: You Run the Software FactoryFree
- The harness: everything the agent runs insideSubscription
- Permissions, sandboxes and approval modes across Claude Code, Codex and CursorSubscription
- Making a codebase agent-readySubscription
- One map: the autonomy ladder, the lifecycle and the factory stationsFree
Related terms
Comparisons
Read the guides in the same words.
Open every guide with the 7-day free trial. Each term is defined once here, and the guides use it the same way.
Then $19.99 a month or $99.99 a year. Cancel from your account page. 30-day refund on a first purchase.