Skip to content

Compound Engineering: The Loop, the Plugin, the Evidence

Compound Engineering is Every’s method, and MIT-licensed plugin, for making each change leave a coding agent better prepared for the next one: brainstorm, plan, work, simplify, review, then compound, which writes the lesson to docs/solutions/ where later plans retrieve it. Version 3.29.0 ships 36 skills for Claude Code, Codex, Cursor and other agent hosts.

The same bug arrives for the third time this quarter. Somebody fixed it in March, the fix was correct, and the reasoning behind it lives in a Slack thread nobody can find and an agent transcript that no longer exists. The fix is not the slow part. The slow part is that neither your team nor your agent remembers it.

This page is for developers and tech leads who want that memory without paying for it in context on every task. It covers the plugin as it ships today, one feature run end to end, how you check the output without reading every line, and where the method sits on this site’s autonomy ladder. For how plugins install and update in general, start with plugins and marketplaces; for how this framework compares with its neighbours, see agentic development frameworks compared.

What you’ll walk away with from Compound Engineering

Section titled “What you’ll walk away with from Compound Engineering”
  • Verified install and update commands for Claude Code, Codex and Cursor, and the one ordering trap that keeps upgrades on the old version
  • One feature run through the six-step loop, with what each skill writes to disk and where a human decides
  • A way to accept the loop’s output from the plan, the review report and CI rather than from the diff
  • A mapping of Every’s stages onto the L0–L5 ladder, and where /lfg actually sits on it
  • The design property that decides whether accumulated learnings help or hurt your agent
  • Five copy-paste prompts that run the method’s key checks with or without the plugin

The README describes six steps. Five of them produce a reviewed change; the sixth produces a system that makes the next change cheaper.

StepSkillWhat it writesLifecycle stage
Brainstormce-brainstormA requirements-only plan, through interactive Q&APlan
Plance-planThe same plan file, enriched into an implementation-ready plan under docs/plans/Design
Workce-workCode and commits, with the host’s tests and checks run as it goesBuild
Simplifyce-simplify-codeA cleanup of the fresh diff for clarity and reuseTest
Reviewce-code-reviewA report-only, multi-persona review against the plan; applying fixes is explicitTest
Compoundce-compoundOne learning under docs/solutions/, plus vocabulary in CONCEPTS.md; an instruction-file pointer only with your consentMaintain

The return arrow is the method: ce-brainstorm and ce-plan read docs/solutions/ as grounding, so a constraint learned in March shows up in September’s plan. Every states the time split as 80% planning and review, 20% execution.

The compound step also has a bar. ce-compound writes a learning only when it holds reasoning that the final code, tests, types, comments or existing docs do not already carry, and when losing it would plausibly cause a repeat. Its own test is a counterfactual: if the document disappeared, would the next engineer still repeat the mistake? If not, the run writes nothing and says why. A routine fix produces no learning, and that is the skill working.

Install Compound Engineering in Claude Code, Codex or Cursor

Section titled “Install Compound Engineering in Claude Code, Codex or Cursor”

All three hosts read the same repository. The marketplace is named compound-engineering-plugin and the plugin compound-engineering.

Inside a session:

/plugin marketplace add EveryInc/compound-engineering-plugin
/plugin install compound-engineering

From a shell or a script:

Terminal window
claude plugin marketplace add EveryInc/compound-engineering-plugin
claude plugin install compound-engineering@compound-engineering-plugin

To upgrade, refresh the marketplace first. /plugin update on its own reads a cached snapshot and leaves you on the old version:

/plugin marketplace update compound-engineering-plugin
/plugin update compound-engineering

Claude Code lists plugin skills with the plugin prefix, so the README’s /ce-plan appears as /compound-engineering:ce-plan.

Then run ce-setup once per repository (/compound-engineering:ce-setup in Claude Code, $ce-setup in Codex). It reports which optional tools are present, offers to create .compound-engineering/config.yaml, checks that a local override file is gitignored, and offers to add a pointer to the knowledge store in your agent instruction file. If your docs/ folder is already tracked content, set docs_root in that config to move every artifact folder under one repo-relative root.

What it costs in context. claude plugin details compound-engineering projected about 2,989 always-on tokens for 3.29.0 on Claude Code 2.1.283: the listing of 36 skills that every session carries. The core ce-* skills cost roughly 2.5k tokens each when they fire, and ce-debug about 5.6k. Run claude plugin details before and after installing any second discipline plugin; two overlapping loops cost twice and fight over which one fires.

Scaling beyond one repository. Tech leads who want the same rules in every repository can declare Compound Packs (experimental): folders of prescriptive rules, local or ref-pinned from git, that planning pulls in and review enforces with a citation back to the rule file. Treat a pack like shared code: an owner, a review, a version pin.

The loop is identical in all three hosts; only the prefix differs (/compound-engineering: in Claude Code, $ in Codex, the form Cursor’s slash menu shows). The example uses the README’s own starting request.

  1. Brainstorm the requirement. Run ce-brainstorm make background job retries safer. Answer its questions. You should see a requirements-only plan file under docs/plans/.

  2. Plan, then review the plan yourself. Run ce-plan. It enriches the same file into an implementation plan, grounded in any matching docs/solutions/ learnings. This is the human gate that matters most: read the plan, not the code that does not exist yet. Use the plan-review prompt below.

  3. Work. Run ce-work. It implements the plan, runs the project’s tests and type checks, and commits. Run it on a branch or in a git worktree, never on main.

  4. Simplify and review. Run ce-simplify-code, then ce-code-review. The review is report-only: findings are ranked by a confidence level and nothing is applied until you ask.

  5. Compound. After the change is verified, run ce-compound. Either it writes one learning under docs/solutions/, or it reports that nothing met the bar.

  6. Maintain the store. Once a quarter, or when plans start citing stale advice, run ce-compound-refresh. It gives every learning one verdict: Keep, Update, Consolidate, Replace or Delete. The ordinary run only judges accuracy and never deletes an accurate document; a pass that deletes documents the code already explains runs only when you ask for it and confirm.

How do you verify the loop’s output without reading every line?

Section titled “How do you verify the loop’s output without reading every line?”

The loop gives you four pieces of evidence, and together they replace reading the diff for most changes.

  • The plan as the spec. Acceptance criteria you approved in step 2 are the standard the rest of the run is judged against. ce-code-review reviews against that plan, not against taste.
  • The review report. Findings carry a discrete confidence level rather than a score. Agreement between reviewer personas raises confidence only when they ran in separate contexts (see the rules below).
  • Tests and CI. ce-work runs the host’s checks as it goes, ce-test-browser exercises the UI, and /lfg watches CI to a decision with a bounded repair loop.
  • Residuals in writing. Under /lfg, every finding it did not fix is recorded in the pull request body or the final report, so what the run left undone is explicit.

A human still signs off on the merge, and still reads code for the change classes where a wrong line is expensive: auth, money, schema and data migrations. The evidence-not-diffs page defines that list, and reviewing an agent’s pull request gives the triage order.

Where do Every’s stages sit on the autonomy ladder?

Section titled “Where do Every’s stages sit on the autonomy ladder?”

Every’s guide describes a stage ladder, 0 to 5, for how one developer’s way of working changes. This site uses one ladder for maturity, L0–L5 per loop, so here is the mapping. This is a reading aid for Every’s material, not an equivalence; one map treats Every’s stages as a separate model. Use Every’s stages to read Every’s material; use the ladder to say where a loop is.

Every’s stage (guide, read 2026-08-24)What the developer doesLadder levelWhat decides it
0Writes every lineL0, by handNo agent code reaches disk
1Asks a chat model and pastes what is usefulL1, assistedThe human drives and reads line by line
2Uses an agent with file access and approves every actionL2, pairedThe human steers live
3Agrees a detailed plan, steps away, receives a pull requestL3, review managerThe agent runs unattended; a human reads every diff
4Describes the outcome; the agent researches, plans, builds, self-reviews and opens the pull requestL3 or L4L4 only when the merge rests on the approved plan, tests and review evidence instead of reading the diff
5Directs several agents in parallel from anywhereNo level of its ownParallelism raises throughput; each loop still sits at L3 or L4. L5 needs merges no human read, which no stage describes

Where /lfg sits. /lfg automates Every’s stage 4: after ce-brainstorm, it plans, works, simplifies, reviews and applies eligible fixes, captures any durable learning, runs browser tests, commits, pushes, opens a pull request and watches CI. It does not merge unless you grant that for the run, and with no git remote it stops at local commits. On the ladder it is the machinery for an L4 loop, not proof of one. If you then read the whole 40-file diff before merging, the loop is L3 with extra automation. It never reaches L5 by default, because the merge stays with a human.

Does accumulating learnings make agents worse?

Section titled “Does accumulating learnings make agents worse?”

It can. The most direct study of repository context files, Gloaguen et al. at ETH Zurich (Evaluating AGENTS.md, arXiv, 2026), reports that such files do not generally improve task success while raising inference cost. Pruning CLAUDE.md and AGENTS.md carries the study’s figures and the vendors’ mid-2026 shift towards deleting context rather than adding it. A method that says “write down every lesson” has to answer that.

The plugin’s answer is one property: retrieved, not always loaded.

  • docs/solutions/ is a retrieved store. Learnings sit on disk and ce-plan pulls only the ones that match the work. Four hundred learnings add nothing to the always-loaded context; a task pays only for the few it retrieves.
  • CLAUDE.md and AGENTS.md are always loaded. Every line is paid for on every task. ce-compound edits an instruction file only to add a pointer to the store, only in interactive mode, and only after you consent. It never creates one.

Two mechanisms keep the store from becoming the same problem in another folder: the durable bar at capture, and ce-compound-refresh afterwards.

Which rules from the plugin’s glossary should you copy?

Section titled “Which rules from the plugin’s glossary should you copy?”

The repository’s CONCEPTS.md is a glossary the maintainers keep for their own agents. Four of its rules generalise well beyond this plugin.

Two reviewers in one context are not two witnesses. Independence is a property of the execution context a reviewer ran in, not of the lens it applied. If you prompt one agent to review “as security, then as performance, then as an architect”, the three findings share one reading of the diff and one set of mistakes. Only separately dispatched contexts earn a confidence promotion, and when dispatch fails the run reports the coverage it lost.

A model’s identity is a receipt, not a request. When a review goes to another model provider for a second opinion, the plugin records the serving backend’s report of which model ran, beside the model it asked for. Outputs without that receipt are labelled requested-but-unverified. A cross-model pipeline that trusts the request parameter has not instrumented silent fallbacks.

Hosts truncate skill bodies from the end, silently. Every known host limit keeps the beginning of a skill body and discards the rest, and none reports an error. So order is load-bearing: the plugin puts the outcome, the done condition and the stop rules at the top, and loads each phase’s mechanics from a reference file at the moment of acting. Do the same in your own skills and long instruction files.

Guidance at the moment of action beats a learning. An agent loads its skill instructions or root instruction file when it acts, so a learning that contradicts them is not merely stale; it is overridden in practice. Resolve the contradiction where the agent reads, and treat a conflict there as more urgent than an outdated path.

What breaks when you run Compound Engineering, and how do you recover?

Section titled “What breaks when you run Compound Engineering, and how do you recover?”

The upgrade did nothing. You ran /plugin update without refreshing the marketplace. Recover: run the refresh command for your host from the install tabs, update again, and check the result with claude plugin details compound-engineering in Claude Code or codex plugin list -m compound-engineering-plugin in Codex.

Learnings are written and nothing reads them. The store exists but the instruction file never points at it, so plans are not grounded in it. Recover: run the reachability prompt above, then accept ce-setup’s offer to add the knowledge-store pointer.

The store contradicts itself. Two documents about the same subsystem disagree and the agent finds the stale one first. Recover: run ce-compound-refresh on that area, and schedule it rather than waiting for the next contradiction.

A learning contradicts the instruction file. The instruction file wins at the moment of acting, whichever is right. Recover: decide which side current code follows, then fix the losing file. The refresh skill reports a wrong guidance file; it does not edit it for you.

Plan review disappears under deadline. The 80/20 split inverts quietly and you are back to reviewing large diffs from plans nobody read, which is the review backlog described in software factories. Recover: make the plan-review prompt a required step before ce-work, and split plans that come back SPLIT.

/lfg runs with broad permissions on your laptop. /lfg is built to proceed without waiting and stops only for irreversible actions you did not grant. Recover: run it in a worktree or a disposable, network-restricted sandbox with a git token scoped to one repository. Never pair it with --dangerously-skip-permissions outside such a sandbox; see sandboxes for coding agents.

Compounding is asserted, not measured. No published measurement shows the loop’s payoff. Recover: measure it on your repository. Count how often per quarter someone re-explains the same convention or re-diagnoses the same failure, and check whether the count falls after the store exists. If it does not, the loop is running but not compounding.

It is the wrong tool. For throwaway prototypes, repositories without a test runner, or teams where nobody will review what ce-compound writes, the ceremony produces no signal. Use a lighter discipline pack or none; the frameworks comparison lists the alternatives.

Where to go next with Compound Engineering

Section titled “Where to go next with Compound Engineering”