Skip to content

Executable acceptance criteria: from story to failing test

Executable acceptance criteria turn each line of a user story into a Given/When/Then scenario with concrete values and an automated test that fails before any implementation exists. The tests are approved by a human, then locked outside the implementing agent’s edit scope, so a green run proves the agreed behavior instead of the agent’s own interpretation of it.

You hand a coding agent a story that says “users can apply discount codes; invalid codes show an error”. Forty minutes later the pull request is green: 14 new tests, all passing. None of them checks the cart total, the expired-code case returns HTTP 200 with a message, and one test was quietly changed from toBe(9000) to toBe(10000) on the third iteration. The agent did what it was asked. Nobody wrote down what “works” means in a form a machine could hold it to.

This page is for developers who hand stories to Claude Code, Codex, or Cursor, and for tech leads who want review to move from the diff to a contract that the team approves before the code exists.

What executable acceptance criteria give you

Section titled “What executable acceptance criteria give you”
  • A rewrite pattern that turns a vague story into three to seven numbered Given/When/Then criteria with concrete values.
  • Failing acceptance tests in three stacks (TypeScript with Playwright, Python with pytest-bdd, Elixir with ExUnit), each shown before and after.
  • A lock on the tests: a Claude Code deny rule for the implementing session, plus a CI guard and CODEOWNERS entry that work for every tool.
  • Four copy-paste prompts: interrogate the story, write the red tests, implement against the locked contract, and audit the tests for weakness.
  • A sign-off model in which a human approves the contract and the evidence instead of reading every changed line.

Why do agents need executable acceptance criteria?

Section titled “Why do agents need executable acceptance criteria?”

An agent optimises for the check it can run. If the only check is “the tests pass” and the agent writes the tests, the loop has no fixed point: the test and the code move together, and green means “consistent with itself”. A prose criterion in a ticket does not help, because the agent reads it once and then interprets it.

The fix is an order of operations. The criteria become tests before implementation, a human approves them while they are red, and the implementing session cannot change them. This is acceptance test-driven development with one addition that matters for agents: the test author and the implementer are separate sessions, and the contract is enforced by tooling rather than by instructions.

This page covers the per-story contract. Keeping a longer-lived spec.md authoritative across many stories is spec-driven development; writing and running Gherkin with Cucumber.js is on behavior-driven development workflows. Where the criteria sit in the intent → spec → plan → tasks sequence is on the artifact chain.

What makes an acceptance criterion executable?

Section titled “What makes an acceptance criterion executable?”

A criterion is executable when a test can fail on it. Four properties get it there:

PropertyVague (before)Executable (after)
Concrete values“a valid code reduces the total”“a cart totalling 100.00 with code SAVE10 (10%) totals 90.00”
Observable at a boundary“the discount is stored”“POST /api/carts/:id/discount returns 200 with totalCents: 9000”
The unhappy path is named“invalid codes show an error”“an expired code returns 422 code_expired and the total stays 100.00”
A stable IDnoneAC-2, used in the test name and checked by CI

Here is the discount story after the rewrite. It lives in the repository at specs/discount-codes.md, next to the code it governs:

# specs/discount-codes.md — approved by: product owner, 2026-09-26
Feature: Discount codes at checkout
# AC-1
Scenario: A valid code takes its percentage off the cart total
Given a cart totalling 100.00
And an active code "SAVE10" worth 10%
When the customer applies "SAVE10"
Then the response is 200
And the cart total is 90.00
# AC-2
Scenario: An expired code is rejected and the total is unchanged
Given a cart totalling 100.00
And a code "SPRING10" that expired on 2026-03-31
When the customer applies "SPRING10"
Then the response is 422 with error "code_expired"
And the cart total is still 100.00
# AC-3
Scenario: A customer can redeem a code only once
Given customer "c-42" has already redeemed "WELCOME15"
And customer "c-42" has a cart totalling 80.00
When they apply "WELCOME15"
Then the response is 409 with error "code_already_used"
And the cart total is still 80.00

Three to seven criteria per story is a useful range. Fewer usually means the unhappy paths are missing; more usually means the story should be split (see shaping a backlog for agents). Edge cases that are rules over many inputs, such as “a total is never negative”, belong in property-based tests rather than in a long list of scenarios.

The workflow has two sessions and one human gate between them. The human approves a small, readable artifact (scenarios plus a red run), not an implementation diff.

  1. Interrogate the story. Give the agent the ticket (paste it, or pull it through the tracker’s MCP server) and ask for numbered Given/When/Then criteria plus a list of open questions. Answer the questions yourself or with the product owner; an agent that guesses an answer here encodes the guess in the contract.

  2. Commit the criteria. Save them to specs/<feature>.md with the IDs. The criteria file is the thing the product owner reads.

  3. Write the failing tests in a separate session. A fresh session, with no implementation in context, writes one test per criterion in tests/acceptance/, names each test with its ID, and drives the system through its public boundary (HTTP, CLI, or UI), never through internal functions.

  4. Prove the tests are red for the right reason. Run the acceptance suite and read the failure messages. “Expected 9000, received 404” is a correct red: the route does not exist yet. A syntax error, an import error, or a missing fixture is a broken test, and it would also go green for the wrong reason later.

  5. Approve the contract. Open a contract pull request with the criteria file, the tests marked as expected failures, and the red output pasted into the description. The marker is what lets a red-by-design pull request pass required CI; see how a red contract merges below. A human with ownership of the behavior (product owner or tech lead) approves it. This review is short because the artifact is small.

  6. Lock the contract. The implementing session gets a deny rule or an instruction for tests/acceptance/ and specs/, and CI fails any implementation pull request whose change to those paths is anything other than deleting the expected-failure markers. The section on locking below has the files.

  7. Implement until green. Before the session starts, the contract owner (a human, not the agent) opens the implementation branch with one commit that only deletes the expected-failure markers from step 5, so the tests are honestly red. A new session then implements against the locked tests and may add as many unit tests of its own as it wants. It stops when the acceptance suite passes, or when it believes a criterion is wrong, in which case it reports instead of editing.

  8. Verify the evidence. CI checks three things: the acceptance suite is green, the contract paths are unchanged apart from the deleted markers, and every criterion ID has a test. The reviewer reads that result, not every line of the diff.

What does a failing acceptance test look like in your stack?

Section titled “What does a failing acceptance test look like in your stack?”

The same story, in three stacks. Each “before” is the kind of test an agent writes when it tests its own implementation after the fact; each “after” is a criterion-level test written before any implementation.

Before. The agent mocked its own lookup function and asserted the mock was called. The test passes whether or not the total changes:

// src/lib/discounts.test.ts: written by the implementer, after the code
vi.mock('./discount-repo');
it('applies discount', async () => {
vi.mocked(findDiscount).mockResolvedValue({ percent: 10 });
await applyDiscount(cart, 'SAVE10');
expect(findDiscount).toHaveBeenCalledWith('SAVE10');
});

After. A Playwright API test (@playwright/test, using the request fixture) at the HTTP boundary, written first:

tests/acceptance/discount-codes.spec.ts
import { test, expect } from '@playwright/test';
import { seedCart, seedCode } from '../support/seed';
test('AC-1: a valid code takes its percentage off the cart total', async ({ request }) => {
const cart = await seedCart(request, { totalCents: 10_000 });
await seedCode(request, { code: 'SAVE10', percent: 10 });
const res = await request.post(`/api/carts/${cart.id}/discount`, { data: { code: 'SAVE10' } });
expect(res.status()).toBe(200);
expect((await res.json()).totalCents).toBe(9_000);
});
test('AC-2: an expired code is rejected and the total is unchanged', async ({ request }) => {
const cart = await seedCart(request, { totalCents: 10_000 });
await seedCode(request, { code: 'SPRING10', percent: 10, expiresAt: '2026-03-31T23:59:59Z' });
const res = await request.post(`/api/carts/${cart.id}/discount`, { data: { code: 'SPRING10' } });
expect(res.status()).toBe(422);
expect((await res.json()).error).toBe('code_expired');
const after = await request.get(`/api/carts/${cart.id}`);
expect((await after.json()).totalCents).toBe(10_000);
});

This assumes use.baseURL (and a webServer block) in playwright.config.ts; without it, the relative URLs fail with an invalid-URL error, which is red for the wrong reason.

Red output before implementation: Expected: 200, Received: 404. That is the correct red.

The “after” tests share three habits: they name the criterion, they seed state through fixtures instead of mocking the application, and they assert the outcome the customer sees. A test that only calls internal functions can pass while the endpoint returns 500.

How do you keep acceptance tests outside the agent’s edit scope?

Section titled “How do you keep acceptance tests outside the agent’s edit scope?”

Instructions alone do not hold. An agent that has iterated for a while on a failing test will eventually consider “the test is wrong” and edit it. Use two layers: a lock in the implementing session where the tool supports one, and a CI guard that works regardless of tool. The full treatment, including holdout scenarios kept out of the repository, is on protecting the oracle.

Add the CI guard and code owners (all tools)

Section titled “Add the CI guard and code owners (all tools)”

The CI guard is the part that makes the contract binding. Unless a pull request is labeled as a contract change, it fails on any change to the contract paths other than deleting the expected-failure markers, and CODEOWNERS routes every pull request that touches those paths to a human owner:

.github/workflows/acceptance.yml
name: acceptance
on:
pull_request:
types: [opened, synchronize, reopened, labeled, unlabeled]
push:
branches: [main]
permissions:
contents: read
jobs:
contract:
if: github.event_name == 'pull_request'
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
with:
fetch-depth: 0
persist-credentials: false
- name: Implementation PRs may only remove expected-failure markers
if: ${{ !contains(github.event.pull_request.labels.*.name, 'acceptance-contract') }}
env:
BASE_REF: ${{ github.base_ref }}
run: |
base=$(git merge-base "origin/$BASE_REF" HEAD)
changes=$(git diff --name-status "$base" HEAD -- specs tests/acceptance)
bad=0
while IFS=$'\t' read -r status path; do
[ -n "$status" ] || continue
if [ "$status" = M ] && [[ "$path" == tests/acceptance/* ]] &&
git show "$base:$path" |
sed -E -e '/^[[:space:]]*@xfail[[:space:]]*$/d' -e 's/@xfail[[:space:]]+//g' -e 's/\btest\.fail\(/test(/g' |
cmp -s - <(git show "HEAD:$path"); then
echo "$path: expected-failure markers removed, nothing else"
else
echo "::error::$status $path changes the contract (or leaves a marker in place)"; bad=1
fi
done <<< "$changes"
if [ "$bad" -ne 0 ]; then
echo "::error::Open a separate PR labeled acceptance-contract for the owner to approve."
exit 1
fi
- name: Every criterion has a test
run: |
for id in $(grep -ohE '([A-Z]+-)?AC-[0-9]+' specs/*.md | sort -u); do
grep -rqE "(^|[^A-Z-])$id([^0-9]|$)" tests/acceptance || { echo "::error::$id has no acceptance test"; exit 1; }
done
open-contracts:
if: github.event_name == 'push'
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
- name: No expected-failure markers left on the default branch
run: |
if grep -rnE '\btest\.fail\(|@xfail\b' tests/acceptance; then
echo "::error::A contract is merged but not implemented; its markers are listed above."
exit 1
fi
# .github/CODEOWNERS
/specs/ @your-org/product-owners
/tests/acceptance/ @your-org/product-owners
/pytest.ini @your-org/product-owners
/.github/ @your-org/product-owners

Turn on Require review from Code Owners in the branch protection rule or ruleset for your default branch; without it, CODEOWNERS only requests a review. Replace main with your default branch. Run the acceptance suite itself in the same workflow or in your existing test job. The criterion IDs above are unique per repository; if your IDs restart per feature, prefix them (DISC-AC-2). The check matches each ID whole, so AC-2 is not satisfied by a DISC-AC-2 test, and the reverse also holds.

The labeled and unlabeled event types re-run the guard when the owner adds the label, so nobody has to push an empty commit. The label only routes the pull request: anyone who can add labels, including an agent holding a gh token, can make the guard skip. The real control is Require review from Code Owners on the contract paths. The same holds for the guard itself: on pull_request, GitHub runs the workflow file from the pull request’s own branch, so an implementation pull request can edit or delete this guard and the run will pass. That is why /.github/ is in CODEOWNERS: with required Code Owner review, a changed workflow cannot merge without the owner. The workflow needs no secrets, holds a read-only token, and checks out with persist-credentials: false.

How does a red contract merge past required CI?

Section titled “How does a red contract merge past required CI?”

A contract pull request is red by design, so if the acceptance suite is a required check, it cannot merge. Do not fix that by making the acceptance job optional on the default branch: an optional check is also what lets an implementation pull request merge red. Two patterns keep the check required.

Strict expected-failure markers (Playwright, pytest). The contract pull request declares each new test as expected to fail. The required job then passes only while those tests really fail, so CI re-proves the red run on every push. When the implementation makes a test pass, the runner reports the unexpected pass as a failure:

// In the contract PR: same title, so the ID check still finds it
test.fail('AC-1: a valid code takes its percentage off the cart total', async ({ request }) => {
// ...body unchanged...
});
RunnerMarker in the contract PRResult once the feature works
Playwrighttest.fail('AC-1: ...', async () => ...)Expected to fail, but passed., job fails
pytest-bdd@xfail tag on the scenario, plus xfail_strict = true under [pytest] in pytest.iniXPASS(strict), job fails

Deleting the markers is a human step, and it comes first. While a marker is present, a correct implementation turns the suite red (Expected to fail, but passed.), which would push an agent told to reach green to un-implement the feature. So before the implementing session starts, the contract owner opens the implementation branch with one commit that does nothing but delete that story’s markers (test.fail( back to test(, the @xfail tags). The session then works against honestly red tests and stops when the suite is green, as the prompts on this page say. That pull request needs no label: the guard compares each changed file under tests/acceptance/ with its base version minus the markers and fails on any other difference, including a marker left in place, so keep one story per test file. CODEOWNERS still routes it to the owner, who sees a test diff that is only removed markers. Both runners were checked on 2026-10-02 (Playwright 1.63.0, pytest 9.1.1 with pytest-bdd 9.0.0): green while red, failed once the stubbed feature returned the expected value. The guard script was run against marker-only, partial, weakened-assertion, added-tag, new-file, deleted-file and spec changes; only the marker-only and application-only changes passed.

Keep the strictness out of reach too. With xfail_strict = false in pytest.ini, a tagged scenario reports 1 xpassed once the feature works and 1 xfailed while it is broken, both with exit code 0, so the markers never have to come off. Run the pytest suite in CI as pytest -o xfail_strict=true tests/acceptance; the command-line override wins over the ini file (with the ini set to false, the working feature reports 1 failed and exit code 1). pytest.ini is in CODEOWNERS above for the same reason. The open-contracts job fails on the default branch while any marker remains there. It gates no pull request; it keeps a contract that is merged but not implemented visible as a red run until the implementation pull request deletes its markers.

ExUnit has no strict expected-failure mode, and a plain @tag :pending exclusion fails open (the implementation merges without the tests ever running), so Elixir projects use the next pattern.

A story branch. The contract pull request targets story/discount-codes instead of the default branch, and implementation pull requests target the story branch, so the guard compares against a base that already holds the contract. Give story/* its own ruleset with Require review from Code Owners, or the contract pull request merges without the owner. In that ruleset, require the guard (the contract job) but not the acceptance suite: the contract pull request into the story branch is red by design, and a required suite would block it exactly as it would on the default branch. The suite is required on the final pull request from the story branch into the default branch, which must be green and carries the acceptance-contract label because it brings the contract paths with it. No markers are involved, so the implementing session starts against red tests directly.

Lock the implementing session in each tool

Section titled “Lock the implementing session in each tool”

Claude Code checks Edit(path) rules for every built-in file-writing tool, for file commands it recognizes in Bash such as sed and tee, and for redirect targets. Deny rules from any settings scope win over allow rules. A deny rule in the shared .claude/settings.json would also block the session that writes the tests, so pass the rules as flags on the implementation session only:

Terminal window
# Terminal: interactive implementation in its own worktree
claude --worktree disc-impl \
--disallowedTools "Edit(/specs/**)" "Edit(/tests/acceptance/**)"
# Terminal or CI: headless implementation with a budget
claude -p "Implement specs/discount-codes.md until npx playwright test tests/acceptance passes." \
--disallowedTools "Edit(/specs/**)" "Edit(/tests/acceptance/**)" \
--permission-mode acceptEdits \
--allowedTools "Bash(npx playwright test *)" \
--max-budget-usd 5

The leading / anchors the pattern at the settings source. For CLI flags and for .claude/settings.json or .claude/settings.local.json, that is the primary working directory (the worktree, here). For a file passed with --settings <file>, it is that file’s own directory, so the same rules in .claude/implementer.json would deny .claude/specs/** and protect nothing; if you prefer a file, use //-absolute paths in it. Anthropic’s permissions page warns that these rules do not cover “arbitrary subprocesses that read or write files indirectly, like a Python or Node script”; for OS-level enforcement use the Bash sandbox, and keep the CI guard as the binding check. Flags checked against Claude Code 2.1.283 on 2026-09-26.

If the story lives in Linear or Jira, connect the tracker so step 1 reads the ticket directly: claude mcp add --transport http linear https://mcp.linear.app/mcp in Claude Code or codex mcp add linear --url https://mcp.linear.app/mcp && codex mcp login linear in Codex (codex mcp login linear runs the OAuth sign-in that Linear’s server requires), and a remote entry under mcpServers in .cursor/mcp.json for Cursor. For UI criteria, the Playwright MCP server (@playwright/mcp) lets the test-writing session explore the running page before it writes selectors; setup is on browser automation MCP servers.

How do you know the contract is being kept?

Section titled “How do you know the contract is being kept?”

The point of the contract is that nobody has to read every changed line to trust the result. Each check has an owner and a signal:

What is checkedHowWho signs off
The criteria describe the right behaviorContract pull request: criteria file, test names, red outputProduct owner or tech lead, before implementation
The tests are red for the right reasonRed run pasted into the contract pull requestSame reviewer
The tests are strong enoughThe audit prompt above, then mutation score on oracle strengthTech lead, per suite rather than per pull request
The implementation meets the contractAcceptance suite green in CICI
The contract was not changed to fit the codeCI guard on specs/ and tests/acceptance/ (unlabeled, only deleted markers pass), CODEOWNERS on those paths, pytest.ini and .github/CI, plus the code owner for marker removal and any contract change
Every criterion is coveredID traceability stepCI

The implementation pull request then carries a short evidence summary: the AC IDs and their results, the guard result, and anything the agent flagged. That summary is one field of the evidence bundle, and reading it instead of the diff is the practice described on reading evidence instead of code. The reviewer still reads the code where the risk class demands it (security, payments, migrations); the contract moves the default, not every case.

What breaks with executable acceptance criteria?

Section titled “What breaks with executable acceptance criteria?”

The agent edits or skips a test. Signal: the CI guard fails, or a .skip/@pytest.mark.skip appears in the acceptance folder. Recovery: reject the pull request, restart the implementation in a fresh session with the lock loaded, and add “do not skip” to the implementation prompt. If it happens twice on the same criterion, the criterion is probably wrong; take it back to the contract pull request.

The agent special-cases the test data. The code checks if (code === 'SAVE10'). Signal: the audit prompt, or a reviewer skimming for literals from the tests. Recovery: add a second example with different values to the scenario (a scenario outline, or a second test) through a contract pull request, and keep one set of examples out of the repository as a holdout that only CI runs; see protecting the oracle.

The tests were red for the wrong reason. A missing fixture or a typo made every test fail, and the implementer “fixed” the test support code until green. Recovery: re-run step 4 and read each failure message before approving; treat tests/support/ as part of the contract if the implementer should not change it.

Tests are coupled to the implementation. The tests import a service class or assert a SQL query, so every refactor breaks them and the agent is tempted to edit them. Recovery: rewrite them against the public boundary. If the behavior has no public boundary yet, the criterion is at the wrong level; move it to a unit test the implementer owns.

The contract churns mid-story. Product changes a rule while the implementation is running. Recovery: stop the run, open a contract pull request with the changed criterion and its new red test, approve it, then restart implementation. Never let the implementing session carry the change.

The acceptance suite gets slow or flaky. Signal: developers rerun CI until green. Recovery: keep acceptance tests at three to seven per story and at the API boundary where possible; move combinatorial cases to unit or property tests; fix flakes before adding scenarios, because a flaky oracle trains everyone, human and agent, to ignore red.

Frequently asked questions

What is an executable acceptance criterion?

A criterion written as a Given/When/Then scenario with concrete values, paired with an automated test that fails before the feature exists and passes only when the behavior is there. It carries an ID such as AC-2 so CI can prove every criterion has a test.

Why must the acceptance tests be outside the agent's edit scope?

An agent told to make tests pass can also make them pass by editing them. If the implementing session cannot change the acceptance tests, and CI fails any implementation pull request that changes them beyond deleting the expected-failure markers, a green run means the approved behavior exists.

What does the human approve if not the diff?

The contract: the scenarios, the test names mapped to criterion IDs, and the red run that shows each test failing for the right reason. After implementation the human checks the evidence (green acceptance run, unchanged contract, traceability) rather than reading every line.