Executable acceptance criteria: from story to failing test
Executable acceptance criteria turn each line of a user story into a Given/When/Then scenario with concrete values and an automated test that fails before any implementation exists. The tests are approved by a human, then locked outside the implementing agent’s edit scope, so a green run proves the agreed behavior instead of the agent’s own interpretation of it.
You hand a coding agent a story that says “users can apply discount codes; invalid codes show an error”. Forty minutes later the pull request is green: 14 new tests, all passing. None of them checks the cart total, the expired-code case returns HTTP 200 with a message, and one test was quietly changed from toBe(9000) to toBe(10000) on the third iteration. The agent did what it was asked. Nobody wrote down what “works” means in a form a machine could hold it to.
This page is for developers who hand stories to Claude Code, Codex, or Cursor, and for tech leads who want review to move from the diff to a contract that the team approves before the code exists.
What executable acceptance criteria give you
Section titled “What executable acceptance criteria give you”- A rewrite pattern that turns a vague story into three to seven numbered Given/When/Then criteria with concrete values.
- Failing acceptance tests in three stacks (TypeScript with Playwright, Python with pytest-bdd, Elixir with ExUnit), each shown before and after.
- A lock on the tests: a Claude Code deny rule for the implementing session, plus a CI guard and
CODEOWNERSentry that work for every tool. - Four copy-paste prompts: interrogate the story, write the red tests, implement against the locked contract, and audit the tests for weakness.
- A sign-off model in which a human approves the contract and the evidence instead of reading every changed line.
Why do agents need executable acceptance criteria?
Section titled “Why do agents need executable acceptance criteria?”An agent optimises for the check it can run. If the only check is “the tests pass” and the agent writes the tests, the loop has no fixed point: the test and the code move together, and green means “consistent with itself”. A prose criterion in a ticket does not help, because the agent reads it once and then interprets it.
The fix is an order of operations. The criteria become tests before implementation, a human approves them while they are red, and the implementing session cannot change them. This is acceptance test-driven development with one addition that matters for agents: the test author and the implementer are separate sessions, and the contract is enforced by tooling rather than by instructions.
This page covers the per-story contract. Keeping a longer-lived spec.md authoritative across many stories is spec-driven development; writing and running Gherkin with Cucumber.js is on behavior-driven development workflows. Where the criteria sit in the intent → spec → plan → tasks sequence is on the artifact chain.
What makes an acceptance criterion executable?
Section titled “What makes an acceptance criterion executable?”A criterion is executable when a test can fail on it. Four properties get it there:
| Property | Vague (before) | Executable (after) |
|---|---|---|
| Concrete values | “a valid code reduces the total” | “a cart totalling 100.00 with code SAVE10 (10%) totals 90.00” |
| Observable at a boundary | “the discount is stored” | “POST /api/carts/:id/discount returns 200 with totalCents: 9000” |
| The unhappy path is named | “invalid codes show an error” | “an expired code returns 422 code_expired and the total stays 100.00” |
| A stable ID | none | AC-2, used in the test name and checked by CI |
Here is the discount story after the rewrite. It lives in the repository at specs/discount-codes.md, next to the code it governs:
# specs/discount-codes.md — approved by: product owner, 2026-09-26Feature: Discount codes at checkout
# AC-1 Scenario: A valid code takes its percentage off the cart total Given a cart totalling 100.00 And an active code "SAVE10" worth 10% When the customer applies "SAVE10" Then the response is 200 And the cart total is 90.00
# AC-2 Scenario: An expired code is rejected and the total is unchanged Given a cart totalling 100.00 And a code "SPRING10" that expired on 2026-03-31 When the customer applies "SPRING10" Then the response is 422 with error "code_expired" And the cart total is still 100.00
# AC-3 Scenario: A customer can redeem a code only once Given customer "c-42" has already redeemed "WELCOME15" And customer "c-42" has a cart totalling 80.00 When they apply "WELCOME15" Then the response is 409 with error "code_already_used" And the cart total is still 80.00Three to seven criteria per story is a useful range. Fewer usually means the unhappy paths are missing; more usually means the story should be split (see shaping a backlog for agents). Edge cases that are rules over many inputs, such as “a total is never negative”, belong in property-based tests rather than in a long list of scenarios.
How do you go from story to failing test?
Section titled “How do you go from story to failing test?”The workflow has two sessions and one human gate between them. The human approves a small, readable artifact (scenarios plus a red run), not an implementation diff.
-
Interrogate the story. Give the agent the ticket (paste it, or pull it through the tracker’s MCP server) and ask for numbered Given/When/Then criteria plus a list of open questions. Answer the questions yourself or with the product owner; an agent that guesses an answer here encodes the guess in the contract.
-
Commit the criteria. Save them to
specs/<feature>.mdwith the IDs. The criteria file is the thing the product owner reads. -
Write the failing tests in a separate session. A fresh session, with no implementation in context, writes one test per criterion in
tests/acceptance/, names each test with its ID, and drives the system through its public boundary (HTTP, CLI, or UI), never through internal functions. -
Prove the tests are red for the right reason. Run the acceptance suite and read the failure messages. “Expected 9000, received 404” is a correct red: the route does not exist yet. A syntax error, an import error, or a missing fixture is a broken test, and it would also go green for the wrong reason later.
-
Approve the contract. Open a contract pull request with the criteria file, the tests marked as expected failures, and the red output pasted into the description. The marker is what lets a red-by-design pull request pass required CI; see how a red contract merges below. A human with ownership of the behavior (product owner or tech lead) approves it. This review is short because the artifact is small.
-
Lock the contract. The implementing session gets a deny rule or an instruction for
tests/acceptance/andspecs/, and CI fails any implementation pull request whose change to those paths is anything other than deleting the expected-failure markers. The section on locking below has the files. -
Implement until green. Before the session starts, the contract owner (a human, not the agent) opens the implementation branch with one commit that only deletes the expected-failure markers from step 5, so the tests are honestly red. A new session then implements against the locked tests and may add as many unit tests of its own as it wants. It stops when the acceptance suite passes, or when it believes a criterion is wrong, in which case it reports instead of editing.
-
Verify the evidence. CI checks three things: the acceptance suite is green, the contract paths are unchanged apart from the deleted markers, and every criterion ID has a test. The reviewer reads that result, not every line of the diff.
What does a failing acceptance test look like in your stack?
Section titled “What does a failing acceptance test look like in your stack?”The same story, in three stacks. Each “before” is the kind of test an agent writes when it tests its own implementation after the fact; each “after” is a criterion-level test written before any implementation.
Before. The agent mocked its own lookup function and asserted the mock was called. The test passes whether or not the total changes:
// src/lib/discounts.test.ts: written by the implementer, after the codevi.mock('./discount-repo');
it('applies discount', async () => { vi.mocked(findDiscount).mockResolvedValue({ percent: 10 }); await applyDiscount(cart, 'SAVE10'); expect(findDiscount).toHaveBeenCalledWith('SAVE10');});After. A Playwright API test (@playwright/test, using the request fixture) at the HTTP boundary, written first:
import { test, expect } from '@playwright/test';import { seedCart, seedCode } from '../support/seed';
test('AC-1: a valid code takes its percentage off the cart total', async ({ request }) => { const cart = await seedCart(request, { totalCents: 10_000 }); await seedCode(request, { code: 'SAVE10', percent: 10 });
const res = await request.post(`/api/carts/${cart.id}/discount`, { data: { code: 'SAVE10' } });
expect(res.status()).toBe(200); expect((await res.json()).totalCents).toBe(9_000);});
test('AC-2: an expired code is rejected and the total is unchanged', async ({ request }) => { const cart = await seedCart(request, { totalCents: 10_000 }); await seedCode(request, { code: 'SPRING10', percent: 10, expiresAt: '2026-03-31T23:59:59Z' });
const res = await request.post(`/api/carts/${cart.id}/discount`, { data: { code: 'SPRING10' } });
expect(res.status()).toBe(422); expect((await res.json()).error).toBe('code_expired'); const after = await request.get(`/api/carts/${cart.id}`); expect((await after.json()).totalCents).toBe(10_000);});This assumes use.baseURL (and a webServer block) in playwright.config.ts; without it, the relative URLs fail with an invalid-URL error, which is red for the wrong reason.
Red output before implementation: Expected: 200, Received: 404. That is the correct red.
Before. The ticket said “codes can only be used once”; the agent tested an in-memory Redemptions class it had written, so the test never touched the API or the database.
After. pytest-bdd binds the approved scenario text to step functions. The .feature file is copied verbatim from the criteria, and client is a FastAPI TestClient fixture from conftest.py:
from pytest_bdd import scenarios, given, when, then, parsers
scenarios("features/discount_codes.feature")
@given(parsers.parse('customer "{cid}" has already redeemed "{code}"'))def already_redeemed(seed, cid, code): seed.redemption(customer_id=cid, code=code)
@given(parsers.parse('customer "{cid}" has a cart totalling {total:f}'), target_fixture="cart")def cart_for(seed, cid, total): return seed.cart(customer_id=cid, total=total)
@when(parsers.parse('they apply "{code}"'), target_fixture="response")def apply_code(client, cart, code): return client.post(f"/carts/{cart['id']}/discount", json={"code": code})
@then(parsers.parse('the response is {status:d} with error "{error}"'))def error_response(response, status, error): assert response.status_code == status assert response.json()["error"] == error
@then(parsers.parse("the cart total is still {total:f}"))def total_unchanged(client, cart, total): assert client.get(f"/carts/{cart['id']}").json()["total"] == totalRed output before implementation: assert 404 == 409. Put the scenario ID in the scenario title (Scenario: AC-3 A customer can redeem a code only once) so the traceability check below finds it.
Before. The agent called Shop.Discounts.apply/2 directly with a struct it built in the test, which skips the controller, the changeset, and the error mapping the customer sees.
After. A ConnCase test through the router, with the criterion ID in the test name:
defmodule ShopWeb.Acceptance.DiscountCodesTest do use ShopWeb.ConnCase, async: true import Shop.Fixtures
test "AC-2: an expired code is rejected and the total is unchanged", %{conn: conn} do cart = cart_fixture(total_cents: 10_000) code_fixture(code: "SPRING10", percent: 10, expires_at: ~U[2026-03-31 23:59:59Z])
conn = post(conn, ~p"/api/carts/#{cart.id}/discount", code: "SPRING10")
assert %{"error" => "code_expired"} = json_response(conn, 422) assert Shop.Carts.get_cart!(cart.id).total_cents == 10_000 endendRed output before implementation: a Phoenix.Router.NoRouteError for the discount route. In Elixir the locked directory is test/acceptance/, not tests/acceptance/; adjust the paths in the lock below.
The “after” tests share three habits: they name the criterion, they seed state through fixtures instead of mocking the application, and they assert the outcome the customer sees. A test that only calls internal functions can pass while the endpoint returns 500.
How do you keep acceptance tests outside the agent’s edit scope?
Section titled “How do you keep acceptance tests outside the agent’s edit scope?”Instructions alone do not hold. An agent that has iterated for a while on a failing test will eventually consider “the test is wrong” and edit it. Use two layers: a lock in the implementing session where the tool supports one, and a CI guard that works regardless of tool. The full treatment, including holdout scenarios kept out of the repository, is on protecting the oracle.
Add the CI guard and code owners (all tools)
Section titled “Add the CI guard and code owners (all tools)”The CI guard is the part that makes the contract binding. Unless a pull request is labeled as a contract change, it fails on any change to the contract paths other than deleting the expected-failure markers, and CODEOWNERS routes every pull request that touches those paths to a human owner:
name: acceptanceon: pull_request: types: [opened, synchronize, reopened, labeled, unlabeled] push: branches: [main]permissions: contents: readjobs: contract: if: github.event_name == 'pull_request' runs-on: ubuntu-latest steps: - uses: actions/checkout@v7 with: fetch-depth: 0 persist-credentials: false - name: Implementation PRs may only remove expected-failure markers if: ${{ !contains(github.event.pull_request.labels.*.name, 'acceptance-contract') }} env: BASE_REF: ${{ github.base_ref }} run: | base=$(git merge-base "origin/$BASE_REF" HEAD) changes=$(git diff --name-status "$base" HEAD -- specs tests/acceptance) bad=0 while IFS=$'\t' read -r status path; do [ -n "$status" ] || continue if [ "$status" = M ] && [[ "$path" == tests/acceptance/* ]] && git show "$base:$path" | sed -E -e '/^[[:space:]]*@xfail[[:space:]]*$/d' -e 's/@xfail[[:space:]]+//g' -e 's/\btest\.fail\(/test(/g' | cmp -s - <(git show "HEAD:$path"); then echo "$path: expected-failure markers removed, nothing else" else echo "::error::$status $path changes the contract (or leaves a marker in place)"; bad=1 fi done <<< "$changes" if [ "$bad" -ne 0 ]; then echo "::error::Open a separate PR labeled acceptance-contract for the owner to approve." exit 1 fi - name: Every criterion has a test run: | for id in $(grep -ohE '([A-Z]+-)?AC-[0-9]+' specs/*.md | sort -u); do grep -rqE "(^|[^A-Z-])$id([^0-9]|$)" tests/acceptance || { echo "::error::$id has no acceptance test"; exit 1; } done open-contracts: if: github.event_name == 'push' runs-on: ubuntu-latest steps: - uses: actions/checkout@v7 with: persist-credentials: false - name: No expected-failure markers left on the default branch run: | if grep -rnE '\btest\.fail\(|@xfail\b' tests/acceptance; then echo "::error::A contract is merged but not implemented; its markers are listed above." exit 1 fi# .github/CODEOWNERS/specs/ @your-org/product-owners/tests/acceptance/ @your-org/product-owners/pytest.ini @your-org/product-owners/.github/ @your-org/product-ownersTurn on Require review from Code Owners in the branch protection rule or ruleset for your default branch; without it, CODEOWNERS only requests a review. Replace main with your default branch. Run the acceptance suite itself in the same workflow or in your existing test job. The criterion IDs above are unique per repository; if your IDs restart per feature, prefix them (DISC-AC-2). The check matches each ID whole, so AC-2 is not satisfied by a DISC-AC-2 test, and the reverse also holds.
The labeled and unlabeled event types re-run the guard when the owner adds the label, so nobody has to push an empty commit. The label only routes the pull request: anyone who can add labels, including an agent holding a gh token, can make the guard skip. The real control is Require review from Code Owners on the contract paths. The same holds for the guard itself: on pull_request, GitHub runs the workflow file from the pull request’s own branch, so an implementation pull request can edit or delete this guard and the run will pass. That is why /.github/ is in CODEOWNERS: with required Code Owner review, a changed workflow cannot merge without the owner. The workflow needs no secrets, holds a read-only token, and checks out with persist-credentials: false.
How does a red contract merge past required CI?
Section titled “How does a red contract merge past required CI?”A contract pull request is red by design, so if the acceptance suite is a required check, it cannot merge. Do not fix that by making the acceptance job optional on the default branch: an optional check is also what lets an implementation pull request merge red. Two patterns keep the check required.
Strict expected-failure markers (Playwright, pytest). The contract pull request declares each new test as expected to fail. The required job then passes only while those tests really fail, so CI re-proves the red run on every push. When the implementation makes a test pass, the runner reports the unexpected pass as a failure:
// In the contract PR: same title, so the ID check still finds ittest.fail('AC-1: a valid code takes its percentage off the cart total', async ({ request }) => { // ...body unchanged...});| Runner | Marker in the contract PR | Result once the feature works |
|---|---|---|
| Playwright | test.fail('AC-1: ...', async () => ...) | Expected to fail, but passed., job fails |
| pytest-bdd | @xfail tag on the scenario, plus xfail_strict = true under [pytest] in pytest.ini | XPASS(strict), job fails |
Deleting the markers is a human step, and it comes first. While a marker is present, a correct implementation turns the suite red (Expected to fail, but passed.), which would push an agent told to reach green to un-implement the feature. So before the implementing session starts, the contract owner opens the implementation branch with one commit that does nothing but delete that story’s markers (test.fail( back to test(, the @xfail tags). The session then works against honestly red tests and stops when the suite is green, as the prompts on this page say. That pull request needs no label: the guard compares each changed file under tests/acceptance/ with its base version minus the markers and fails on any other difference, including a marker left in place, so keep one story per test file. CODEOWNERS still routes it to the owner, who sees a test diff that is only removed markers. Both runners were checked on 2026-10-02 (Playwright 1.63.0, pytest 9.1.1 with pytest-bdd 9.0.0): green while red, failed once the stubbed feature returned the expected value. The guard script was run against marker-only, partial, weakened-assertion, added-tag, new-file, deleted-file and spec changes; only the marker-only and application-only changes passed.
Keep the strictness out of reach too. With xfail_strict = false in pytest.ini, a tagged scenario reports 1 xpassed once the feature works and 1 xfailed while it is broken, both with exit code 0, so the markers never have to come off. Run the pytest suite in CI as pytest -o xfail_strict=true tests/acceptance; the command-line override wins over the ini file (with the ini set to false, the working feature reports 1 failed and exit code 1). pytest.ini is in CODEOWNERS above for the same reason. The open-contracts job fails on the default branch while any marker remains there. It gates no pull request; it keeps a contract that is merged but not implemented visible as a red run until the implementation pull request deletes its markers.
ExUnit has no strict expected-failure mode, and a plain @tag :pending exclusion fails open (the implementation merges without the tests ever running), so Elixir projects use the next pattern.
A story branch. The contract pull request targets story/discount-codes instead of the default branch, and implementation pull requests target the story branch, so the guard compares against a base that already holds the contract. Give story/* its own ruleset with Require review from Code Owners, or the contract pull request merges without the owner. In that ruleset, require the guard (the contract job) but not the acceptance suite: the contract pull request into the story branch is red by design, and a required suite would block it exactly as it would on the default branch. The suite is required on the final pull request from the story branch into the default branch, which must be green and carries the acceptance-contract label because it brings the contract paths with it. No markers are involved, so the implementing session starts against red tests directly.
Lock the implementing session in each tool
Section titled “Lock the implementing session in each tool”Claude Code checks Edit(path) rules for every built-in file-writing tool, for file commands it recognizes in Bash such as sed and tee, and for redirect targets. Deny rules from any settings scope win over allow rules. A deny rule in the shared .claude/settings.json would also block the session that writes the tests, so pass the rules as flags on the implementation session only:
# Terminal: interactive implementation in its own worktreeclaude --worktree disc-impl \ --disallowedTools "Edit(/specs/**)" "Edit(/tests/acceptance/**)"
# Terminal or CI: headless implementation with a budgetclaude -p "Implement specs/discount-codes.md until npx playwright test tests/acceptance passes." \ --disallowedTools "Edit(/specs/**)" "Edit(/tests/acceptance/**)" \ --permission-mode acceptEdits \ --allowedTools "Bash(npx playwright test *)" \ --max-budget-usd 5The leading / anchors the pattern at the settings source. For CLI flags and for .claude/settings.json or .claude/settings.local.json, that is the primary working directory (the worktree, here). For a file passed with --settings <file>, it is that file’s own directory, so the same rules in .claude/implementer.json would deny .claude/specs/** and protect nothing; if you prefer a file, use //-absolute paths in it. Anthropic’s permissions page warns that these rules do not cover “arbitrary subprocesses that read or write files indirectly, like a Python or Node script”; for OS-level enforcement use the Bash sandbox, and keep the CI guard as the binding check. Flags checked against Claude Code 2.1.283 on 2026-09-26.
Put the rule in AGENTS.md, which Codex reads before it starts work, and let the CI guard enforce it. Since Codex 0.150.0 an untrusted project does not supply its project-level AGENTS.md, so trust the repository: accept the trust prompt the first time you run codex in it. Add this section to AGENTS.md:
## Acceptance contract- Never edit files under specs/ or tests/acceptance/. They are approved by the product owner.- If an acceptance test looks wrong, stop and explain why in your final message.Run implementation in its own managed worktree, and run the test audit read-only:
# Terminal (Codex CLI 0.157.1)codex exec --worktree --approve-for-me \ -o impl-report.md "Implement specs/discount-codes.md until the acceptance suite passes."
codex exec --sandbox read-only \ "Audit tests/acceptance against specs/discount-codes.md. Do not modify files."Codex 0.157.1 has no documented per-path write-deny rule (checked against codex --help and codex features list, 2026-09-26), so the lock is AGENTS.md plus the CI guard; a PreToolUse hook (one of 12 Codex hook events) can add a local block, covered on protecting the oracle.
Start the implementation in Plan Mode, which “creates detailed implementation plans before writing any code”, and check that the plan names the acceptance command and leaves specs/ and tests/acceptance/ alone. Add the same two rules as the Codex tab to a project rule so every agent session sees them; the file location and format are on project rules.
Cursor hooks “can observe, block, or modify behavior” at defined stages of the agent loop, so a hook can refuse edits to the contract paths. This page does not name a specific hook event; see protecting the oracle for a hook that blocks writes to the contract paths, and keep the CI guard as the enforcement.
For a Cloud Agent (including one started by an Automation on a GitHub event), the same CI guard runs on the pull request it opens, so the contract holds even with no local configuration.
If the story lives in Linear or Jira, connect the tracker so step 1 reads the ticket directly: claude mcp add --transport http linear https://mcp.linear.app/mcp in Claude Code or codex mcp add linear --url https://mcp.linear.app/mcp && codex mcp login linear in Codex (codex mcp login linear runs the OAuth sign-in that Linear’s server requires), and a remote entry under mcpServers in .cursor/mcp.json for Cursor. For UI criteria, the Playwright MCP server (@playwright/mcp) lets the test-writing session explore the running page before it writes selectors; setup is on browser automation MCP servers.
How do you know the contract is being kept?
Section titled “How do you know the contract is being kept?”The point of the contract is that nobody has to read every changed line to trust the result. Each check has an owner and a signal:
| What is checked | How | Who signs off |
|---|---|---|
| The criteria describe the right behavior | Contract pull request: criteria file, test names, red output | Product owner or tech lead, before implementation |
| The tests are red for the right reason | Red run pasted into the contract pull request | Same reviewer |
| The tests are strong enough | The audit prompt above, then mutation score on oracle strength | Tech lead, per suite rather than per pull request |
| The implementation meets the contract | Acceptance suite green in CI | CI |
| The contract was not changed to fit the code | CI guard on specs/ and tests/acceptance/ (unlabeled, only deleted markers pass), CODEOWNERS on those paths, pytest.ini and .github/ | CI, plus the code owner for marker removal and any contract change |
| Every criterion is covered | ID traceability step | CI |
The implementation pull request then carries a short evidence summary: the AC IDs and their results, the guard result, and anything the agent flagged. That summary is one field of the evidence bundle, and reading it instead of the diff is the practice described on reading evidence instead of code. The reviewer still reads the code where the risk class demands it (security, payments, migrations); the contract moves the default, not every case.
What breaks with executable acceptance criteria?
Section titled “What breaks with executable acceptance criteria?”The agent edits or skips a test. Signal: the CI guard fails, or a .skip/@pytest.mark.skip appears in the acceptance folder. Recovery: reject the pull request, restart the implementation in a fresh session with the lock loaded, and add “do not skip” to the implementation prompt. If it happens twice on the same criterion, the criterion is probably wrong; take it back to the contract pull request.
The agent special-cases the test data. The code checks if (code === 'SAVE10'). Signal: the audit prompt, or a reviewer skimming for literals from the tests. Recovery: add a second example with different values to the scenario (a scenario outline, or a second test) through a contract pull request, and keep one set of examples out of the repository as a holdout that only CI runs; see protecting the oracle.
The tests were red for the wrong reason. A missing fixture or a typo made every test fail, and the implementer “fixed” the test support code until green. Recovery: re-run step 4 and read each failure message before approving; treat tests/support/ as part of the contract if the implementer should not change it.
Tests are coupled to the implementation. The tests import a service class or assert a SQL query, so every refactor breaks them and the agent is tempted to edit them. Recovery: rewrite them against the public boundary. If the behavior has no public boundary yet, the criterion is at the wrong level; move it to a unit test the implementer owns.
The contract churns mid-story. Product changes a rule while the implementation is running. Recovery: stop the run, open a contract pull request with the changed criterion and its new red test, approve it, then restart implementation. Never let the implementing session carry the change.
The acceptance suite gets slow or flaky. Signal: developers rerun CI until green. Recovery: keep acceptance tests at three to seven per story and at the API boundary where possible; move combinatorial cases to unit or property tests; fix flakes before adding scenarios, because a flaky oracle trains everyone, human and agent, to ignore red.
Where to go next with acceptance criteria
Section titled “Where to go next with acceptance criteria”- Lock the tests properly, including holdouts and test-weakening detection: protecting the oracle.
- Run the test loop the implementing agent works inside: the test stage of the lifecycle.
- Measure whether your acceptance tests would catch a real bug: how strong is your oracle?
- Keep a feature-level spec authoritative across many stories: spec-driven development.
- Configure the permission modes and sandboxes the implementer runs under: permissions, sandboxes and approval modes.
- Shape tickets so each one arrives with testable criteria: shaping a backlog for agents.
Frequently asked questions
What is an executable acceptance criterion?
A criterion written as a Given/When/Then scenario with concrete values, paired with an automated test that fails before the feature exists and passes only when the behavior is there. It carries an ID such as AC-2 so CI can prove every criterion has a test.
Why must the acceptance tests be outside the agent's edit scope?
An agent told to make tests pass can also make them pass by editing them. If the implementing session cannot change the acceptance tests, and CI fails any implementation pull request that changes them beyond deleting the expected-failure markers, a green run means the approved behavior exists.
What does the human approve if not the diff?
The contract: the scenarios, the test names mapped to criterion IDs, and the red run that shows each test failing for the right reason. After implementation the human checks the evidence (green acceptance run, unchanged contract, traceability) rather than reading every line.