Product management when build time collapses
Product management when build time collapses means the product manager’s output shifts from tickets to verifiable intent: an intent.md with a measurable outcome and acceptance criteria that become failing tests before an agent writes code. Roadmaps become ranked outcome bets gated by evidence, and prototypes stay in a separate lane that cannot merge to production.
On Monday a designer and an agent built a clickable “pause subscription” flow in three hours. On Wednesday the CEO saw it in the all-hands demo and asked why it was not live yet. The roadmap still lists features by sprint, the prioritization meeting still argues about story points, and nobody has written down what the pause flow must do when an annual customer opens it.
This page is for executives who own the product organization and for tech leads who work beside product managers at Level 4 of the autonomy ladder.
What you can adopt for product management with agents
Section titled “What you can adopt for product management with agents”- A product
intent.mdtemplate with outcome, kill criterion, acceptance criteria and a proposed risk class, plus a filled example. - Three copy-paste prompts a product manager runs in Claude Code, Codex or Cursor: draft the intent, stress-test the criteria, and build a throwaway prototype.
- An outcome roadmap format whose capacity unit is review and rollout slots, not engineer-weeks.
- A 45-minute prioritization agenda and a scoring rule in which build effort is replaced by verification and ownership cost.
- A prototype-to-production gate list with a CI guard you can paste into GitHub Actions.
- Five metric definitions that tell you whether the product side of the system works.
Why does discovery become the bottleneck when agents build?
Section titled “Why does discovery become the bottleneck when agents build?”When an agent can produce a working feature in an afternoon, the slow step is deciding what is worth building, stating what “correct” means, and learning whether the result helped a user.
The people closest to the tools say so. A line Andrej Karpathy quotes approvingly in his Sequoia Ascent 2026 summary (2026-04-30) reads: “I am becoming the bottleneck of even knowing what we are trying to build, why it is worth doing, and how to direct my agents”. Dan Shapiro’s description of Level 4 has the human “write a spec”, then “leave for 12 hours, and check to see if the tests pass”, working as “more of an engineering manager or product/program/project manager” (The Five Levels, 2026-01-23). DORA’s AI Capabilities Model (Google Cloud, 2025-09-23) lists user-centric focus among the seven capabilities that amplify the benefits of AI: faster building pays off only where someone keeps the work tied to a user outcome.
More building does not mean more shipping. Faros AI’s vendor telemetry from 22,000 developers on its platform (AI Engineering Report 2026, April 2026) shows epics completed per developer up 66.2% while deployments per week fell 11.7% and incidents per pull request rose 242.7%. Built-but-unreleased or released-then-broken work is inventory; product management decides how much of it starts.
Cheap building raises the volume of optional work. In Anthropic’s internal study (2025-12-02, a vendor’s self-report from 132 engineers and researchers), “27% of Claude-assisted work consists of tasks that wouldn’t have been done otherwise”. Some is valuable; all of it needs a decision, a verifier and an owner.
What does a product manager write when agents do the building?
Section titled “What does a product manager write when agents do the building?”The product manager moves from tickets engineers interpret to artifacts machines can check. The artifact chain names the files; this table shows which ones product owns.
| Artifact | Product’s role | Engineering’s role | Why this split |
|---|---|---|---|
intent.md: problem, evidence, outcome, kill criterion | Writes and accepts | Reviews for feasibility, proposes a risk class | Product owns the why and the success measure |
| Acceptance criteria (Given/When/Then, with IDs) | Writes the first draft, approves the final set | Turns them into failing tests, flags untestable lines | The criteria become the oracle, so both sides must agree on them |
spec.md: design, data, interfaces | Reads and flags product conflicts | Writes and accepts | Architecture decides how safely the team can move |
Implementation and plan.md | None | Agents build; engineers operate and review | The human decision is acceptance, not authorship |
| Release and measurement | Reads the metric and decides keep or kill | Ships behind a flag, instruments, rolls back | Separation of duties survives the change |
The biggest change is the second row. “Users can pause instead of cancelling” is a wish; five numbered Given/When/Then criteria with concrete dates and amounts are a contract a test can hold an agent to. How engineers turn that contract into locked, failing tests is on executable acceptance criteria.
A product intent.md template you can adopt as-is
Section titled “A product intent.md template you can adopt as-is”This extends the artifact-chain template with the four fields product owns at Level 4: evidence, outcome metric, kill criterion and acceptance criteria. The example uses illustrative numbers; replace every value with yours.
# Intent: Pause a subscription instead of cancellingAuthor: Maria Nowak (Product, Billing). Status: proposed. Record: BILL-412
## ProblemMonthly customers who want a short break have only one option: cancel.Exit-survey answer "temporary break" is the top cancellation reason this quarter.
## Evidence- Exit survey export, 2026-07-01 to 2026-09-20 (link)- Six customer calls, notes in research/pause-calls.md
## Outcome and metricShare of accounts that open the cancel flow and are still paying 60 days later.Baseline: measured from the last full quarter. Read on day 30 and day 60 after release.
## Kill criterionIf the 60-day metric does not move, or paused accounts resume at a lower rate thancancelled accounts return, remove the pause option.
## Acceptance criteriaAC-1 Given an active monthly plan billed on the 5th, when the owner pauses for one month on 2026-10-01, then no invoice is issued on 2026-10-05 and the next invoice is issued on 2026-11-05.AC-2 Given a paused subscription, when any team member signs in, then the workspace is read-only and a banner shows the resume date.AC-3 Given an annual plan, when the owner opens the cancel flow, then pause is not offered.AC-4 Given an account that paused in the last 12 months, when the owner opens the cancel flow, then pause is not offered.AC-5 Given a paused subscription, when the owner clicks Resume on 2026-10-20, then billing restarts that day and the invoice is prorated to the 5th.
## Out of scopePausing annual plans. Pauses longer than one month. Changes to the dunning flow.
## Proposed risk classCritical (money movement). Engineering lead confirms.
## Open questionsDoes a paused workspace keep its API tokens active?Two fields do most of the work. The kill criterion is written before anything is built, so “we already built it” never keeps a feature alive. The proposed risk class uses the four classes (low, medium, high, critical) defined on the one-map page and assigned by the policy in governance and autonomy. Billing is money movement, so it is critical even when the diff is small, and everyone knows up front that it needs a named owner and a second approver at the production gate.
How does a product manager draft intent with an agent?
Section titled “How does a product manager draft intent with an agent?”The agent needs to read the tracker and stay out of application code. Connect the tracker once, then run the session in a planning mode.
Connect Jira and Confluence through the Atlassian Rovo MCP server, then start a session in plan mode so Claude can read but not edit application code:
# Terminal, in the product repositoryclaude mcp add --transport http atlassian https://mcp.atlassian.com/v2/mcpclaude --permission-mode planInside the session, run /mcp to complete the OAuth sign-in. For Linear, the server URL is https://mcp.linear.app/mcp. Leave plan mode (Shift+Tab) only to save the finished intent/*.md file. Flags checked against Claude Code 2.1.283 on 2026-09-26.
Add the same server and sign in, then use /plan so Codex proposes a file instead of changing the product:
# Terminal (Codex CLI 0.157.1)codex mcp add atlassian --url https://mcp.atlassian.com/v2/mcpcodex mcp login atlassiancodexStart the conversation with /plan. When the draft is final, leave plan mode and ask Codex to save it under intent/. For Linear, use --url https://mcp.linear.app/mcp. Codex reads AGENTS.md, so add one line there: “Files under intent/ are written by product; do not edit application code in intent sessions.”
Install the Atlassian plugin from Cursor’s marketplace, or add a remote entry to .cursor/mcp.json:
{ "mcpServers": { "atlassian": { "url": "https://mcp.atlassian.com/v2/mcp" } } }Open the agent in Plan Mode, which “creates detailed implementation plans before writing any code”, and attach research notes or screenshots next to the prompt. Switch to Agent mode only to save the finished markdown file. Cursor details were checked against its docs on 2026-08-28.
If you want a reusable interviewer rather than a prompt, the grill-me skill from mattpocock/skills questions you one decision at a time and writes no code (about 1.2 million installs on skills.sh, snapshot read 2026-09-26, secondary). Install it with claude plugin install mattpocock-skills in Claude Code, or npx skills add mattpocock/skills --skill grill-me grilling setup-matt-pocock-skills -a codex -a cursor for Codex and Cursor. grill-me delegates to grilling, which is why both names appear, and setup-matt-pocock-skills runs once per repository to record your issue tracker. Install one or the other, not both: the plugin and the skills.sh copy register duplicate skills. Setup is on Matt Pocock’s skills.
Where do prototypes fit, and which gates keep them out of production?
Section titled “Where do prototypes fit, and which gates keep them out of production?”A clickable prototype answers in an afternoon what a document answered in a fortnight. The risk is that it looks finished: it has no acceptance tests, it runs on fake data, and nobody chose its risk class. Keep two lanes and one-way doors between them.
| Prototype lane | Production lane | |
|---|---|---|
| Purpose | Answer a product question | Deliver an accepted outcome |
| Starts from | A question in intent.md → Open questions | An accepted intent.md with approved criteria |
| Where it runs | A prototype/* branch, a local or throwaway environment, fake or seeded data | Normal branches, CI, the release pipeline |
| Credentials | No production credentials or customer data | Scoped by risk class |
| Output | What was learned, recorded in intent.md; the code is deleted | A merged change with an evidence bundle |
| Who decides it is done | The product manager | The code owner; for critical class, a named owner and a second approver |
The prototype skill in the same mattpocock/skills collection packages this behaviour (“Build a throwaway prototype to answer a design question”); install it with npx skills add mattpocock/skills --skill prototype -a codex -a cursor, or through the mattpocock-skills plugin in Claude Code.
The prototype-to-production gates
Section titled “The prototype-to-production gates”A prototype never merges. What crosses into the production lane is the learning, as edits to intent.md, and the build restarts from the approved contract. Agents make that rebuild cheap enough to afford.
-
G1: Intent accepted. The product manager marks
intent.mdaccepted, with the prototype’s findings recorded under Evidence. -
G2: Contract approved. Engineering turns the acceptance criteria into failing tests with matching AC IDs, and the product manager approves them while they are red. The tests are then locked outside the implementing agent’s edit scope.
-
G3: Risk class confirmed. The engineering lead confirms or raises the proposed class. A critical change names the owner and the second approver who authorize the release at G5.
-
G4: Evidence complete. The pull request carries the evidence bundle: the acceptance run, type and lint gates, security scan and review-agent findings, plus the human review its risk class requires.
-
G5: Released reversibly. The change ships behind a feature flag through progressive delivery, with the outcome metric instrumented and a rollback path tested.
-
G6: Measured and decided. On the dates in
intent.md, the product manager reads the metric against the kill criterion and records keep or remove.
G1–G3 and G6 are human decisions; G4 and G5 are enforced by tooling (CI checks, the evidence bundle, the flag and rollback). The simplest mechanical gate refuses prototype branches outright. Save this as .github/workflows/prototype-guard.yml and make the block-prototypes job a required status check on main:
name: prototype-guardon: pull_request: branches: [main] types: [opened, synchronize, reopened, edited, labeled, unlabeled]
permissions: {}
jobs: block-prototypes: runs-on: ubuntu-latest steps: - name: Refuse prototype branches and labels if: startsWith(github.head_ref, 'prototype/') || contains(github.event.pull_request.labels.*.name, 'prototype') run: | echo "Prototype work does not merge. Record the learning in intent/*.md and build from the approved contract." exit 1 - name: Require a linked intent if: github.actor != 'dependabot[bot]' env: BODY: ${{ github.event.pull_request.body }} run: | printf '%s' "$BODY" | grep -Eq 'intent/[a-z0-9-]+\.md' || { echo "Link the accepted intent/*.md file in the pull request description." exit 1 }The workflow reads only event data, so it runs with permissions: {}, checks out no code, and passes the pull request body through an environment variable instead of interpolating it. The second step is a lightweight G1 check: every change names the intent it serves. Exempt other bots and infrastructure-only changes as your repository needs. The acceptance-test lock that enforces G2 is on executable acceptance criteria.
What does a roadmap look like at Level 4?
Section titled “What does a roadmap look like at Level 4?”A dated feature list assumes that building is scarce and that a built feature stays. At Level 4 building is cheap, verifying and owning are not, and some shipped bets should be removed. The roadmap becomes a ranked list of outcome bets, each carried by an intent.md, each at a visible gate.
| Horizon | Outcome bet | Metric and read date | Intent | Gate | Owner | Review slots |
|---|---|---|---|---|---|---|
| Now | Fewer voluntary cancellations on monthly plans | 60-day retention after cancel flow; 2026-12-15 | intent/pause-subscription.md | G2: contract approved | M. Nowak | 1 critical-class slot |
| Now | Faster first invoice export for finance admins | Median time to first export; 2026-11-30 | intent/csv-export.md | G5: behind flag | J. Kowal | — |
| Next | Self-serve seat reduction | Support tickets tagged “seats” per month | intent/seat-reduction.md | G1: prototype answered two questions | M. Nowak | 1 critical-class slot |
| Later | Usage-based add-on | To be defined | none yet | Discovery | — | — |
Three rules make this format work:
- Capacity is review and rollout slots, not engineer-weeks. A team that can review one critical-class change at a time can run one critical-class bet at a time, however fast agents build. Org design shows how to measure that capacity.
- “Now” holds only items past G1. An idea without an accepted intent is discovery, shown as its own row, not as a feature with a date.
- Every “Now” row has a read date. The read date, not the ship date, is the commitment to leadership.
How does the prioritization meeting change?
Section titled “How does the prioritization meeting change?”The meeting stops negotiating story points and becomes a weekly decision forum about evidence, contracts and capacity. Adopt this agenda for a 45-minute session with the product manager, the engineering lead and the eval owner for the domain.
| Minutes | Item | Decision recorded |
|---|---|---|
| 0–10 | Read-outs. Bets whose read date has passed: metric against kill criterion | Keep, iterate or remove, per bet |
| 10–20 | Prototype results. What each prototype answered | Update intent.md, run another prototype, or drop |
| 20–30 | Contracts. Intents whose acceptance criteria are ready | Approve for G2, or return with named gaps |
| 30–40 | Capacity. Review and rollout slots free per risk class this week | Which approved contracts start, in what order |
| 40–45 | Kill list. Anything in “Now” without progress for two cycles | Remove or re-justify |
Rank the contracts competing for slots with a rule the whole room can apply:
priority = (expected outcome impact × confidence in the evidence) ÷ (verification cost + ongoing ownership cost)Build effort is deliberately absent. Verification cost is the review minutes its risk class needs plus the tests and evals it must add. Ownership cost is what it costs to keep correct: on-call, support, a dependency, a regulated data path. A cheap feature that needs a critical-class reviewer every release ranks below a harder low-class one. Score each factor from 1 to 5 and write the score into the roadmap row.
How do you know product management is working at Level 4?
Section titled “How do you know product management is working at Level 4?”These five metrics come from git history, the tracker and your analytics tool.
| Metric | Definition | Source | Owner | Healthy direction |
|---|---|---|---|---|
| Intent lead time | Days from first commit of intent.md to its accepted status | Git history of intent/ | Head of product | Down, without a rise in contract rework |
| Contract rework rate | Share of changes whose acceptance tests changed after G2 approval | Git history of the locked test paths | Engineering lead | Low and stable; a spike means criteria were vague |
| Read-out rate | Share of shipped bets whose metric was read on its date and a keep or remove decision recorded | Roadmap and intent.md status | Head of product | Toward 100% |
| Removal rate | Share of read-out bets removed under their kill criterion | intent.md status | Head of product | Above zero; zero means kill criteria are not real |
| Prototype leak count | Changes merged to main that trace to a prototype/* branch or skipped G1–G2 | CI guard logs and pull request audit | Tech lead | Zero |
Read them beside the delivery metrics on the metrics frameworks page: a falling intent lead time with a rising change failure rate means product is feeding the pipeline faster than engineering can verify.
What breaks when product managers write the intent?
Section titled “What breaks when product managers write the intent?”| Failure | What you see | Recovery |
|---|---|---|
| The demo becomes the commitment | A prototype shown to leadership is expected in production next week | Show prototypes only with their gate status on screen; the answer to “when” is “after G2”, with the read date |
| Criteria written as prose | Engineers ask the same clarifying questions on every intent; contract rework rate climbs | Run the stress-test prompt before handoff; reject intents without concrete values at the contracts item |
| The product manager writes the solution | intent.md names tables, endpoints or components | Move design detail to spec.md; keep intent to problem, outcome and observable behaviour |
| Moving the oracle to make it pass | Acceptance criteria are edited after G2 so an agent’s build passes | Any change to approved criteria goes back to the contracts item and re-approval; the CI lock described in protecting the oracle makes it visible |
| Roadmap flood | “Now” grows every week because building feels free | Cap “Now” by review slots; enforce the kill list |
| Kill criteria that never fire | Removal rate stays at zero for two quarters | Ask at every read-out what result would have removed the feature; if nobody can answer, rewrite the criterion |
| Review queue saturates | Contracts wait at G4 while intents keep arriving | Stop approving new contracts until the queue drains; see running the review queue |
Questions to ask your head of product and your CTO
Section titled “Questions to ask your head of product and your CTO”- Which of our last ten shipped features had a written kill criterion, and how many did we remove?
- Who writes acceptance criteria today, and could a test fail on each of them?
- What stops a prototype branch from merging to
main, mechanically rather than by agreement? - How many critical-class changes can we verify per week, and does the roadmap respect that number?
- Which metric is our roadmap committed to: ship dates or read dates?