Skip to content

Building the business case for agentic engineering

A business case for agentic engineering is a decision memo that compares four options (do nothing, seats only, a shared harness, a factory pilot) against a measured baseline, with full cost, a break-even figure instead of a vendor multiplier, and stop gates written before money moves. It asks for a bounded experiment, not a productivity forecast.

Your CTO wants budget for coding agents. The vendor deck promises a multiplier; your CFO has read that METR’s 2025 randomised trial found experienced developers took 19% longer with AI. Neither number describes your codebase, review capacity or incident rate. This page gives the executive who approves the spend, and the CTO who defends it, a memo that survives that CFO.

What this business-case template gives you

Section titled “What this business-case template gives you”
  • A one-page decision memo in Markdown, with its options table, full-cost checklist, stop gates and risk register
  • A break-even calculation that replaces “3x faster” with “this pays back if accepted changes rise by N% at constant headcount and quality”
  • Two prompts: one pulls the baseline from your repository, one red-teams the memo before your CFO does

Why a multiplier-based business case fails with a CFO

Section titled “Why a multiplier-based business case fails with a CFO”

A multiplier loses the room once someone asks for its source, and third-party evidence points both ways:

SourceWhat it measuredWhat it found
METR, randomised trial, 2025-07-1016 experienced open-source developers, 246 issues, early-2025 toolsTasks took 19% longer with AI (interval +2% to +39%), yet the developers believed AI had sped them up by 20%.
METR, follow-up, 2026-02-24The same design, late-2025 toolsPoint estimates favour AI (about 18% and 4% less time in two cohorts), but both intervals cross zero; METR calls the data “an unreliable signal”, hard to interpret because of selection.
DORA, 2025 report, 2025-09-23Survey of software professionals“A positive relationship between AI adoption on both software delivery throughput and product performance”, and “a negative relationship with software delivery stability”.
Faros AI, April 2026Vendor telemetry: two years, 22,000 developers, 4,000+ teamsEpics per developer +66.2% and tasks per developer +33.7%, while bugs per developer rose 54%, incidents per pull request 242.7% and median time in review 441.5%.

Three conclusions follow:

  1. Self-reports are not a baseline. METR’s developers misjudged their own speed, so “hours saved” surveys cannot fund a budget line.
  2. Throughput gains arrive with a quality bill. DORA and Faros both show more output with less stability. DORA names the mechanism: “Without robust control systems, like strong automated testing, mature version control practices, and fast feedback loops, an increase in change volume leads to instability.”
  3. The honest ask is for an experiment. No study measures your codebase, and vendor or consultancy multipliers without a stated method measure it even less. So the memo asks for a bounded trial with a written decision rule and states the cost of being wrong. The state of agentic engineering page collects the dated evidence.

Which four options should the memo compare?

Section titled “Which four options should the memo compare?”

Compare the same four options every time, so “do nothing” is costed as seriously as the ask. Levels refer to the autonomy ladder, which rates each delivery loop, not the company.

OptionWhat you buyLadder reachWhat it can proveMain riskReversible?
A. Do nothingNo sanctioned toolsL0–L2, unmanagedNothingShadow use with no data controls and no measurementYes
B. Seats onlyLicences for every engineerL2–L3 per personAdoption and satisfaction; weak evidence on outcomesMore diffs than reviewers can read; the Faros patternYes, at renewal
C. Shared harnessSeats plus shared rules, skills, hooks, CI gates and review agents, owned by a named personL3, with a path to L4 on chosen loopsCost per accepted change and quality against a matched control teamHarness work competes with feature workYes; the harness is files in the repository
D. Factory pilotOne narrow loop run unattended against an oracle (such as dependency upgrades)L4 on one loopWhether a loop can run with evidence instead of line-by-line reviewRunaway usage; a weak oracle passes bad changesYes, if the loop is switched off by config

“Do nothing” is not the zero-risk option. DORA’s 2025 report found that 90% of survey respondents use AI at work. Without a sanctioned tool, engineers use personal accounts with no data terms, audit trail or measurement.

Recommend C and D together: C for feature teams, D for one loop whose oracle is already strong. B alone adds generation without verification: the throughput-up, stability-down pattern.

How to build the business case, step by step

Section titled “How to build the business case, step by step”
  1. Write the decision question and the decision date. For example: “Do we fund option C for 12 teams in Q1, based on an eight-week, two-team pilot ending on 2026-12-18?” A memo without a date becomes a standing budget line.

  2. Measure the baseline from systems, not surveys. For the 8–12 weeks before any change, take accepted changes per month (merged and not reverted within 30 days), lead time, change-fail rate, median review time, incidents and rework, with the definitions frozen from the metrics frameworks guide. Compute today’s cost per accepted change (the economics guide) and the fully loaded monthly cost of the delivery organisation.

  3. Cost each option in full with the checklist in the next section. Finance supplies the loaded cost per engineer; engineering supplies usage, CI and review hours.

  4. State the case as a break-even figure, as in the worked example below.

  5. Write the stop gates and the decision rule before the pilot starts. Pilot design covers cohorts, confounders and sample size.

  6. Fill in the risk register: an owner, an early signal and a funded control for every risk.

  7. Get three signatures: the CFO for the baseline and cost model, the CTO for the gates and oracle, and a named analyst outside the pilot team for the result.

What goes into the full cost of each option?

Section titled “What goes into the full cost of each option?”

Licences are the smallest line that moves; usage, review time and the quality bill decide the case.

Cost lineWhere the number comes fromOwnerOptions it applies to
Seats and plan feesVendor quote; list prices on the pricing analysisProcurementB, C, D
Metered usage and overageTool analytics and exports (see the tabs below); model rates on the models hubEngineeringB, C, D
CI computeCI billing, before and afterPlatformC, D
Harness ownershipHours of the named owner for rules, skills, hooks and gatesCTOC, D
Oracle buildEngineering hours to bring tests or evals up to the bar the loop needsTech leadD
Review timeMedian review time × accepted changes, costed at the loaded rateEngineeringB, C, D
ReworkChanges reverted or rewritten within 30 daysEngineeringAll
IncidentsIncidents attributed to changes × average cost of an incidentSREAll
Training and the adoption dipHours in training; a planned period of lower throughputEngineeringB, C, D
Security, legal and procurement reviewOne-off hours plus any new toolingCISO, legalB, C, D

For metered usage, Anthropic’s Claude Code documentation reports that across enterprise deployments “the average cost is around $13 per developer per active day and $150-250 per developer per month, with costs remaining below $30 per active day for 90% of users” (checked 2026-09-26). Treat it as a vendor planning input until pilot data replaces it. Unattended loops cost more: DoltHub’s Tim Sehn reported about $100 in Claude tokens for one hour of a multi-agent orchestrator (2026-01-15) and $3,000 for a week (2026-03-24). One user’s anecdote, but it is why option D needs a hard budget cap.

Budget the adoption dip: DORA’s ROI report describes a J-curve, an initial dip before returns, according to InfoQ’s May 2026 coverage (secondary source; the report itself was not reviewed). Write its expected length into the memo as an assumption, for example four to six weeks, and test it.

Where each tool reports the cost data you need

Section titled “Where each tool reports the cost data you need”

Cost per accepted change needs a cost record from every agent run, and each tool keeps one differently.

The Team and Enterprise analytics dashboard shows adoption and PR attribution but not per-user cost; for per-user, per-model dollars export the OpenTelemetry metric claude_code.cost.usage (see cost governance).

For option D, run each unattended task headless with a budget cap. claude -p starts in Manual permission mode, so pre-approve every tool the loop needs, or the run records a cost without changing anything:

Terminal window
# Terminal or CI: one unattended task, capped at $5, cost recorded with the result
claude -p "Upgrade the date-fns dependency to the latest minor version, fix any failing tests, and stop when npm test passes" \
--permission-mode acceptEdits --allowedTools "Bash(npm install *)" "Bash(npm test *)" \
--output-format json --max-budget-usd 5 > run.json
jq '{cost_usd: .total_cost_usd, is_error: .is_error}' run.json

total_cost_usd is a client-side estimate at list prices; reconcile it monthly against the invoice.

How much improvement does each option need to pay back?

Section titled “How much improvement does each option need to pay back?”

Replace the multiplier with a break-even figure:

break-even rise in accepted changes = added monthly cost of the option ÷ fully loaded monthly cost of the delivery organisation

That is how much accepted changes must rise, at the same headcount and with the quality gates holding, to cover the option.

The example below uses illustrative inputs for a 40-engineer organisation; replace each with your finance team’s figures.

Input (illustrative)Value
Fully loaded cost per engineer per month$15,000 (about $94 an hour at 160 hours)
Monthly cost of the delivery organisation40 × $15,000 = $600,000
Option C: seats plus usage, costed at the top of Anthropic’s reported $150–250 monthly usage range40 × $250 = $10,000
Option C: harness owner, half an engineer$7,500
Option C: extra CI compute$1,500
One-off: training, 8 hours per engineer40 × 8 × $94 ≈ $30,000
One-off: security, legal and procurement review$15,000
One-off: adoption dip, assumed four weeks (the low end of the range above) at 10% lower throughput0.10 × $600,000 = $60,000
One-offs amortised over 12 months$105,000 ÷ 12 = $8,750
Option C: added monthly cost$27,750
Break-even rise in accepted changes$27,750 ÷ $600,000 ≈ 4.6%

If you buy seats, replace the usage line with the quoted seat price from the pricing analysis plus expected overage. Review time has no line because reviewers are already in the $600,000; extra review load shows up as fewer accepted changes.

The bar looks low, and that is the point for a CFO: the risk is the quality bill, not the tool cost. Rework and incidents rising as Faros measured would erase a 5% gain unnoticed, so the stop gates below are quality gates first.

Write the outcome as three scenarios, not one number:

  • Null: no measurable rise. The organisation spent the stated pilot cost and learned where its verification is weak.
  • Below break-even: continue only if the quality gates held and the trend is rising.
  • Case met: above break-even with quality gates held. Scale to the next cohort, not to everyone.

Time saved is not cash saved. The memo says where freed capacity goes (backlog shipped earlier, a hire not made, a contract ended), or the saving never reaches the P&L.

Stop gates: write them before the money moves

Section titled “Stop gates: write them before the money moves”

The gates turn a budget request into an experiment. Adjust the thresholds to your baseline; keep the structure.

GateMetricStarting thresholdCheckedAction if breached
StabilityChange-fail rate in the pilot cohortNo more than 2 points above baselineWeeklyTwo consecutive breaches stop the pilot
ReworkAccepted changes reverted or rewritten within 30 daysNo higher than baselineEvery two weeksPause new loops; find the missing test
Review loadMedian review time per accepted changeNo more than 20% above baselineWeeklyCap pull request size; add a review agent before adding seats
Protected changesUnreviewed merges touching auth, payments, schema or migrationsZeroContinuous, enforced in CIStop the loop that produced it
CostCost per accepted change (all six terms in the economics guide)No higher than baseline by week 8MonthlyExtend once for four weeks, then stop
Spend ceilingUsage spend per loop per weekA fixed dollar cap per loopContinuous, by budget flag or gatewayThe run stops; the operator reviews

The decision rule, written into the memo before the pilot starts, reads like this: “Scale option C to the next six teams if, after eight weeks, accepted changes per engineer in the pilot cohort are at least 10% above the matched control cohort, cost per accepted change is no higher than the control’s, and no quality gate was breached in two consecutive weekly checks. Stop if any stability or protected-change gate is breached twice. Otherwise extend once for four weeks, then stop.”

The 10% bar sits above the 4.6% break-even as a detection margin, so noise between two small cohorts cannot pass for a gain. Set it from your own weekly variance with the pilot design guide.

The risk register for an agentic engineering business case

Section titled “The risk register for an agentic engineering business case”
RiskEarly signalControlOwner
Stability falls as change volume rises (DORA 2025; Faros 2026)Change-fail rate and incidents per change trend upQuality gates above; tests and fitness functions before autonomyCTO
Review becomes the bottleneckReview queue age and pull request size growPR size budget; review agents; review queue practicesTech lead
Usage cost overrunsWeekly spend per loop above planPer-run budget caps; spend alerts; cost governanceEngineering finance
Prompt injection or secrets exposureBroad agent credentials; untrusted issue text in promptsScoped agent identities; the agent threat modelCISO
IP and licence exposureNo policy on code provenance or vendor data termsVendor terms reviewed; legal and IP checklistGeneral counsel
Vendor lock-in or price changeWorkflows depend on one vendor’s proprietary featuresPortable AGENTS.md, skills and MCP; lock-in and portabilityCTO
Skills erode and juniors stop learningSeniors cannot explain merged changesDeliberate practice and review rotation; Anthropic’s 2025-12-02 internal study calls it the “paradox of supervision”Engineering managers
The measurement is gamed or confounded“Accepted change” redefined; volunteers onlyDefinitions frozen in the memo; matched control cohort; analyst outside the teamCFO

The memo is checked like code, against evidence the pilot team did not produce alone.

  • The baseline comes from version control, CI and the incident system, signed by the CFO first.
  • Every pilot change carries its evidence bundle: tests, review record, run cost.
  • The result is computed by an outside analyst against a matched control cohort over the same weeks.
  • The decision follows the pre-written rule; overrides are recorded with a reason.

Copy this into your document system: one page plus the risk register.

# Decision memo: <option> for agentic engineering
Decision requested: <fund option X for N teams, amount, period>
Decision date: <date> Pilot ends: <date>
Signatories: CFO (baseline, cost model) · CTO (gates, oracle) · Analyst (result): <name>
## 1. Options considered
| Option | Added monthly cost | What it can prove | Main risk |
| --- | --- | --- | --- |
| A. Do nothing | | | |
| B. Seats only | | | |
| C. Shared harness | | | |
| D. Factory pilot | | | |
Recommended: <option>, because <one sentence>.
## 2. Baseline (last 12 weeks, from systems of record)
Accepted changes per month: ___ Lead time (median): ___
Change-fail rate: ___ Median review time per change: ___
Rework (reverted within 30 days): ___ Incidents: ___
Monthly delivery cost: ___ Cost per accepted change: ___
## 3. Full cost of the recommended option
Seats ___ · Usage ___ · CI ___ · Harness owner ___ · Oracle build ___
Review ___ · Training and adoption dip ___ · Security and legal review ___
## 4. Break-even and scenarios
Break-even rise in accepted changes: ___%
Null: <what we spend and learn> Below break-even: <action> Case met: <action>
Where freed capacity goes: <backlog items / avoided hire / ended contract>
## 5. Stop gates (see table) and decision rule
<decision rule, verbatim>
## 6. Risk register (see table)
## 7. Evidence cited
<publisher, title, date for every external number>

What sinks an agentic engineering business case in review?

Section titled “What sinks an agentic engineering business case in review?”

The memo quotes a vendor multiplier. The CFO asks for population and method; use the break-even calculation instead.

The baseline was taken after rollout. Personal accounts contaminated “before”; compare against the control cohort and say so.

Only volunteers joined the pilot. Enthusiasts are faster with any tool; METR says selection made its 2026 data hard to interpret. Match cohorts on team type and codebase (pilot design).

Output rose and nobody counted the quality bill. Recompute cost per accepted change with rework and incidents.

Seat plans hid the marginal cost. Cost pilot usage at API rates for the scale scenario.

The dip was read as failure. Judge the last four weeks, not the first.

Edit page

Last updated:

Cite this page — https://developertoolkit.ai/en/strategy/business-case/, developertoolkit.ai