Skip to content

Designing a pilot that proves something

A pilot that proves something about AI coding agents compares a treatment cohort with a matched control cohort over the same 4–8 weeks, against a baseline taken from version control, CI and incident data. It names its confounders, sizes its sample in advance, and commits to a written decision rule, so the result, not opinion, decides the budget.

A volunteer team got licences six weeks ago. They say they are “much faster”, pull requests per engineer went up, and a reviewer on another team says the changes are harder to trust. Nobody can say whether delivery improved, and the CTO’s budget meeting is next month.

  • Four pilot designs, matching criteria and a nine-row confounder register
  • A sample-size calculation you rerun on your own baseline
  • A decision rule and a one-page charter you sign before week one
  • Cohort tagging for Claude Code, Codex and Cursor, plus two prompts: baseline variance and a pre-launch design audit

The pilot feeds the business case: the case asks for the experiment, and this page makes the experiment worth its cost. It assumes you have already chosen your metric definitions with the metrics frameworks guide.

The typical pilot hands licences to volunteers, asks them after a month whether they feel faster, and counts pull requests. Each choice alone makes the result uninterpretable:

SourceWhat it showsWhat it means for your pilot
METR randomised trial, 2025-07-10 (16 experienced open-source developers, 246 issues)Tasks took 19% longer with AI (interval +2% to +39%), yet developers “still believed AI had sped them up by 20%”.A survey measures belief, not speed. Measure from systems.
METR follow-up, 2026-02-24Point estimates favour AI (about 18% and 4% less time for two cohorts), but both intervals cross zero. METR says the new results are “subject to increased selection” and that time-on-task is unreliable for developers “who use multiple AI agents concurrently”.Who joins changes the answer; per-task timing breaks with parallel agents.
DORA 2025 report, 2025-09-23“A positive relationship between AI adoption on both software delivery throughput and product performance”, and “a negative relationship with software delivery stability”.Count quality in the same window as throughput.
Faros AI telemetry report, April 2026 (22,000 developers, more than 4,000 teams)Tasks per developer +33.7% while bugs per developer rose 54% and median time in review rose 441.5%.Output can rise while delivery gets worse. Guardrails are not optional.

Write one question, one population and one primary metric. A pilot that tries to answer “is AI good for us?” answers nothing.

Name the intervention precisely, because the control cohort almost certainly uses AI already (DORA 2025: 90% of survey respondents use AI at work). You are testing a specific setup against current practice, not “AI versus no AI”:

“Does the shared harness (Claude Code with our rules, skills, CI gates and a review agent) reduce cost per accepted change by at least 10% against current practice, for product teams working in the monorepo, without raising the change-fail rate?”

Pick the primary metric by the option the business case funds:

Business-case optionPrimary metricGuardrails (must not get worse)
B. Seats onlyAccepted changes per engineer-weekChange-fail rate, rework within 30 days, median review time
C. Shared harnessCost per accepted change, all cost linesChange-fail rate, rework, review time, incidents
D. Factory pilot on one loopShare of loop tasks that pass the oracle and are accepted without reworkEscaped defects from the loop, spend per task, protected-path violations

An accepted change is a change merged to the main branch and not reverted or rewritten within 30 days. Freeze that definition in the charter, and count per ticket where you can: agents make splitting one change into five pull requests cheap.

  1. Write the question, the primary metric and the guardrails. One primary metric only; everything else is a guardrail or an exploratory note.

  2. Take the baseline from systems of record. Pull the last 8–12 weeks for every candidate team: accepted changes, lead time, change-fail rate, median review time, rework and incidents. Record the variance as well as the mean, because the sample size depends on it. The first prompt below computes PR open-to-merge time, a proxy for lead time, from gh; if you define lead time as first commit to production, measure that instead and keep one definition throughout.

  3. Choose the design and the unit of assignment. For most organisations: matched concurrent cohorts at team level (compared in the next section).

  4. Build the cohorts by matching, then assign by coin toss. Pair teams on the matching criteria, then randomly pick the treatment team in each pair. That removes the “we chose the best team” objection.

  5. Fill in the confounder register. For each of the nine confounders below, write the control or why it does not apply.

  6. Size the sample. Run the calculation below on your baseline variance. If the pilot cannot detect the effect the business case needs, change the design now, not after the readout.

  7. Write and sign the decision rule. Scale, extend or stop, with thresholds. The CTO signs the guardrails, the CFO signs the cost definitions, and a named analyst outside the pilot team computes the result.

  8. Instrument both cohorts before week one. Tag every agent run with its cohort (see the tabs below) and confirm the control cohort’s metrics reach the same dashboards.

  9. Run for the planned weeks, then read out twice. Interim readout at pilot end; final readout 30 days later, when the rework window for the last changes has closed.

Which pilot design fits your organisation?

Section titled “Which pilot design fits your organisation?”
DesignHow assignment worksUse it whenWhat it cannot do
Matched concurrent cohorts (default)Teams paired on similar work; a coin toss per pair picks treatmentYou have at least four comparable teamsDetect small effects with few teams; see sample size
Staggered rollout (stepped wedge)Every team gets the treatment, at randomly ordered start dates 2–4 weeks apartWithholding the tool is politically impossibleFinish fast
Task-level randomisation (METR-style)Each ticket is randomly marked “agents allowed” or not before work startsYou want the strongest causal answer for one teamSurvive parallel agents, which make time-on-task unreliable (METR)
Before and after on one teamNo control; compare with the team’s own baselineOnly as a supplement to one of the aboveSeparate the tool from the season, the roadmap or a reorganisation

Never compare volunteers with non-volunteers: enthusiasts differ in ways the tool did not cause, which is METR’s 2026 selection caveat.

For a factory pilot on one loop (option D), the unit is the task, not the team. Randomly route half of the loop’s incoming tasks (for example, dependency upgrades) to the unattended agent and half to the usual process, and compare acceptance, rework and cost per task.

Match on what drives delivery metrics more than any tool does, at least the first four rows.

Match onWhy it mattersHow to check
Codebase and languageCoverage and build speed decide how much an agent can verify itselfSame repository or stack; coverage within a few points
Work typeGreenfield features, maintenance and incident work move differentlyShare of ticket types over the baseline weeks
Team size and seniority mixSeniors and juniors adopt agents differentlyHeadcount and a rough seniority split
Baseline metric levelRegression to the mean makes the worst team look improved anywayPrimary metric within the same quartile
Release calendarA launch freeze or a quarter-end crunch swamps any tool effectNo known freeze or major launch inside the pilot window

Which confounders break an AI coding pilot?

Section titled “Which confounders break an AI coding pilot?”

A confounder is anything other than the intervention that changes the metric during the pilot. List each in the charter with its control.

ConfounderHow it shows upControl
Volunteer selectionThe keenest team gets the tool and was already the fastestMatched pairs, coin toss inside each pair
Novelty and observationEveryone works harder because they are being watchedTell both cohorts they are measured; judge the last four weeks
The adoption dipThroughput falls in the first weeks while people learnReport weeks 1–2 separately; do not stop on throughput alone before week 4
ContaminationControl engineers use personal AI accounts or copy the treatment team’s rulesRecord the control’s current tools; in a shared repository, ship the treatment harness through managed or user-level settings (or a plugin) enabled only on treatment machines, not through files committed to the repo
Work-mix shiftTreatment engineers pick tickets that suit agentsAssign tickets through the normal backlog; record a size estimate before work starts, as METR did
Metric gamingPull requests get smaller and more numerous; “accepted” gets redefinedCount per ticket; freeze definitions in the charter; the analyst sits outside the team
Concurrent changeA CI migration, reorganisation or new on-call rota lands mid-pilotFreeze other process changes in both cohorts, or log them with dates
Shared reviewersAgent pull requests lengthen a shared review queueKeep review inside cohorts, or measure review time per reviewer
Parallel agent runsPer-task timing breaks when one engineer runs three agentsMeasure per engineer-week and per accepted change

Most pilots are too small to detect the effect they claim. The standard two-group formula, two-sided at 5% significance and 80% power:

changes per cohort = 15.7 × σ² ÷ δ²

σ is the standard deviation of the metric per change and δ the smallest effect worth detecting. Changes from one engineer are correlated, so multiply by the design effect 1 + (m − 1) × ICC, where m is changes per engineer over the pilot and ICC the share of variance between engineers.

A worked example for PR open-to-merge time, the lead-time proxy, with illustrative inputs: replace them with the values the first prompt computes.

Input (illustrative)Value
σ of log open-to-merge time per change1.0
Effect to detect: 30% shorter timeδ = |ln 0.7| ≈ 0.357
Changes per cohort before clustering15.7 × 1.0 ÷ 0.127 ≈ 124
Changes per engineer over 8 weeks (m)30
ICC between engineers0.1
Design effect1 + 29 × 0.1 = 3.9
Changes per cohort after clustering124 × 3.9 ≈ 484, about 17 engineers per cohort

For a 20% effect the same inputs need about 1,230 changes, or 41 engineers, per cohort. That is the arithmetic behind the most common pilot mistake: one team of eight cannot detect anything short of a large effect.

When the numbers do not fit, change the design:

  • Add teams, not weeks. Because of clustering, more engineers help far more than more weeks from the same engineers.
  • Adjust for the baseline. Compare each engineer with their own baseline; it removes stable differences between people and shrinks σ.
  • Pick a higher-volume unit. A factory loop with hundreds of tasks a month reaches a useful size in four weeks.
  • Decide on intervals and guardrails, not on a p-value alone: guardrails hold and the confidence interval clears a threshold.

The decision rule turns the readout into an action nobody can renegotiate. Adjust the thresholds to your baseline and business case.

OutcomeConditionAction
ScalePrimary metric improves by at least the business-case threshold (for example, cost per accepted change 10% below control), and no guardrail was breached in two consecutive weeksExtend to the next cohort of teams, not to everyone
Extend oncePrimary metric improves but below the threshold, or the interval is too wide to decide, and guardrails heldRun four more weeks with the same cohorts, then apply the rule again with no further extension
StopAny stability or protected-path guardrail is breached twice, or the primary metric is flat or worse at the final readoutStop the treatment and record what the pilot learned about your verification gaps

Write the rule into the charter verbatim; any leadership override goes in too, with its reason. The board report cites that record.

How to tag pilot runs in Claude Code, Codex and Cursor

Section titled “How to tag pilot runs in Claude Code, Codex and Cursor”

Quality metrics come from git, CI and the incident system, which are tool-neutral: the analyst joins them to the cohort list. What differs by tool is how you get cost and usage per cohort, which option C needs.

Claude Code exports OpenTelemetry metrics, including claude_code.cost.usage (USD), claude_code.token.usage, claude_code.pull_request.count and claude_code.commit.count. OTEL_RESOURCE_ATTRIBUTES adds your own keys to every metric and event, so a cohort label becomes a filter in your dashboards. Set it in the managed settings file on the treatment engineers’ machines:

{
"env": {
"CLAUDE_CODE_ENABLE_TELEMETRY": "1",
"OTEL_METRICS_EXPORTER": "otlp",
"OTEL_LOGS_EXPORTER": "otlp",
"OTEL_EXPORTER_OTLP_PROTOCOL": "grpc",
"OTEL_EXPORTER_OTLP_ENDPOINT": "http://collector.example.com:4317",
"OTEL_RESOURCE_ATTRIBUTES": "pilot.id=harness-2026q4,pilot.cohort=treatment,team.id=payments"
}
}

Managed settings apply to every user the file is deployed to, so deploy this file only to treatment machines (or put the block in each engineer’s ~/.claude/settings.json). Give control engineers the same block with pilot.cohort=control, so both cohorts reach the same dashboards.

Two traps, checked against Claude Code 2.1.283. Claude Code ignores the exporter variables in a repository’s .claude/settings.json: use managed settings or each engineer’s ~/.claude/settings.json. And OTEL_RESOURCE_ATTRIBUTES values cannot contain spaces (team.id=payments_core, not quoted text).

On Team and Enterprise plans with the GitHub app installed, the analytics dashboard labels merged pull requests with Claude Code contributions claude-code-assisted. Contribution metrics are in public beta and unavailable with zero data retention: a cross-check, not the cohort definition.

Telemetry setup is covered in depth on the agent telemetry page.

Fill in every field of this one-page charter before week one.

# Pilot charter: <intervention> — <pilot id>
Question: Does <intervention> change <primary metric> by at least <threshold>
against current practice, for <population>, without worsening <guardrails>?
Business case: <link> Decision owner: <name> Analyst (outside the team): <name>
## Design
Design: matched concurrent cohorts | staggered rollout | task randomisation
Unit of assignment: team | engineer | task
Pairs and coin-toss result: <team A vs team B → A treatment>, ...
Pilot weeks: <start> to <end> (4–8 weeks) Final readout: <end + 30 days>
## Metrics (definitions frozen at signature)
Primary: <metric, exact definition, source system>
Guardrails: change-fail rate | rework within 30 days | median review time | incidents
Accepted change: merged to main and not reverted or rewritten within 30 days
## Baseline (last 8–12 weeks, per cohort)
Primary: mean ___ sd ___ Guardrails: ___
## Sample size
sigma ___ delta ___ ICC ___ m ___ → engineers per cohort ___
## Confounder register
| Confounder | Control | Owner |
## Decision rule (verbatim)
Scale if ___. Extend once if ___. Stop if ___.
## Signatures
CTO (guardrails) ___ CFO (cost definitions) ___ Analyst ___ Date ___

Nobody on the pilot team grades their own pilot.

  • The baseline comes from git, CI and incident data before the coin toss; the CFO signs it with the cost definitions.
  • Every treatment change carries its evidence: tests passed, review record and the cost of the run. The evidence bundle format lets the analyst sample changes instead of trusting a dashboard.
  • The analyst outside the team computes the metrics from the frozen definitions, reports the interval, and applies the decision rule.
  • Protected paths (auth, payments, schema, migrations) are enforced in CI in both cohorts.
  • The final readout waits for the 30-day rework window.

What goes wrong in AI coding pilots, and how to recover

Section titled “What goes wrong in AI coding pilots, and how to recover”

The pilot team was the volunteers. The result will be dismissed as selection, rightly. Keep them as the first treatment team, find a matched control, and run a second, randomly assigned staggered wave.

The readout was a survey. METR’s developers felt faster while measuring slower. Rebuild the comparison from git and CI data for the same weeks.

Pull requests per engineer rose, and the pilot declared success. Recompute per ticket with rework and incidents. If the guardrails moved, that is the Faros pattern, not a win.

The control team adopted the treatment’s rules halfway through. Date the leak, analyse the weeks before it, and treat the rest as a staggered rollout.

Throughput fell in weeks two and three, and a sponsor wants to stop. That is the planned adoption dip. Stopping before week 4 on throughput alone is outside the rule; a guardrail breach is inside it.

The result is inside the noise. Apply the “extend once” branch, add teams rather than weeks, and do not extend twice.

The dashboards cannot separate cohorts. Move the treatment cohort onto the tagged telemetry above and use only the tagged weeks for the cost metric.

After a “scale” decision: the team adoption roadmap per repository, the transformation roadmap across the organisation. Dated evidence: the state of agentic engineering.