Designing a pilot that proves something
A pilot that proves something about AI coding agents compares a treatment cohort with a matched control cohort over the same 4–8 weeks, against a baseline taken from version control, CI and incident data. It names its confounders, sizes its sample in advance, and commits to a written decision rule, so the result, not opinion, decides the budget.
A volunteer team got licences six weeks ago. They say they are “much faster”, pull requests per engineer went up, and a reviewer on another team says the changes are harder to trust. Nobody can say whether delivery improved, and the CTO’s budget meeting is next month.
What this pilot design gives you
Section titled “What this pilot design gives you”- Four pilot designs, matching criteria and a nine-row confounder register
- A sample-size calculation you rerun on your own baseline
- A decision rule and a one-page charter you sign before week one
- Cohort tagging for Claude Code, Codex and Cursor, plus two prompts: baseline variance and a pre-launch design audit
The pilot feeds the business case: the case asks for the experiment, and this page makes the experiment worth its cost. It assumes you have already chosen your metric definitions with the metrics frameworks guide.
Why most AI coding pilots prove nothing
Section titled “Why most AI coding pilots prove nothing”The typical pilot hands licences to volunteers, asks them after a month whether they feel faster, and counts pull requests. Each choice alone makes the result uninterpretable:
| Source | What it shows | What it means for your pilot |
|---|---|---|
| METR randomised trial, 2025-07-10 (16 experienced open-source developers, 246 issues) | Tasks took 19% longer with AI (interval +2% to +39%), yet developers “still believed AI had sped them up by 20%”. | A survey measures belief, not speed. Measure from systems. |
| METR follow-up, 2026-02-24 | Point estimates favour AI (about 18% and 4% less time for two cohorts), but both intervals cross zero. METR says the new results are “subject to increased selection” and that time-on-task is unreliable for developers “who use multiple AI agents concurrently”. | Who joins changes the answer; per-task timing breaks with parallel agents. |
| DORA 2025 report, 2025-09-23 | “A positive relationship between AI adoption on both software delivery throughput and product performance”, and “a negative relationship with software delivery stability”. | Count quality in the same window as throughput. |
| Faros AI telemetry report, April 2026 (22,000 developers, more than 4,000 teams) | Tasks per developer +33.7% while bugs per developer rose 54% and median time in review rose 441.5%. | Output can rise while delivery gets worse. Guardrails are not optional. |
What question should the pilot answer?
Section titled “What question should the pilot answer?”Write one question, one population and one primary metric. A pilot that tries to answer “is AI good for us?” answers nothing.
Name the intervention precisely, because the control cohort almost certainly uses AI already (DORA 2025: 90% of survey respondents use AI at work). You are testing a specific setup against current practice, not “AI versus no AI”:
“Does the shared harness (Claude Code with our rules, skills, CI gates and a review agent) reduce cost per accepted change by at least 10% against current practice, for product teams working in the monorepo, without raising the change-fail rate?”
Pick the primary metric by the option the business case funds:
| Business-case option | Primary metric | Guardrails (must not get worse) |
|---|---|---|
| B. Seats only | Accepted changes per engineer-week | Change-fail rate, rework within 30 days, median review time |
| C. Shared harness | Cost per accepted change, all cost lines | Change-fail rate, rework, review time, incidents |
| D. Factory pilot on one loop | Share of loop tasks that pass the oracle and are accepted without rework | Escaped defects from the loop, spend per task, protected-path violations |
An accepted change is a change merged to the main branch and not reverted or rewritten within 30 days. Freeze that definition in the charter, and count per ticket where you can: agents make splitting one change into five pull requests cheap.
How to design the pilot, step by step
Section titled “How to design the pilot, step by step”-
Write the question, the primary metric and the guardrails. One primary metric only; everything else is a guardrail or an exploratory note.
-
Take the baseline from systems of record. Pull the last 8–12 weeks for every candidate team: accepted changes, lead time, change-fail rate, median review time, rework and incidents. Record the variance as well as the mean, because the sample size depends on it. The first prompt below computes PR open-to-merge time, a proxy for lead time, from
gh; if you define lead time as first commit to production, measure that instead and keep one definition throughout. -
Choose the design and the unit of assignment. For most organisations: matched concurrent cohorts at team level (compared in the next section).
-
Build the cohorts by matching, then assign by coin toss. Pair teams on the matching criteria, then randomly pick the treatment team in each pair. That removes the “we chose the best team” objection.
-
Fill in the confounder register. For each of the nine confounders below, write the control or why it does not apply.
-
Size the sample. Run the calculation below on your baseline variance. If the pilot cannot detect the effect the business case needs, change the design now, not after the readout.
-
Write and sign the decision rule. Scale, extend or stop, with thresholds. The CTO signs the guardrails, the CFO signs the cost definitions, and a named analyst outside the pilot team computes the result.
-
Instrument both cohorts before week one. Tag every agent run with its cohort (see the tabs below) and confirm the control cohort’s metrics reach the same dashboards.
-
Run for the planned weeks, then read out twice. Interim readout at pilot end; final readout 30 days later, when the rework window for the last changes has closed.
Which pilot design fits your organisation?
Section titled “Which pilot design fits your organisation?”| Design | How assignment works | Use it when | What it cannot do |
|---|---|---|---|
| Matched concurrent cohorts (default) | Teams paired on similar work; a coin toss per pair picks treatment | You have at least four comparable teams | Detect small effects with few teams; see sample size |
| Staggered rollout (stepped wedge) | Every team gets the treatment, at randomly ordered start dates 2–4 weeks apart | Withholding the tool is politically impossible | Finish fast |
| Task-level randomisation (METR-style) | Each ticket is randomly marked “agents allowed” or not before work starts | You want the strongest causal answer for one team | Survive parallel agents, which make time-on-task unreliable (METR) |
| Before and after on one team | No control; compare with the team’s own baseline | Only as a supplement to one of the above | Separate the tool from the season, the roadmap or a reorganisation |
Never compare volunteers with non-volunteers: enthusiasts differ in ways the tool did not cause, which is METR’s 2026 selection caveat.
For a factory pilot on one loop (option D), the unit is the task, not the team. Randomly route half of the loop’s incoming tasks (for example, dependency upgrades) to the unattended agent and half to the usual process, and compare acceptance, rework and cost per task.
How to choose matched cohorts
Section titled “How to choose matched cohorts”Match on what drives delivery metrics more than any tool does, at least the first four rows.
| Match on | Why it matters | How to check |
|---|---|---|
| Codebase and language | Coverage and build speed decide how much an agent can verify itself | Same repository or stack; coverage within a few points |
| Work type | Greenfield features, maintenance and incident work move differently | Share of ticket types over the baseline weeks |
| Team size and seniority mix | Seniors and juniors adopt agents differently | Headcount and a rough seniority split |
| Baseline metric level | Regression to the mean makes the worst team look improved anyway | Primary metric within the same quartile |
| Release calendar | A launch freeze or a quarter-end crunch swamps any tool effect | No known freeze or major launch inside the pilot window |
Which confounders break an AI coding pilot?
Section titled “Which confounders break an AI coding pilot?”A confounder is anything other than the intervention that changes the metric during the pilot. List each in the charter with its control.
| Confounder | How it shows up | Control |
|---|---|---|
| Volunteer selection | The keenest team gets the tool and was already the fastest | Matched pairs, coin toss inside each pair |
| Novelty and observation | Everyone works harder because they are being watched | Tell both cohorts they are measured; judge the last four weeks |
| The adoption dip | Throughput falls in the first weeks while people learn | Report weeks 1–2 separately; do not stop on throughput alone before week 4 |
| Contamination | Control engineers use personal AI accounts or copy the treatment team’s rules | Record the control’s current tools; in a shared repository, ship the treatment harness through managed or user-level settings (or a plugin) enabled only on treatment machines, not through files committed to the repo |
| Work-mix shift | Treatment engineers pick tickets that suit agents | Assign tickets through the normal backlog; record a size estimate before work starts, as METR did |
| Metric gaming | Pull requests get smaller and more numerous; “accepted” gets redefined | Count per ticket; freeze definitions in the charter; the analyst sits outside the team |
| Concurrent change | A CI migration, reorganisation or new on-call rota lands mid-pilot | Freeze other process changes in both cohorts, or log them with dates |
| Shared reviewers | Agent pull requests lengthen a shared review queue | Keep review inside cohorts, or measure review time per reviewer |
| Parallel agent runs | Per-task timing breaks when one engineer runs three agents | Measure per engineer-week and per accepted change |
How big does the pilot need to be?
Section titled “How big does the pilot need to be?”Most pilots are too small to detect the effect they claim. The standard two-group formula, two-sided at 5% significance and 80% power:
changes per cohort = 15.7 × σ² ÷ δ²
σ is the standard deviation of the metric per change and δ the smallest effect worth detecting. Changes from one engineer are correlated, so multiply by the design effect 1 + (m − 1) × ICC, where m is changes per engineer over the pilot and ICC the share of variance between engineers.
A worked example for PR open-to-merge time, the lead-time proxy, with illustrative inputs: replace them with the values the first prompt computes.
| Input (illustrative) | Value |
|---|---|
| σ of log open-to-merge time per change | 1.0 |
| Effect to detect: 30% shorter time | δ = |ln 0.7| ≈ 0.357 |
| Changes per cohort before clustering | 15.7 × 1.0 ÷ 0.127 ≈ 124 |
| Changes per engineer over 8 weeks (m) | 30 |
| ICC between engineers | 0.1 |
| Design effect | 1 + 29 × 0.1 = 3.9 |
| Changes per cohort after clustering | 124 × 3.9 ≈ 484, about 17 engineers per cohort |
For a 20% effect the same inputs need about 1,230 changes, or 41 engineers, per cohort. That is the arithmetic behind the most common pilot mistake: one team of eight cannot detect anything short of a large effect.
When the numbers do not fit, change the design:
- Add teams, not weeks. Because of clustering, more engineers help far more than more weeks from the same engineers.
- Adjust for the baseline. Compare each engineer with their own baseline; it removes stable differences between people and shrinks σ.
- Pick a higher-volume unit. A factory loop with hundreds of tasks a month reaches a useful size in four weeks.
- Decide on intervals and guardrails, not on a p-value alone: guardrails hold and the confidence interval clears a threshold.
Write the decision rule before week one
Section titled “Write the decision rule before week one”The decision rule turns the readout into an action nobody can renegotiate. Adjust the thresholds to your baseline and business case.
| Outcome | Condition | Action |
|---|---|---|
| Scale | Primary metric improves by at least the business-case threshold (for example, cost per accepted change 10% below control), and no guardrail was breached in two consecutive weeks | Extend to the next cohort of teams, not to everyone |
| Extend once | Primary metric improves but below the threshold, or the interval is too wide to decide, and guardrails held | Run four more weeks with the same cohorts, then apply the rule again with no further extension |
| Stop | Any stability or protected-path guardrail is breached twice, or the primary metric is flat or worse at the final readout | Stop the treatment and record what the pilot learned about your verification gaps |
Write the rule into the charter verbatim; any leadership override goes in too, with its reason. The board report cites that record.
How to tag pilot runs in Claude Code, Codex and Cursor
Section titled “How to tag pilot runs in Claude Code, Codex and Cursor”Quality metrics come from git, CI and the incident system, which are tool-neutral: the analyst joins them to the cohort list. What differs by tool is how you get cost and usage per cohort, which option C needs.
Claude Code exports OpenTelemetry metrics, including claude_code.cost.usage (USD), claude_code.token.usage, claude_code.pull_request.count and claude_code.commit.count. OTEL_RESOURCE_ATTRIBUTES adds your own keys to every metric and event, so a cohort label becomes a filter in your dashboards. Set it in the managed settings file on the treatment engineers’ machines:
{ "env": { "CLAUDE_CODE_ENABLE_TELEMETRY": "1", "OTEL_METRICS_EXPORTER": "otlp", "OTEL_LOGS_EXPORTER": "otlp", "OTEL_EXPORTER_OTLP_PROTOCOL": "grpc", "OTEL_EXPORTER_OTLP_ENDPOINT": "http://collector.example.com:4317", "OTEL_RESOURCE_ATTRIBUTES": "pilot.id=harness-2026q4,pilot.cohort=treatment,team.id=payments" }}Managed settings apply to every user the file is deployed to, so deploy this file only to treatment machines (or put the block in each engineer’s ~/.claude/settings.json). Give control engineers the same block with pilot.cohort=control, so both cohorts reach the same dashboards.
Two traps, checked against Claude Code 2.1.283. Claude Code ignores the exporter variables in a repository’s .claude/settings.json: use managed settings or each engineer’s ~/.claude/settings.json. And OTEL_RESOURCE_ATTRIBUTES values cannot contain spaces (team.id=payments_core, not quoted text).
On Team and Enterprise plans with the GitHub app installed, the analytics dashboard labels merged pull requests with Claude Code contributions claude-code-assisted. Contribution metrics are in public beta and unavailable with zero data retention: a cross-check, not the cohort definition.
Codex reads an [otel] table from ~/.codex/config.toml. span_attributes adds the cohort label to every exported trace span, and only to spans: it does not touch metrics_exporter output, so it gives per-cohort traces, not per-cohort cost. Field names were read from the OtelConfigToml source at the rust-v0.157.1 tag:
[otel]environment = "prod"log_user_prompt = falsespan_attributes = { "pilot.id" = "harness-2026q4", "pilot.cohort" = "treatment", "team.id" = "payments" }trace_exporter = { otlp-http = { endpoint = "https://otel.example.com/v1/traces", protocol = "binary" } }For per-cohort token usage, record it from codex exec --json, where each turn.completed event carries a usage object, or join usage by user to the cohort list. The task edits files, so the run needs a writable sandbox; this pilot standardises on the permission profile -c default_permissions=":workspace" (beta in Codex 0.157.1; the legacy --sandbox workspace-write also works, but do not combine the two):
# Terminal or CI: one pilot task, usage appended with its task id and cohortcodex exec --json -c default_permissions=":workspace" "Fix the failing test in src/billing and stop when npm test passes" \ | jq -c 'select(.type == "turn.completed") | {task: "BIL-142", cohort: "treatment", usage: .usage}' >> pilot-usage.jsonlCursor’s team analytics and admin APIs could not be checked on 2026-09-26 (cursor.com was unreachable from our environment). Before week one, confirm that you can export usage and spend per user per day, and join it to the cohort list. If you cannot, run the Cursor cohort on a dedicated team account so the account total is the cohort total.
Telemetry setup is covered in depth on the agent telemetry page.
Prompts and templates to run the pilot
Section titled “Prompts and templates to run the pilot”Fill in every field of this one-page charter before week one.
# Pilot charter: <intervention> — <pilot id>
Question: Does <intervention> change <primary metric> by at least <threshold>against current practice, for <population>, without worsening <guardrails>?Business case: <link> Decision owner: <name> Analyst (outside the team): <name>
## DesignDesign: matched concurrent cohorts | staggered rollout | task randomisationUnit of assignment: team | engineer | taskPairs and coin-toss result: <team A vs team B → A treatment>, ...Pilot weeks: <start> to <end> (4–8 weeks) Final readout: <end + 30 days>
## Metrics (definitions frozen at signature)Primary: <metric, exact definition, source system>Guardrails: change-fail rate | rework within 30 days | median review time | incidentsAccepted change: merged to main and not reverted or rewritten within 30 days
## Baseline (last 8–12 weeks, per cohort)Primary: mean ___ sd ___ Guardrails: ___
## Sample sizesigma ___ delta ___ ICC ___ m ___ → engineers per cohort ___
## Confounder register| Confounder | Control | Owner |
## Decision rule (verbatim)Scale if ___. Extend once if ___. Stop if ___.
## SignaturesCTO (guardrails) ___ CFO (cost definitions) ___ Analyst ___ Date ___How the pilot result is verified
Section titled “How the pilot result is verified”Nobody on the pilot team grades their own pilot.
- The baseline comes from git, CI and incident data before the coin toss; the CFO signs it with the cost definitions.
- Every treatment change carries its evidence: tests passed, review record and the cost of the run. The evidence bundle format lets the analyst sample changes instead of trusting a dashboard.
- The analyst outside the team computes the metrics from the frozen definitions, reports the interval, and applies the decision rule.
- Protected paths (auth, payments, schema, migrations) are enforced in CI in both cohorts.
- The final readout waits for the 30-day rework window.
What goes wrong in AI coding pilots, and how to recover
Section titled “What goes wrong in AI coding pilots, and how to recover”The pilot team was the volunteers. The result will be dismissed as selection, rightly. Keep them as the first treatment team, find a matched control, and run a second, randomly assigned staggered wave.
The readout was a survey. METR’s developers felt faster while measuring slower. Rebuild the comparison from git and CI data for the same weeks.
Pull requests per engineer rose, and the pilot declared success. Recompute per ticket with rework and incidents. If the guardrails moved, that is the Faros pattern, not a win.
The control team adopted the treatment’s rules halfway through. Date the leak, analyse the weeks before it, and treat the rest as a staggered rollout.
Throughput fell in weeks two and three, and a sponsor wants to stop. That is the planned adoption dip. Stopping before week 4 on throughput alone is outside the rule; a guardrail breach is inside it.
The result is inside the noise. Apply the “extend once” branch, add teams rather than weeks, and do not extend twice.
The dashboards cannot separate cohorts. Move the treatment cohort onto the tagged telemetry above and use only the tagged weeks for the cost metric.
Where to go next with the pilot
Section titled “Where to go next with the pilot”After a “scale” decision: the team adoption roadmap per repository, the transformation roadmap across the organisation. Dated evidence: the state of agentic engineering.