Bringing skeptics and senior engineers along
Bringing skeptics and senior engineers along to agentic engineering works when their objections are treated as a list of verification gaps, not as resistance. Each senior owns the design of the checks that decide when agent output is trustworthy, then runs a measured pilot on tasks from their own backlog, with a decision rule written before the first task starts.
Your staff engineer said it in the retro: “I tried it for a week, it wrote plausible code that was wrong in ways I only found because I know this system, and fixing it took longer than writing it.” Two other seniors have quietly stopped opening the agent at all. Your CTO wants an adoption number by the end of the quarter, and the rollout guide you were handed says to pair the holdouts with a champion and publish productivity charts until they come round.
That playbook treats the people who know best what can go wrong as the obstacle. This page is for the tech lead who has to get those people on board without losing their judgment, and ideally by putting that judgment to work.
What you get from this page
Section titled “What you get from this page”- An objection-to-check table that turns the eight objections you will hear into a verification gap and a check a senior can own.
- A four-week pilot protocol a senior runs on their own tasks, with an estimate first, random assignment, and a decision rule fixed in advance.
- A pilot log template to paste into the repository, plus a prompt that applies the decision rule to it without editorialising.
- Four copy-paste prompts for Claude Code, Codex, and Cursor: turn an objection into a check, gate a pilot task, attack the agent’s change, and analyse the pilot.
- An honest reading of the METR results you can put in front of the skeptic, including the part that supports them.
- The failure modes of skeptic conversions, from rigged pilots to seniors buried in review, and how to recover from each.
Why is senior skepticism worth listening to?
Section titled “Why is senior skepticism worth listening to?”Because the best evidence available agrees with part of it. Senior skepticism usually rests on two observations: the agent’s code looks right more often than it is right, and nobody has measured whether the agent saves an experienced engineer time on familiar code. The research supports the second outright and is consistent with the first.
What METR measured, read honestly
Section titled “What METR measured, read honestly”METR’s randomised controlled trials are the only trials in this site’s evidence base that measure experienced developers on their own repositories. Put these results in front of your skeptic yourself, before they find them:
| Study | What it found | What it means for your team |
|---|---|---|
| METR, July 2025: 16 experienced open-source developers, 246 issues in repositories they knew well | “The use of AI causes tasks to take 19% longer, with a confidence interval between +2% and +39%.” Developers expected a 24% speedup and afterwards “still believed AI had sped them up by 20%.” | A statistically significant slowdown for experts on familiar code, on early-2025 tooling. Self-reported speed is not evidence, including your own. |
| METR, February 2026 update | The point estimates now favour AI: about 18% less time for returning developers (interval from 38% less to 9% more) and about 4% less for new recruits (15% less to 9% more). Both intervals cross zero. | Not a measured speedup. METR calls this data “an unreliable signal of the current productivity effect of AI tools”, partly because developers declined to work without AI and time is hard to measure when one person runs several agents at once. |
| METR’s analysis repository, checked 2026-09-26 | “Please interpret these results with a major grain of salt.” | Neither METR study shows that agents make experienced engineers faster on their own code; no trial in this site’s evidence base does either. |
The METR design is also the template for your pilot. Its analysis code compares the time each task actually took against the developer’s estimate of the time it would take without AI, recorded before the work started. Your senior can run the same comparison on 10 of their own tasks.
What the rest of the evidence adds
Section titled “What the rest of the evidence adds”Seniors are not alone in their distrust. In Sonar’s developer survey (January 2026, more than 1,100 developers), “96% of developers do not fully trust AI-generated code, and only 48% always verify it before committing.” The 2025 DORA report (Google Cloud, 23 September 2025) found a positive relationship between AI adoption and delivery throughput, but a continued “negative relationship with software delivery stability”, and it names the cause: without “strong automated testing, mature version control practices, and fast feedback loops, an increase in change volume leads to instability.”
That last finding is where seniors come in. The controls DORA names are exactly what senior engineers build and maintain. Anthropic’s study of its own engineers (December 2025) names the risk on the other side: the “paradox of supervision”, in which overseeing an agent needs the very skills that over-delegation erodes. Your seniors hold those skills. The job is to point them at verification, not to talk them out of their judgment.
How do you turn each objection into a check?
Section titled “How do you turn each objection into a check?”Every serious objection names a failure the agent can cause and your current checks cannot catch. Write the objection down in the senior’s own words, then ask which check would have caught the failure behind it. The senior who raised it owns that check.
| What you hear | The verification gap behind it | The check the senior owns | Where it is built |
|---|---|---|---|
| “It writes plausible code that is subtly wrong.” | The tests pass on wrong code: the oracle is weak. | Acceptance criteria written before the task, plus mutation testing on changed files. | Oracle strength |
| “It changes the tests until they pass.” | The agent can edit its own judge. | Protected test paths in CODEOWNERS, and a CI check that fails when a pull request modifies or deletes existing test, fixture or snapshot files without approval from the CODEOWNERS owner (new test files are allowed). | Protecting the oracle |
| “I’m faster without it.” | Nobody has measured it, and METR’s 2025 result says they may be right. | The four-week pilot on their own tasks, below. | This page |
| “Reviewing its pull requests takes longer than writing the code.” | Review cost is uncapped: large diffs, no evidence. | A pull request size budget and an evidence bundle the agent must produce. | Evidence bundle, review queue |
| “It ignores our architecture.” | The conventions live in people’s heads, not in a machine check. | Two or three fitness functions for the rules violated most, plus the shared rules file. | Fitness functions, shared agent rules |
| “It will rot the codebase.” | Nobody tracks duplication, churn, or complexity over time. | A monthly codebase-health report with a threshold that stops agent work in a module. | Codebase health |
| “I’ll be accountable for code I never read.” | No rule says which changes a human must still read. | The escalation classes (auth, money, schema, migrations, the oracle itself), routed through CODEOWNERS. | Reading evidence instead of code |
| “Juniors will stop learning.” | The team has no plan for skills the agent now exercises. | A written list of what juniors still do by hand, and review-to-learn sessions. | Growing junior developers |
Two objections are not on the table because they are not verification gaps: “it’s hype” and “I don’t want to”. The first dissolves when you stop making claims the evidence does not support. For the second, ask what they would need to see. The answer is almost always one of the eight rows.
Give seniors the verification design, not an adoption target
Section titled “Give seniors the verification design, not an adoption target”The old rollout advice makes enthusiasts the champions and measures seniors on usage. Reverse it. The senior who distrusts the agent most is the right person to decide what “verified” means for the part of the system they know best.
In practice, each skeptical senior gets one loop (a repeatable kind of change in one area, such as “API endpoint changes in the orders service”) and owns four artifacts for it:
- The acceptance criteria template the agent’s spec must fill before work starts.
- The escalation list for that area: the paths where a human reads the code every time.
- The protected oracle: which tests, fixtures, and CI steps the agent may not edit.
- The go/no-go call on whether that loop moves to evidence-based review. See helping the team stop reading every diff for the stages.
What you, as tech lead, stop doing matters as much. Remove every usage target: seat activations, sessions per week, and share of agent-authored lines all reward the wrong behaviour and are trivially gamed. Replace them with the outcome measures in the pilot log below. The metrics frameworks page defines the team-level versions.
Run a measured pilot on the senior’s own tasks
Section titled “Run a measured pilot on the senior’s own tasks”A demo on a task you picked proves nothing to a skeptic, because you picked it. A pilot on tasks from their own backlog, assigned at random, with the rule for success written before the first task starts, is an experiment they can respect. It will not be statistically significant with 10 tasks. It does not need to be; it needs to produce a decision both of you agreed to in advance.
-
Choose 10 to 12 tasks from the senior’s own backlog for the next four weeks. They choose, not you. Exclude anything in the escalation classes. Mix sizes, but keep each task under two days.
-
Record an estimate for every task before assigning it. The senior writes down how many hours each task would take by hand. This is the METR baseline, and it has to exist before anyone knows which arm the task falls into.
-
Assign each task to “agent” or “by hand” at random. Flip a coin or run
shufover the list. Random assignment is what stops the pilot from quietly giving the agent the easy tasks. -
Write the decision rule and sign it. Put it at the top of the pilot log before task one. A rule that works for most teams: adopt the agent for this kind of task if the median ratio of actual to estimated time in the agent arm is no worse than in the by-hand arm, and no agent task needs rework or causes a defect within 14 days of merge. Add one stop rule: pause if any agent change reaches production with a defect in an escalation class.
-
Have the senior write the gate before each agent task. Acceptance checks come first, in their words, using the gate prompt below. The agent implements only after they approve the checks. For “actual time”, count wall-clock hours from start to merge, including review and fixes: that is the time the objection was about.
-
Log every task in both arms the same day it merges. Estimate, actual hours, whether the gate passed first time, and a 14-day rework field you fill in later.
-
Let the senior present the result. Apply the decision rule with the analysis prompt, then the senior presents to the team. There are three legitimate outcomes: adopt for this kind of task, adopt with a stronger gate, or keep it manual and re-run the pilot next quarter.
For a pilot that has to convince an organisation rather than one engineer, with concurrent cohorts, confounders, and sample sizes, use designing a pilot that proves something. This page’s pilot is the smallest version that still earns a senior’s trust.
Paste this pilot log into the repository
Section titled “Paste this pilot log into the repository”Keep the log in the repository (for example docs/pilots/anna-orders.md) so the result is reviewable and survives the conversation about it.
# Pilot: Anna, orders-service API changes (2026-09-28 to 2026-10-23)
Decision rule (signed 2026-09-26, Anna + Marek):Adopt the agent for orders-service API changes if the median actual/estimateratio in the agent arm is <= the by-hand arm AND no agent task needs rework orcauses a defect within 14 days of merge. Stop rule: pause if an agent changereaches production with a defect in auth, billing, schema or migrations.
| # | Task | Estimate (h) | Arm | Actual (h, start to merge) | Ratio | Gate passed first time | Rework or defect within 14 days | Note || -- | ---------------------------- | ------------ | ------- | -------------------------- | ----- | ---------------------- | ------------------------------- | ---------------------------- || 1 | Cursor pagination on /orders | 6 | agent | 4.5 | 0.75 | yes | no | wrote 7 acceptance checks || 2 | Remove legacy v1 serializer | 3 | by hand | 3.5 | 1.17 | n/a | no | || 3 | Retry policy for outbox | 8 | agent | 11 | 1.38 | no | yes: race in retry counter | concurrency: add to escalation list? |The last column is where the pilot pays for itself even if the result is negative. Row 3 above is the pilot doing its job: it found a class of change, concurrency, where the senior’s skepticism was right, and that class goes on the escalation list.
How does each tool run the agent arm of the pilot?
Section titled “How does each tool run the agent arm of the pilot?”The gate and the log are the same in every tool. What differs is how the senior keeps the agent in planning until the checks are approved, isolates the work, and gets an independent review before merge.
Start the session in Plan mode in its own worktree, so nothing is written until the senior approves the acceptance checks and the plan:
# Terminal, from the repository root (Claude Code 2.1.283)claude --permission-mode plan -w pilot-task-1After implementation, run /code-review in the session for a second opinion on the local diff, and /usage to note what the session consumed. If the team wants the gate enforced rather than remembered, a hook can block edits to the protected test paths; see shared hooks governance.
Run the task in a managed worktree and switch to Plan mode with /plan before the first message, so Codex proposes the checks and the plan first:
# Terminal, from the repository root (Codex CLI 0.157.1)codex --worktreeBefore opening the pull request, get a non-interactive review. In Codex CLI 0.157.1, codex review accepts either a target flag (--base, --uncommitted, --commit) or custom instructions, not both, so pick one of these:
# Terminal (Codex CLI 0.157.1)# Option 1: a standard review of the branch against main, no custom briefcodex review --base main
# Option 2: the senior's objection as the brief, with no target flagcodex review "Review the changes on this branch against main. Look for the failure I expect in this area: state mutated across retries without a lock."For the full hostile review, paste the “attack the agent’s change” prompt below into the session instead.
Switch the agent to Plan Mode before the first message, so it “creates detailed implementation plans before writing any code”, and run the task in a worktree (“Worktrees let Agent work in isolated Git checkouts”) so it stays isolated from the senior’s other work. Both quotes are from Cursor’s Plan Mode and worktrees documentation, checked on 2026-08-28. Where the mode picker and the worktree option sit in the Agents Window are covered in Cursor agent modes and the Agents Window.
On the pull request, Bugbot “reviews pull requests and identifies bugs, security issues, and code quality problems”. Treat its findings as a reviewer’s comments that the senior triages, not as the gate: the gate is the senior’s acceptance checks. See Bugbot in Cursor.
Copy-paste prompts for skeptics and seniors
Section titled “Copy-paste prompts for skeptics and seniors”The prompts work unchanged in Claude Code, Codex, and Cursor. The first two are the senior’s; the last two serve the pilot as a whole.
What should you say in the one-to-one?
Section titled “What should you say in the one-to-one?”The conversation decides whether the senior runs the pilot in good faith. Adopt these substitutions as-is:
| Instead of saying | Say |
|---|---|
| “Everyone else is using it, you’ll fall behind.” | “You know where this system breaks. I want you to decide what counts as verified here.” |
| “Studies show it makes people much faster.” | “METR’s trial of experienced developers found a slowdown in 2025, and its 2026 data is too noisy to say. Let’s measure it on your tasks.” |
| “It’s a prompting issue; pair with Kasia for a week.” | “Which failure did you see? Let’s write the check that would have caught it.” |
| “Try it on something small first.” | “Pick 10 tasks from your own backlog. We’ll flip a coin for which ones the agent gets.” |
| “Usage is part of your goals this quarter.” | “Your goal is the gate for the orders loop and a signed pilot result, whatever it says.” |
How do you verify the seniors’ checks actually work?
Section titled “How do you verify the seniors’ checks actually work?”A check a senior writes is only useful if it fails when the code is wrong. Hold every check from the objection table to the same test you would hold the agent’s code to:
- It fails on a known-bad example. The “turn my objection into a check” prompt requires a failing run before the check is accepted. Keep that output in the pull request that adds the check.
- It runs in CI and blocks the merge. A check that only runs on the senior’s machine is a habit, not a gate. Required status checks in branch protection make it binding.
- The agent cannot edit it. The check’s files are in
CODEOWNERSwith the senior as owner, so any change to them needs the senior’s approval. - The pilot log shows it working. Over the four weeks, the “gate passed first time” and “rework within 14 days” columns show whether the gate catches problems before merge or after.
The senior signs off on the loop’s move to evidence-based review. You sign off on the decision rule and on accepting its outcome, including a negative one.
When bringing skeptics along goes wrong
Section titled “When bringing skeptics along goes wrong”The pilot tasks were easy, and everyone knows it. A pilot on typo fixes and README updates converts nobody and misleads everyone else. Recovery: re-run it with tasks the senior chose from their own backlog and assigned at random, and say publicly why the first run did not count.
The senior ran the pilot to prove it fails. Signs: agent-arm tasks with no gate written, abandoned sessions, or estimates padded in the by-hand arm. Recovery: do not argue about motive. Add a second reviewer who reads the log weekly, and require the gate prompt’s output for every agent task before it counts.
The result was negative, and you overrode it. You lose this senior and everyone watching. Recovery: honour the outcome for that kind of task, record the reason and a re-run date one quarter out, and re-run it when the models or the gate change.
The seniors became everyone’s reviewers. Agent pull requests route to the people who distrust them most, and the seniors spend their weeks reading diffs. Recovery: cap review load per person, require the evidence bundle before a senior is assigned, and use the review queue limits.
A senior’s objection was right, and nobody wrote it down. The concurrency failure in the example log is valuable only if it changes the system. Recovery: every “rework or defect” note becomes either a new check or a new escalation class within a week, owned by the senior who found it.
Usage targets came back through the side door. A dashboard of sessions per engineer reappears in a leadership review, and seniors start logging sessions to be left alone. Recovery: take the dashboard out of performance conversations and report outcome measures only, per loop.
Where to go next with skeptics and seniors
Section titled “Where to go next with skeptics and seniors”Frequently asked questions
How do you convince a senior engineer who is skeptical of coding agents?
You do not argue or demo. You translate each objection into a verification gap, give the senior ownership of the check that closes it, and let them run a measured pilot on tasks from their own backlog with a decision rule written before the first task starts.
Does the METR study prove AI slows experienced developers down?
It proved it for one setting: in METR's 2025 trial, 16 experienced open-source developers took 19% longer with AI, a statistically significant result. The 2026 follow-up's point estimates favour AI but the intervals cross zero, and METR calls that data an unreliable signal.
What if the pilot shows the agent does not help a senior engineer?
Accept the result for that task class, keep it manual, record the reason and a re-run date, and move on. A pilot that can only end in adoption is a demo, and skeptics can tell the difference.