Skip to content

Why AI Coding Tools? From Autocomplete to Agentic Engineering

AI coding tools matter because they stopped suggesting lines and started completing tasks: Claude Code, Codex, and Cursor read a repository, plan a change, edit files, run tests, and iterate until a check passes. The constraint moves from writing code to verifying it. Agentic engineering is the discipline built on that shift: specifying behavior and proving it with checks.

Your team rolled out an agent three months ago. Pull requests per week went up, and so did the review queue: five days to first review, a senior engineer who reads every diff until 9 p.m., and a production incident traced to a test the agent rewrote to go green. Nobody is typing less carefully. The system around the typing has not changed yet, and that system is what this page is about.

Developers get the loop to run today, tech leads get the bottleneck to fix, and CTOs and executives get the decision the tools force on them. If you have not picked your path yet, start on Start here.

What you’ll walk away with from this agentic engineering primer

Section titled “What you’ll walk away with from this agentic engineering primer”
  • A one-screen answer to “what actually changed”, mapped to the autonomy ladder your team is already on
  • A decision table for which work to hand to an agent first, based on whether a machine can say “done”
  • Your first verified agent loop in Claude Code, Codex, or Cursor, with three copy-paste prompts
  • The four artifacts you read instead of the diff, and the short list of changes a human still reads line by line
  • The failure modes teams hit in the first quarter, each with its recovery
  • The next page for your role

What changed when AI coding tools became agents?

Section titled “What changed when AI coding tools became agents?”

Autocomplete predicts the next few tokens inside a function you are already writing. An agent takes a whole task, reads the code it needs, runs your commands, and reports back when a stop condition is met. Andrej Karpathy put the change in one line in his Sequoia Ascent summary (30 April 2026): “The unit of programming changed from typing lines of code to delegating larger ‘macro actions’…”

Dan Shapiro’s autonomy ladder (January 2026) turns that into levels you can place a workflow on. The level is decided by who reads the output, not by which tool you bought:

LevelWhat the agent doesWhat the human readsThe tools in this mode
L0 By handSuggests the next few tokens (spicy autocomplete)Every character, before it reaches diskInline completion in any editor
L1 AssistedTakes small, discrete tasks you hand itAll of itAn agent chat for one-off tasks
L2 PairedWrites while you steer liveEvery line, as it is writtenAn agent chat beside your editor
L3 Review managerRuns a task unattendedEvery diffClaude Code, Codex, or Cursor with a task and a test command
L4 Spec managerImplements from a specTests, evidence, and outcomesThe same tools plus acceptance criteria and CI gates
L5 Dark factorySpecs in, releases outNo human reads the codeA pipeline of agents behind a strong oracle

Shapiro places most AI-native developers at Level 2 and calls Level 3 the level almost everyone tops out at. The reason is arithmetic: the agent writes faster than people can read. The ladder, its self-test, and what each level requires live on the autonomy ladder. How the ladder, the lifecycle stages, and the factory stations fit together is on one map.

Why the bottleneck moves from writing code to verifying it

Section titled “Why the bottleneck moves from writing code to verifying it”

The evidence does not show agents making every developer faster. It shows generation scaling while verification does not. Three sources, each dated:

  • DORA 2025 report (Google Cloud blog, 23 September 2025): “AI doesn’t fix a team; it amplifies what’s already there.”
  • Faros AI, The Acceleration Whiplash (April 2026; vendor telemetry from 22,000 developers on Faros’s self-selected customer base): epics completed per developer rose 66.2%, while incidents per pull request rose 242.7% and median time in review rose 441.5%.
  • METR: its 2025 randomized trial found experienced open-source developers took 19% longer with AI tools (10 July 2025). Its February 2026 update has point estimates in AI’s favor, but both intervals cross zero and METR calls that data “an unreliable signal”.

Read together: output goes up, and whether that becomes value depends on the checks around it. The full evidence, with counter-evidence and caveats, is in the state of agentic engineering.

What agentic engineering asks of each role

Section titled “What agentic engineering asks of each role”

The tools are the same for everyone. The work they create is different for each reader.

You areThe problem you now ownWhat to adopt firstYour next page
DeveloperHanding over tasks you can check without re-reading themAn instruction file with runnable commands, and acceptance criteria on every taskDeveloper track
Tech leadA review queue that grows faster than the teamEvidence-based review and a short list of must-read change classesTech lead track
CTO / VP EngineeringWhich loops may run at which level, and how to prove itA risk-class policy per loop, and outcome metrics instead of ”% of code”CTO track
ExecutiveCost, speed, and risk claims you cannot yet checkThree questions for your CTO (below) and a dated evidence baseExecutive track

Which work should you hand to an agent first?

Section titled “Which work should you hand to an agent first?”

Hand over work where a machine can decide “done”. The question is not how hard the task is. It is how strong the oracle is: the tests, types, linters, and runtime checks that fail when the result is wrong.

WorkOracle you usually haveWhere to start itWhy
Dependency bumps, lint and type fixesCompiler, type checker, existing testsL4: agent runs, you read the checksThe check is deterministic and existed before the agent
Tests for untested code (characterization tests)The current behavior itselfL3: read the tests, not the codeNew tests are the oracle, so a human checks them
Features behind acceptance criteriaContract and end-to-end tests you write firstL3 moving to L4Criteria become checks; see acceptance criteria
Codebase comprehension and docsA human’s judgment of the answerL2Useful at once; nothing to merge, so low risk
Auth, payments, schema, data migrationsTests cover allowed paths, not the missing oneL2 or L3, a named person reads the codeA mistake is expensive and slow to surface
Novel architecture and product directionNoneHuman-led, agent as a sparring partnerThere is nothing to check against yet

A level belongs to a loop, not to a company. Your dependency-bump loop can run at Level 4 while feature work beside it stays at Level 2.

The loop is the same in every tool: write down how the repository is built and tested, give the agent a task with acceptance criteria, let it run the checks, and read the evidence it returns. The commands differ.

  1. Create the instruction file. The agent needs your real build, test, and lint commands, or it guesses. See what goes in CLAUDE.md and AGENTS.md.
  2. Plan before editing. Ask for a plan in plan mode and correct it. A wrong plan costs one message; a wrong diff costs a review.
  3. Run the task with acceptance criteria and a stop condition. “Done” is “these checks pass”, not “the agent says so”.
  4. Read the evidence, then the diff only where the evidence is weak. The section after the prompts shows what to read.

Run these in the repository root (Claude Code 2.1.283, checked 2026-09-26):

Terminal window
claude # start a session; type /init to generate CLAUDE.md
claude --permission-mode plan # start in plan mode; Shift+Tab also cycles modes
claude --worktree # run the task in its own git worktree

On the latest channel (v2.1.283) an interactive session starts in auto mode on supported models, so a classifier approves actions; use --permission-mode plan or Shift+Tab to plan first.

Claude Code reads CLAUDE.md. On the latest release channel (v2.1.277 and later) it also reads AGENTS.md when a project has no CLAUDE.md (v2.1.281 on Bedrock, Google Cloud, Microsoft Foundry and LLM gateways). Inside the session, /code-review reviews the current changes before you open a pull request.

How do you know the agent’s output is right without reading every line?

Section titled “How do you know the agent’s output is right without reading every line?”

You read evidence first and code only where evidence cannot decide. Reading evidence instead of code sets out four artifacts, in this order:

  1. Spec delta: which behavior changed, in sentences, compared with what you asked for.
  2. Acceptance results: each criterion mapped to a check that ran and passed.
  3. Oracle strength: whether the pull request touched tests, snapshots, CI, or lint config, and whether the tests would fail if the code were wrong. See how strong your oracle is.
  4. Runtime signals: a preview or staging run of each acceptance path, then canary metrics after deploy.

Some changes still get a named human reader every time, whatever the evidence says: authentication and authorization, money, schema, data migrations, changes to the oracle itself, and anything no check covers. Route those through CODEOWNERS so the forge enforces the reader. The developer who ran the agent owns the evidence; the code owner signs off on the must-read classes; the organization writes the list down once in its autonomy and risk-class policy. The template that makes CI reject an incomplete pull request is the evidence bundle.

All three run the loop above. They differ in where the work happens:

  • Claude Code is terminal-first, with the same agent in a desktop app, VS Code, JetBrains, the web, and headless runs (claude -p) for CI.
  • Codex spans a CLI, an IDE extension, a desktop app, and Codex cloud, with GitHub review through @codex review.
  • Cursor is editor-first, with Cloud Agents, Automations, and Bugbot for work that runs outside your editor.

You can run two: one for interactive work and one in CI. The tool matters less than the loop around it. Model choices live on the models hub, and the side-by-side comparison is on comparing AI coding tools.

Will agents replace developers, and is the code safe to ship?

Section titled “Will agents replace developers, and is the code safe to ship?”

Replacement. The work moves from typing to specifying, verifying, and owning outcomes. What that means for skills and careers is on your craft and career when agents write the code, and what stays human is on the human’s job.

Code quality. Agent code is as trustworthy as the checks it passed, no more. Treat “the agent says it is done” as a claim and the evidence bundle as the proof.

Security and IP. An agent runs commands with your credentials. Start with sandboxing and approval settings in permissions and sandboxing, then threat-model the whole setup with the agent threat model. Code and data retention, training use, and IP questions are covered in privacy and data handling.

Cost. Plan prices change often and live on pricing analysis. Whether the spend pays back depends on which loops you can move up the ladder, which the economics page turns into a model.

What goes wrong when teams adopt AI coding tools?

Section titled “What goes wrong when teams adopt AI coding tools?”

Pick your role track on Start here, or go straight to the page that fits your next step.