Your craft and career when agents write the code
When coding agents write most of the code, the compounding skills are specification, verification design, architecture, and taste. Syntax recall and boilerplate can be let go; unassisted debugging erodes most visibly (the largest gap in Anthropic’s January 2026 trial). Anthropic’s study of its own engineers names the trap the “paradox of supervision”: supervising an agent takes the same coding skills that heavy delegation wears away.
You merged 30 pull requests last month and typed almost none of them. Your performance review still asks what you built, and an incident just landed in a module no human has read in full since March. This page is for developers who want an honest answer to “what is my craft now”, and for tech leads who have to answer it for other people.
What this page gives you for your craft and career
Section titled “What this page gives you for your craft and career”- The skills that compound and the skills that erode, with a source for each claim.
- The “paradox of supervision”, and the habits the one controlled study links to keeping what you learn.
- Setup for a learning mode in Claude Code, Codex, and Cursor.
- A portfolio plan for each level of the autonomy ladder: what to keep sharp, what to hand over, and what evidence to collect.
- A one-page skills plan you can copy, and four prompts that turn your own git history into that evidence.
- Self-checks that show whether your skills are holding, and a list of what no source can tell you yet.
What the evidence says about skills when agents write the code
Section titled “What the evidence says about skills when agents write the code”Two studies speak directly to the question. Both come from Anthropic, a model vendor, so read them as a vendor’s research about its own product and staff. No independent study of comparable design was available when this page was written.
The internal workplace study. In August 2025 Anthropic surveyed 132 of its engineers and researchers and ran 53 in-depth interviews. The results were published on 2 December 2025 by Saffron Huang and colleagues. Most respondents said they could “fully delegate” only 0–20% of their work. One engineer described their work as having shifted “70%+ to being a code reviewer/reviser rather than a net-new code writer.” Some worried about “skills atrophying as [they] delegate more” and about losing the “collateral” learning that happens during manual problem-solving. The report names the risk directly:
“effectively using Claude requires supervision, and supervising Claude requires the very coding skills that may atrophy from AI overuse.”
A senior engineer in the same study explained why seniority protects them: “I’m primarily using AI in cases where I know what the answer should be or should look like. I developed that ability by doing SWE ‘the hard way’.” A junior engineer, in their view, would need “a lot of deliberate effort to continue growing my own abilities rather than blindly accepting the model output.”
The skill-formation trial. On 29 January 2026 Judy Hanwen Shen and Alex Tamkin published a randomized study of 52 mostly junior engineers. Each engineer learned Trio, a Python library for asynchronous programming, by building two features. The group that used an AI assistant averaged 50% on the follow-up quiz. The group that coded by hand averaged 67%. The AI group finished about two minutes faster, but that difference was not statistically significant. The biggest gap was on debugging questions. The authors note the limit themselves: the quiz measured comprehension shortly after the task, and the study “does not resolve” whether that predicts long-term skill.
Participants who handed over all the code writing or the debugging scored lowest. Those who asked conceptual questions, or asked follow-up questions about generated code, scored highest. The finding is not “AI makes you worse”. It is that how you delegate decides what you keep.
Which skills compound when agents write the code?
Section titled “Which skills compound when agents write the code?”These four skills get more valuable as agents write more of the code. Each one either decides what the agent builds or decides whether its output can be trusted. The human’s job covers them as responsibilities; this section covers them as skills you practise.
| Skill | What it looks like in practice | Why it compounds (source) | How to practise it this month |
|---|---|---|---|
| Specification | Turning intent into acceptance criteria and a stop condition an agent cannot misread | Karpathy: “You are in charge of the spec and plan” (Sequoia Ascent summary, 30 April 2026). Shapiro’s Level 4 developer “write[s] a spec” and checks whether the tests pass (January 2026) | Write the acceptance criteria before every agent task, then count how often the result matched them on the first run |
| Verification design | Choosing the oracle that decides “done” and keeping it out of the agent’s reach | Sonar, January 2026: “96% of developers do not fully trust AI-generated code, and only 48% always verify it before committing.” That gap is the skill | For each feature, write the one check the agent cannot edit, and read how strong your oracle is |
| Architecture | Boundaries, dependency direction, and what may change without a human decision | DORA 2025: “Teams working in loosely coupled architectures with fast feedback loops see gains, while those constrained by tightly coupled systems and slow processes see little or no benefit” | Record one architecture decision a week, as a file the agent reads |
| Taste | Seeing that a working implementation is the wrong one | Karpathy: “You still have to be in charge of aesthetics, judgment, taste, and oversight.” SlopCodeBench (March 2026) measured “structural erosion” as agents extend their own solutions, decay that a passing test suite does not report | Weekly, reject one green agent result on design grounds and write down why |
A fifth skill sits under all four: a working model of the system in your head. An engineer in Anthropic’s study described building it by reading “docs and code that isn’t directly useful for solving your problem”. It is the first thing heavy delegation removes.
Which skills erode, and which of them can you let go?
Section titled “Which skills erode, and which of them can you let go?”Not everything that erodes is a loss. The honest split is between skills the agent now does better and more cheaply, and skills that only look like typing.
| Skill | Verdict | Reason |
|---|---|---|
| Syntax and standard-library recall | Let it go | The agent has it and gates catch the errors. |
| Boilerplate, scaffolding, and glue code | Let it go | Agents handle it best. Keep an eye out for duplicated code. |
| Exploring a new tool’s configuration by hand | Let most of it go | One engineer in Anthropic’s study: “now I rely on AI to tell me how to use new tools and so I lack the expertise.” Keep hands-on depth for the two or three tools you are accountable for. |
| Debugging from first principles | Protect it | The largest gap in the skill-formation trial, and what incidents need once the agent’s fix has failed twice. |
| The “collateral” learning from struggling with a problem | Protect it deliberately | Anthropic’s study names its loss as the core worry. No gate replaces it. |
| Reading every diff line by line | Change its shape | It does not scale: Faros AI (April 2026, 22,000 developers) measured 31.3% more pull requests merging with no review at all. Replace it with reading evidence instead of code, keeping code reading for high-risk changes. |
The paradox of supervision, and how you stay out of it
Section titled “The paradox of supervision, and how you stay out of it”The paradox works like a ratchet: delegating erodes the skill you need to check what you delegated, weaker checking means more trust, and more trust means more delegation. The way out is not to delegate less. It is to change how you delegate, and to schedule the practice that delegation removes.
The skill-formation trial and Anthropic’s workplace study point to four habits:
- Ask for understanding, not only for output. Participants who asked conceptual questions or asked follow-up questions about generated code kept the most. Turn that into a rule for the session (the setup below does this).
- Predict before you read. Before you open the agent’s diff, write two lines saying what you expect it to change. A wrong prediction points at a gap in your model of the system, and that gap is worth more than the diff.
- Solve some problems by hand on purpose. One engineer in Anthropic’s study: “Every once in a while, even if I know that Claude can nail a problem, I will not ask it to. It helps me keep myself sharp.” Put a fixed slot on your calendar for this, not “when there is time”.
- Debug first, delegate second. When a failure lands, spend 15 minutes forming your own hypothesis before you hand it to the agent. Then compare your hypothesis with what the agent found.
Set up a learning mode in your tool
Section titled “Set up a learning mode in your tool”Only Claude Code ships a built-in style for this; in the other two tools you give the agent the same contract as an instruction.
Claude Code has built-in Explanatory and Learning output styles. The skill-formation study names “Claude Code Learning and Explanatory mode” as tools “designed to foster understanding”. In a session, run /output-style (added in Claude Code 2.1.269) and pick Learning for a codebase or library you are learning, or Explanatory when you want reasons without extra exercises. Switch back for familiar work: the explanations cost tokens and time.
We found no built-in learning or explanatory mode in codex --help or codex features list (Codex CLI 0.157.1, checked 26 September 2026). Paste the learning contract below at the start of the session. To make it standing for a shared repository, agree with your tech lead before adding it to AGENTS.md, which Codex reads before any work, because it changes every teammate’s sessions.
Paste the learning contract below at the start of an Agent chat. To make it stick for a project, save it as a Rule, Cursor’s mechanism for giving the agent standing instructions. Because Rules can be shared across a team, keep a personal learning rule out of the team’s shared set unless the team agrees.
A portfolio plan for each ladder level
Section titled “A portfolio plan for each ladder level”When agents write the code, portfolio evidence moves to the artifacts that decided what got built and proved it worked. Each ladder level adds one kind of artifact. Keep the earlier ones, because a portfolio that shows only the top level looks untested.
| Ladder level | Keep sharp by hand | Hand over to the agent | Portfolio evidence to collect | Signal that you are ready for the next level |
|---|---|---|---|---|
| Level 1–2: assisted and paired | Debugging, reading unfamiliar code, the core of your main language | Boilerplate, test scaffolding, lookups | Explain-back notes on one agent change a week; one problem a week solved by hand | You can predict most of an agent diff before you read it |
| Level 3: you review the diffs | Code review judgment, spotting design drift | Implementation of well-scoped tasks | A review log: defects you caught, the class of each, and the check that would have caught it automatically | Most defects in your log have a check that now catches them |
| Level 4: you write the specs | Specification, oracle design, architecture decisions | Multi-hour runs against a stop condition | Pairs of spec and evidence: the spec, the acceptance criteria, the test that failed first, and the merged result; architecture decision records | Your specs come back right on the first run more often than not, and the misses teach you something new |
| Level 5: you run the factory | Taste, the decision on what may run unattended, incident debugging | Whole loops, including their review station | Harness contributions: gates, fitness functions, and the written case for each loop that runs unattended; incidents where a gate held | Other people’s loops run on your gates. Level 5 is not a default career goal: Shapiro places only “a handful of people” there, in “small teams, less than five people”, and a strong Level 4 portfolio is a complete one |
Copy the one-page skills plan
Section titled “Copy the one-page skills plan”Save this as a private file, fill it in, and bring it to your next one-to-one. It is the exit artifact of the developer track: a written plan for the skills you keep sharp and the ones you hand over.
# Skills plan: <your name>, <quarter>
## Current ladder level, per loop- Feature work in <main repo>: Level <n>- Bug fixes: Level <n>- Refactors and migrations: Level <n>
## Compounding skills: one goal each this quarter- Specification: <e.g. acceptance criteria before every task; first-run match rate tracked>- Verification design: <e.g. one agent-proof check per feature>- Architecture: <e.g. one decision record a week>- Taste: <e.g. one rejected green result a week, with the reason written down>
## Protected practice (calendar slots, not intentions)- Unassisted problem: <day, time, 60 min>- Debug-first rule: 15 minutes of my own hypothesis before I delegate a failure- Learning mode on for: <library or subsystem>
## Handed over on purpose- <skill>: handed over because <gate that now covers it>
## Evidence collected this quarter- Spec and evidence pairs: <links>- Review or trust log: <link>- Decision records: <links>- Gates or checks I added: <links>Prompts that turn your history into evidence
Section titled “Prompts that turn your history into evidence”How do you prove your skills are holding?
Section titled “How do you prove your skills are holding?”Your judgment of your own skill is the weakest check available, so measure it. In a study of 16 experienced open-source developers on 246 issues with early-2025 tools (METR, July 2025), tasks took 19% longer with AI, yet afterwards they “still believed AI had sped them up by 20%”. Use checks that produce a number or an artifact:
- Unassisted kata, monthly. Solve one timed problem in your main language with no agent. Track the time and the number of mistakes. A steady drift upward is the early warning.
- Explain-back score. After the explain-back prompt, count the points you missed. If the misses keep growing on familiar modules, your model of the system is thinning.
- Prediction hit rate. Compare your two-line prediction with the actual diff. Record hit or miss.
- First-run spec match. For Level 4 work, record whether the agent’s first result met your acceptance criteria. This measures your specification skill, not the model.
- Catches that became checks. Count the entries in your review log that now have an automated check. That count is verification design, measured.
The code does not rely on your self-assessment either: it goes through tests, types, lint, an evidence bundle on each pull request, and a human read for high-risk changes. You own the practice; your tech lead reviews the plan and evidence with you each quarter and signs off on the level you work at in each loop.
What tech leads should change on the team
Section titled “What tech leads should change on the team”The Anthropic study quotes a senior engineer: “It’s been sad that more junior people don’t come to me with questions as often, though they definitely get their questions answered more effectively and learn faster.” Those questions used to pass on judgment. Three changes restore the channel without slowing the team:
- Pair juniors on specs, not on code. Review a junior’s acceptance criteria before the run, not only the diff after it. The spec shows what they understood.
- Make the portfolio the unit of review. Ask for spec and evidence pairs, review logs, and decision records in one-to-ones and performance reviews, instead of pull request or line counts. Career ladders when output is cheap covers the matrix.
- Protect practice time in the plan. A learning-mode slot and an unassisted problem each week cost less than one incident that nobody on the team can debug. Junior developers on an agentic team and upskilling turn this into a curriculum.
What goes wrong with a craft plan when agents write the code
Section titled “What goes wrong with a craft plan when agents write the code”You stop being able to debug your own system. The symptom is that every failure goes straight to the agent, and you cannot say what the fix changed. To recover, apply the debug-first rule for a month, and run the “find where your model is thin” prompt on the module behind the last incident.
Your portfolio shows only throughput. Merged pull request counts say nothing about judgment, and reviewers now discount them. To recover, rebuild the quarter’s evidence with the review-log prompt, and attach a spec and evidence pair to each achievement you claim.
The learning mode becomes permanent. Explanations on every task slow shipping work, and you start skimming them. Turn it on per subsystem or library, and off when your explain-back score stays high for a month.
You protect the wrong skills. Hours drilling syntax are hours not spent on specification. If an agent and a gate cover a skill, hand it over and record the gate.
Reassurance replaces evidence. “Engineers will always be needed” and “juniors are finished” are both unsourced. Treat any career forecast without a publisher and a date as noise.
What no source can tell you yet
Section titled “What no source can tell you yet”No study available on 26 September 2026 measures what happens to an engineer’s skills over years of agent-heavy work. The skill-formation trial measured comprehension right after the task. Anthropic’s workplace study is self-report from one company. We could not retrieve and verify the employment data for early-career developers, so this page prints no number about it. What the evidence does support is narrower and more useful: how you delegate changes what you keep, and the skills that decide what gets built and whether it can be trusted are the ones agents do not take over.
Where to go next with your craft
Section titled “Where to go next with your craft”Frequently asked questions
Which engineering skills compound when agents write the code?
Specification, verification design, architecture and taste. Each one decides what the agent builds or whether its output can be trusted, and each one is named as a human responsibility by third-party sources such as Karpathy's Sequoia Ascent summary and DORA's 2025 report.
What is the paradox of supervision?
A term from Anthropic's December 2025 study of its own engineers: using an agent well requires supervising it, and supervising it requires the coding skills that heavy delegation can wear away.
Does using an AI assistant stop developers learning?
In Anthropic's January 2026 trial with 52 mostly junior engineers, the group that used AI scored 50% on a follow-up quiz against 67% for the group that coded by hand, with the biggest gap on debugging. Participants who asked conceptual questions or asked for explanations kept most of the learning. The study measured comprehension right after the task, not long-term skill.
What should a developer's portfolio show when agents write most of the code?
Specs and acceptance criteria you wrote, tests and checks that caught agent mistakes, architecture decisions you recorded, and a review or trust log. Each ladder level adds one kind of evidence.