Skip to content

Growing junior developers when agents write the code

Growing junior developers on a team where agents write the code means replacing the practice that typing used to provide with deliberate practice: debugging by hand, reviewing agent changes to learn, and writing specs that a senior checks before the agent runs. Without that plan, juniors ship fast but do not build the judgment needed to supervise agents.

Your newest engineer merged 14 pull requests in their first month, more than anyone on the team. Then a payment retry failed in staging, and they could not say what their own change did, because an agent wrote it and the tests were green. The senior who used to answer their questions noticed they had stopped asking. This page is for the tech lead who has to turn that junior into a future senior, and for the CTO or executive deciding whether to keep hiring juniors at all.

What a junior development plan for an agentic team gives you

Section titled “What a junior development plan for an agentic team gives you”
  • The evidence on juniors and agents, with dates and limits.
  • Four practices that replace what typing used to teach, and what a lead checks for each.
  • A learning setup for Claude Code, Codex, and Cursor, plus four copy-paste prompts.
  • A weekly review-to-learn session and a spec-review loop.
  • A 90-day plan and four growth measures that are not pull request counts.
  • A decision table for CTOs on the risk of hiring no juniors.

What does the evidence say about juniors and coding agents?

Section titled “What does the evidence say about juniors and coding agents?”

Two sources speak directly to juniors. Both come from Anthropic, a model vendor studying its own product and its own staff, so read them as that. We found no independent study of comparable design when this page was written.

The workplace study. In August 2025 Anthropic surveyed 132 of its engineers and researchers and ran 53 in-depth interviews. The results were published on 2 December 2025. A senior engineer said: “It’s been sad that more junior people don’t come to me with questions as often, though they definitely get their questions answered more effectively and learn faster.” The report names the underlying risk as the “paradox of supervision”:

“effectively using Claude requires supervision, and supervising Claude requires the very coding skills that may atrophy from AI overuse.”

Another senior explained where their own judgment came from: “I developed that ability by doing SWE ‘the hard way’.” Earlier in a career, they said, “it would take a lot of deliberate effort to continue growing my own abilities.” This is self-report from one company.

The skill-formation trial. On 29 January 2026 Judy Hanwen Shen and Alex Tamkin published a randomized trial with 52 mostly junior engineers learning a new Python library. The group that used an AI assistant averaged 50% on the follow-up quiz. The group that coded by hand averaged 67%. The AI group finished about two minutes faster, a difference that was not statistically significant. The largest gap was on debugging questions. Participants who only asked conceptual questions, or who generated code and then asked follow-up questions to understand it, scored highest. Participants who handed over all the coding, or leaned on the assistant to debug, scored lowest. The authors note that the sample was small and that the quiz measured comprehension shortly after the task, not long-term skill.

The practical reading for a lead: access to an agent does not stop a junior learning, but unstructured delegation does. The plan below structures it. Your craft and career when agents write the code covers the same evidence from the individual developer’s side.

The four practices that replace what typing used to teach

Section titled “The four practices that replace what typing used to teach”

Juniors used to learn by writing code that failed, then debugging it. When an agent writes the code, that loop disappears unless you rebuild it on purpose. Each practice below rebuilds one part of it, and each produces an artifact that a lead can check in minutes without reading the junior’s diffs.

PracticeWhat the junior doesWhat the agent doesEvidence the lead checks
Hand practiceDebugs from their own hypothesis first; solves one task a week unassistedNothing during the practice; afterwards compares its answerKata log: time, mistakes, and whether the hypothesis was right
Review to learnPredicts the change, reviews the agent’s pull request, and explains it backWrites the change, then grades the explanation against the codeExplain-back misses per week on familiar modules
Spec writingWrites acceptance criteria before every agent taskImplements against the criteria and reports which ones passFirst-run spec match rate, and the senior’s spec comments
Owning a checkAdds one automated check a month that an agent cannot editProposes candidates from recent defectsChecks merged, and defects each one has caught since

Hand practice and review to learn come first, because a junior cannot spec a system they cannot yet explain; spec writing starts in month two, owning a check in month three.

Keep these unassisted, and put them in the sprint plan as named items rather than intentions:

  • Debug first, delegate second. When a test fails, the junior writes a hypothesis and spends 15 minutes testing it before handing the failure to the agent. Debugging was the largest gap in the skill-formation trial.
  • One unassisted task a week. Pick a small, real ticket, 60 to 90 minutes, in the module the junior owns. No agent, with a senior available for questions.
  • Explain-back before approval. Before a junior marks their own agent-written pull request ready, they explain the change without the diff open.
  • Reading one module end to end. Once a month, the junior reads one module the agent changes often and writes a half-page note on its data flow and failure modes.

Syntax recall, boilerplate, scaffolding and configuration lookups can go to the agent from day one. The gates catch those errors, and drilling them costs time the four practices need.

Set up a learning configuration in each tool

Section titled “Set up a learning configuration in each tool”

Claude Code ships a built-in Learning style. Codex (checked in 0.157.1) and Cursor have no equivalent, so you give the agent the same contract as an instruction.

Claude Code ships four built-in output styles besides the default: Proactive, Concise, Explanatory, and Learning. In a session, run /output-style (added in Claude Code 2.1.269) and pick Learning while the junior works in a subsystem they are learning. Pick Explanatory when they want the reasons without exercises. Use plan mode (/plan) for spec practice: the agent drafts a plan, and edits stay blocked until the plan is approved, so the junior compares the plan with their own acceptance criteria before any code exists.

Review used to be where juniors learned from seniors. When agents open most pull requests, a junior can review far more code than before, and learn from each review, if the review has a structure and a grader. Run this session once a week for 45 minutes, with one senior and up to three juniors.

  1. Pick one merged agent pull request from the last week, in a module the juniors work in. Prefer one that touched error handling or state, not a rename.
  2. Predict (5 minutes). From the ticket and the acceptance criteria only, each junior writes two lines: what the change must touch, and what could go wrong.
  3. Review (15 minutes). Each junior reviews the diff and writes findings, classed as logic, security, performance, design, or test gap.
  4. Compare with the machine reviewer (10 minutes). Run the team’s review agent on the same change: /code-review in Claude Code, /review or codex review in Codex, or read the Bugbot comments on the pull request if your team uses Cursor’s reviewer. List what the juniors found that the agent missed, and the reverse.
  5. Senior debrief (15 minutes). The senior names the one finding that matters most and says why. The session ends with one action: a new check, a note in the module’s docs, or nothing, stated explicitly.

The comparison step is the point: a junior who finds what the review agent missed is learning judgment; one who finds only what it found is learning to be a slower review agent.

A spec shows what a junior understood before any code exists, which makes it the cheapest thing a senior can review. Ten minutes on acceptance criteria teaches more than an hour on a finished diff, and it moves the senior’s attention to where agents need humans most. Executable acceptance criteria: from story to failing test is the canonical guide to the format.

  1. The junior writes the criteria for the ticket: behaviour in Given / When / Then form, the edge cases, what must not change, the command that proves “done”, and, where the team uses executable criteria, the failing test.
  2. A senior reviews the spec, not the code (10 minutes). Use the checklist below. The senior returns comments, not a rewrite.
  3. The agent runs against the revised spec in plan mode first, then implements.
  4. The junior records the first-run result: which criteria passed on the first run, which failed, and why.
  5. In the one-to-one, discuss the misses. A miss caused by a vague criterion is a spec lesson. A miss caused by the agent is a verification lesson: which check would have caught it?

Adopt this spec-review checklist as-is for step 2:

## Spec review (10 minutes, senior)
- [ ] Every criterion is observable: a test or a command can decide it
- [ ] At least one edge case the junior did not get from the ticket
- [ ] "Must not change" names real files, APIs or behaviours
- [ ] The done command runs in CI, not only on a laptop
- [ ] The risk class is stated (auth, money, schema, migration = senior reads the code)
- [ ] One question back to the junior about why a criterion exists

A 90-day plan for a new junior on an agentic team

Section titled “A 90-day plan for a new junior on an agentic team”

Copy this into the junior’s onboarding doc. Each month adds a practice and keeps the earlier ones.

PeriodPractices addedAgent useExit evidence the lead signs off
Days 1–30Hand practice, explain-back, weekly review-to-learn sessionLearning style or learning contract on; small, low-risk ticketsFour kata entries; explain-back misses trending down on the junior’s module; can walk a senior through one request path end to end
Days 31–60Spec writing with senior spec review on every ticketPlan first, then implement; still no high-risk change classesFirst-run spec match recorded on every ticket; at least one spec miss discussed and understood
Days 61–90Owning a check: one agent-proof check mergedLearning contract only for new subsystems; normal delivery elsewhereOne check merged that caught or would have caught a real defect; a half-page note on one module’s failure modes

After day 90, the junior’s plan becomes the individual skills plan in Your craft and career when agents write the code, reviewed every quarter.

The senior runs this prompt, not the junior, after replacing the placeholder. The senior keeps the original fix commit as the answer key, and the junior solves the kata with no agent. Ask the junior not to read the branch history; the revert commit shows the answer.

How do you verify a junior is growing, not only shipping?

Section titled “How do you verify a junior is growing, not only shipping?”

Merged pull requests and lines changed measure the agent, not the junior. Use these four measures instead, recorded by the junior and reviewed by the lead every two weeks:

MeasureDefinitionHealthy direction
Explain-back missesMissed or wrong points in the graded explanation, per change, on modules the junior has worked in beforeFalls over the first 60 days, then stays low
First-run spec matchShare of tickets where the agent’s first result met every acceptance criterionRises from day 31; misses come with a written reason
Kata trendTime and mistakes on the weekly unassisted task, in the same moduleFlat or improving; a steady rise is the early warning of atrophy
Catches that became checksDefects the junior found in review that now have an automated checkAt least one per month from day 61

These measures are for coaching conversations, not for ranking people. Keep them in the junior’s own file, and do not put them on a team dashboard. Career ladders when output is cheap covers why output counts turn into targets.

The code itself does not depend on the junior’s growth. A junior’s agent-written pull requests go through the same gates as everyone’s: tests, types, lint, and an evidence bundle on each pull request. Two rules are specific to juniors. A junior does not approve changes in the high-risk classes (auth, money, schema, migrations) during the first 90 days, and a senior reads the code for those changes, as reading evidence instead of code prescribes. The tech lead signs off the exit evidence in the 90-day table, and the junior’s line manager signs off the level they work at afterwards.

What is the talent-pipeline risk of hiring no juniors?

Section titled “What is the talent-pipeline risk of hiring no juniors?”

This section is for CTOs and executives, and argues from the evidence above, not a forecast. Agents need supervision. The people who supervise well are seniors who built their judgment by doing the work “the hard way”, in their own words. A company that stops hiring juniors keeps its current seniors for now, but it stops producing the next ones. The cost shows up years later as too few people who can own verification design, architecture decisions and incident response, which are the jobs agents do not take over.

We could not retrieve and verify the published data on early-career developer employment, so this page prints no number about the market. Plan around your own headcount and attrition, not a headline.

OptionCost nowRisk laterWhen it fits
Hire no juniorsLowest headcount costNo internal path to senior; dependence on hiring seniors from a market where every company wants the same peopleA short-lived product or a team that will not exist in three years
Hire juniors, let them delegate freelySalary and seats; looks productive at onceJuniors who ship but cannot supervise: the pattern in the skill-formation trial’s lowest scorersNever as a plan; it is what happens by default
Hire juniors with an apprenticeship planSalary and seats, plus about two to three senior hours a week per junior for spec review and the weekly session (assuming ~10 tickets a week per junior and the weekly session shared by three juniors)Slower visible output in the first 90 daysAny team expected to own its systems for years

The senior-hours figure in the last row is this page’s own estimate from the session formats above (a weekly 45-minute session shared by up to three juniors, plus about 10 minutes of spec review per ticket at roughly 10 tickets a week per junior), not a measured benchmark. With fewer tickets or a session for one junior, the figure changes. Measure it on your own team in the first month.

  • How many engineers joined in the last 12 months at junior level, and what is the plan for each one’s first 90 days?
  • Who reviews juniors’ specs before agents run, and how many senior hours a week does that take?
  • Which change classes may a junior approve, and where is that written down?
  • What evidence shows our juniors can debug a production incident without an agent?
  • If our three most senior engineers left, who would own verification design, and how long would it take them to be ready?

What goes wrong when you grow juniors on an agentic team

Section titled “What goes wrong when you grow juniors on an agentic team”

The junior ships a lot and cannot explain any of it. The symptom is a high merge count and a blank look in the incident review. To recover, make the explain-back a condition for marking a pull request ready, and restart the 90-day plan at day 1 for that junior’s main module.

Seniors stop mentoring because the agent answers the questions. The Anthropic study describes this directly. To recover, move mentoring to spec review and the weekly session, where the senior’s time is scheduled and the junior’s questions are about judgment, not syntax.

Protected practice is the first thing cut under a deadline. Katas and review sessions disappear in the crunch sprint and never come back. To recover, put them in the sprint plan as tickets with an owner, and have the lead report skipped sessions in the retrospective.

The learning style stays on forever. Explanations on every task slow delivery, and the junior starts skimming them. To recover, turn the learning style or contract on per subsystem, and off once explain-back misses stay low for a month on that subsystem.

Growth measures become a leaderboard. Once kata times or spec match rates appear on a shared dashboard, juniors start picking easy tickets. Keep the measures in the junior’s own file and discuss them only in one-to-ones.

A junior approves a high-risk change. An agent-written migration or auth change passes the gates and a junior approves it. To recover, add the risk-class rule to the pull request template and to your code-owner rules, so a senior’s approval is required for those paths.

Frequently asked questions

Should junior developers use coding agents?

Yes, under a plan. In Anthropic's January 2026 trial with 52 mostly junior engineers, the group that used AI scored 50% on a follow-up quiz against 67% for the group that coded by hand, but participants who asked conceptual questions or asked for explanations averaged 65% or higher, close to the hand-coding group. How a junior delegates decides what they keep.

What should junior developers still practise by hand?

Debugging from their own hypothesis, one unassisted task a week, explaining an agent's change back without the diff, and writing acceptance criteria before any agent run. Syntax recall and boilerplate can go to the agent.

What is the paradox of supervision?

A term from Anthropic's December 2025 study of its own engineers: using an agent well requires supervising it, and supervising it requires the coding skills that heavy delegation can wear away. It is why juniors need deliberate practice, not only access to an agent.

What is the risk of hiring no junior developers?

The people who supervise agents are senior engineers who built their judgment by doing the work by hand. A team that hires no juniors stops producing the next group of supervisors, and the cost arrives years later as a shortage of people who can own verification, architecture and incidents.

Edit page

Last updated:

Cite this page — https://developertoolkit.ai/en/teams/junior-developers/, developertoolkit.ai