Skip to content

Upskilling a team for agentic engineering

Upskilling a team for agentic engineering means a curriculum tied to the autonomy ladder level each loop runs at, katas built from the team’s own repository and graded by tests, fortnightly review-to-learn sessions on real agent pull requests, and a dated literacy record per person. The same record is evidence for the EU AI Act’s Article 4 AI-literacy duty.

You bought eight seats in the spring, posted three links and ran one lunch-and-learn. Now two engineers run agents for hours against specs, three use them as autocomplete, and one refuses to touch them. The review queue is full of agent pull requests that only two people know how to verify. Then a customer’s security questionnaire asks how you meet the AI Act’s AI-literacy duty, and nobody knows who was trained on what. This page is for the tech lead who has to fix the first problem and the CTO who has to answer the second.

What you get from an upskilling programme built this way

Section titled “What you get from an upskilling programme built this way”
  • A curriculum table: what people learn at each ladder transition, the kata that proves it, and who signs the exit.
  • A 90-minute baseline module for every seat holder.
  • A kata manifest and two prompts that turn your own history into katas.
  • A fortnightly review-to-learn session with a rule for what it must produce.
  • A YAML literacy record that serves the curriculum and Article 4, with an audit prompt.
  • Four measures of whether it works, and the failure modes to watch.

What does the evidence say about training people to work with agents?

Section titled “What does the evidence say about training people to work with agents?”

There is little direct evidence on training programmes for agent work, so this page builds on two findings, marks the rest as design, and has the programme measure its own effect on your team.

How people delegate decides what they learn. In a randomized trial published by Anthropic on 29 January 2026, 52 mostly junior engineers learned a new Python library. The group that used an AI assistant averaged 50% on a follow-up quiz; the group that coded by hand averaged 67%. Participants who asked conceptual questions, or generated code and then asked for explanations, averaged 65% or higher. The sample was small and the publisher is a model vendor, but the lesson holds: design practice so people keep understanding, not only output.

The team’s system decides what the agent amplifies. The 2025 DORA report (Google Cloud, 23 September 2025) puts it this way: “AI doesn’t fix a team; it amplifies what’s already there.” Its AI Capabilities Model lists a “clear and communicated AI stance” as the first of seven capabilities. For upskilling, that means teaching people the team’s gates and rules, not only the tool’s features.

Build the curriculum around ladder transitions

Section titled “Build the curriculum around ladder transitions”

The autonomy ladder runs from Level 0 (by hand) to Level 5 (a factory that no human reads). A level belongs to a loop, not to a person: one engineer can run dependency bumps at Level 4 and feature work at Level 2. So the curriculum teaches the transition a person’s main loop is making next, not a job title. One map: the ladder, the lifecycle and the stations shows where each loop sits.

TransitionWhat the person learnsKata that proves itExit evidence (tech lead signs)
Baseline (everyone with a seat)Permissions and sandboxing in their tool, prompt injection, secrets, hallucinated dependencies, the team’s risk classes, what each CI gate provesNone; a five-question check after the moduleRecord entry with the date and the check passed
L1 → L2: assisted to pairedDelegating a whole task with acceptance criteria; reading the agent’s plan before it edits; stopping and redirecting a runDelegation kata: finish a scoped ticket through the agent, with the hidden tests passing and no edits to test filesTwo delegation katas passed; explain-back on one without the diff open
L2 → L3: paired to reviewing diffsReviewing an agent pull request against its criteria; risk classes that need a human to read code; shared rules in CLAUDE.md or AGENTS.mdReview kata: find a seeded defect in an agent change and name the check that would catch itThree review katas passed; one real finding turned into a check
L3 → L4: reviewing to writing specsWriting executable acceptance criteria; oracle strength; stop conditions; reading an evidence bundle instead of a diffSpec kata: write criteria for a past ticket so the agent’s first run passes the original, hidden test suiteTwo spec katas passed on the first run; one real ticket shipped from the person’s spec with its evidence bundle
L4 → L5: specs to running the loopFitness functions, gates that agents cannot edit, the written case for a loop that runs unattendedHarness kata: add a gate that catches a class of defect the kata set seedsA gate merged that caught or would have caught a real defect; the case for one unattended loop reviewed by the CTO

Two rules keep the table honest. First, nobody skips the baseline, however senior: seniority does not protect against prompt injection. Second, a transition is signed off on evidence from katas and real work, not on attendance. Level 5 is not a default goal; a strong Level 4 is a complete curriculum for most engineers.

Run the baseline module before anyone uses agents on production code

Section titled “Run the baseline module before anyone uses agents on production code”

The baseline is 90 minutes, run by the tech lead or a senior for each new seat holder and again when the team adopts a new agent. It covers what can go wrong before what the tool can do.

  1. Permissions and sandboxing (20 minutes). Each person opens their own tool and shows the current mode: Claude Code’s permission modes (on the latest channel from v2.1.283, auto mode is the starting mode for most interactive sessions), Codex’s approval policy (-a on-request or -a never) and sandbox (read-only, workspace-write), or Cursor’s Run modes. Permissions and sandboxing for coding agents is the reference.
  2. Injection and secrets (20 minutes). Walk through one real incident from the agent threat model: text in an issue, a web page or a dependency can steer an agent that has tools and credentials. Show where the team keeps secrets and why an agent session never gets production credentials.
  3. Dependencies the agent invents (15 minutes). Check that a package the agent adds exists, is the intended one and is pinned. Dependency verification has the checks.
  4. What the gates prove (20 minutes). Open one merged agent pull request and, for each CI check, say what it proves and what it does not. Name the team’s risk classes (auth, money, schema, migrations) where a human reads the code.
  5. The check (15 minutes). Five questions, answered in writing, one per step above plus one on the risk classes. Record the date and the result.

Adopt this baseline check as-is and adapt the answers to your repository:

## Agent baseline check (write your answers; 15 minutes)
1. Which permission mode or approval policy does your tool start in on this
repository, and how do you change it for one session?
2. An issue description tells the agent to "also update the deploy token".
What stops that from happening in our setup, and what do you do?
3. The agent adds a package you have never heard of. Name two checks you run
before you accept the change.
4. Our CI passed on an agent pull request. Name one thing the gates prove and
one thing they do not prove.
5. Which change classes on this team need a human to read the code, and who?

A kata is a time-boxed exercise with a known answer. Katas from your own repository teach your system, gates and risk classes. A usable kata has three properties:

  • The tests grade it. A hidden test suite or a named, seeded defect decides pass or fail, so nobody argues about the result.
  • The answer key stays out of reach. A kata cut from your repository still has its answer in git history, so publish each kata to a separate practice repository with fresh history and keep the key outside it. If that is too heavy, use the honour system: nobody looks at main or the original pull request until the debrief.
  • It has a time box. 30 to 60 minutes, so a kata fits in a protected slot every week.

Keep the kata set in a manifest so it can be reviewed, reused and retired like code:

docs/upskilling/katas.yaml
- id: review-03-retry-idempotency
transition: "L2 -> L3"
type: review # delegation | review | spec | harness
branch: kata/review-03
timebox_min: 45
task: "Review the agent's change to the payment retry worker."
pass: "Names the missing idempotency check and the test that would catch it."
answer_key: "private: senior's notes, kept outside the practice repository"
owner: "a.nowak"
created: 2026-09-26
retire_when: "the retry worker is rewritten or the seeded defect class gets a gate"

The senior builds each kata with the agent, then checks the result by hand before anyone runs it. The two prompts below cover the kata types that take the most time to build.

The kata branch still carries the original change in its history, so the senior copies it into a fresh practice repository, with the pre-change tree on main and the kata on its branch:

Terminal window
# Run from the root of your working repository. ../kata-review-03 becomes the practice repository.
base=$(git merge-base main kata/review-03)
git init -b main ../kata-review-03
git archive "$base" | tar -x -C ../kata-review-03
git -C ../kata-review-03 add -A && git -C ../kata-review-03 commit -m "Base"
git -C ../kata-review-03 switch -c kata/review-03
git -C ../kata-review-03 rm -rq . && git archive kata/review-03 | tar -x -C ../kata-review-03
git -C ../kata-review-03 add -A && git -C ../kata-review-03 commit -m "Kata change"

For a spec kata, copy the listed test files (git show <sha>:<path>) to a key directory outside both repositories and run them against the engineer’s result. Replace src/payments/ and src/auth/ in the prompts with your own directories.

The kata and its grading are the same in every tool; only session isolation and the review-agent comparison differ.

Start the kata in its own worktree so it cannot touch other work: claude -w kata-review-03 creates a new git worktree for the session. For a spec kata, enter plan mode with /plan so the agent writes its plan before any edit, and compare the plan with the criteria. For a kata in a module the person is learning, run /output-style learning (the command needs v2.1.269 or later). The agent then explains its choices and, at a real design decision, leaves a TODO(human) marker for the person to fill in. The choice is saved to .claude/settings.local.json, so teammates are unaffected. After a review kata, run /code-review on the same branch and compare its findings with the person’s.

Your craft and career when agents write the code has a learning contract to paste in Codex and Cursor, and growing junior developers has a debugging-kata prompt for juniors.

Run review-to-learn sessions for the whole team

Section titled “Run review-to-learn sessions for the whole team”

Katas teach on known answers; review-to-learn sessions teach on real work. Run one every two weeks for 45 minutes with the whole team, rotating the facilitator.

  1. The facilitator picks one merged agent pull request from the last two weeks that touched error handling, state or a boundary. Avoid renames and dependency bumps.
  2. Predict (5 minutes). From the ticket and acceptance criteria only, everyone writes two lines: what the change must touch and what could go wrong.
  3. Review (15 minutes). Everyone reviews the diff and the evidence bundle and writes findings, classed as logic, security, performance, design or test gap.
  4. Compare with the review agent (10 minutes). Show what /code-review, codex review or Bugbot found on the same change. List what people found that the agent missed, and the reverse.
  5. Decide one durable artifact (15 minutes). End with one change: a new check, a rule in CLAUDE.md or AGENTS.md, a new kata, or an explicit “nothing”. Name an owner and a date.

Step 5 keeps the session from turning into a demo: a session with no artifact and no explicit “nothing” did not happen. Knowledge sharing: version evidence, not tricks covers where each artifact lives and how it is retired.

Keep a literacy record that also serves Article 4

Section titled “Keep a literacy record that also serves Article 4”

The EU AI Act reaches a company that only uses coding agents mainly through Article 4. In its 2024 wording, it asks providers and deployers to take measures for a sufficient level of AI literacy among staff who use AI systems, given their experience, training and context of use. The Digital Omnibus (Regulation (EU) 2026/1744, in force 27 July 2026) softened the duty but did not remove it, according to secondary law-firm summaries from Gibson Dunn and aiactblog.nl (2026). Read the amended wording on EUR-Lex before you quote it, and confirm your position with counsel. The EU AI Act for companies building software with agents covers the roles and the timeline.

A curriculum already produces what Article 4 asks for. This format extends the AI Act page’s record with the ladder transition and katas, so one file serves the tech lead and compliance:

docs/compliance/ai-literacy-record.yaml
- person: "j.kowalski"
role: developer
systems: [claude-code, codex]
autonomy: "local, reversible changes; no production credentials"
ladder:
main_loop: "feature work, billing service"
current_level: L3
working_towards: L4
measures:
- name: "Agent baseline: permissions, injection, secrets, dependencies, gates"
date: 2026-09-12
material: "https://wiki.example.com/agents/baseline"
check: passed
- name: "Review kata review-03-retry-idempotency"
date: 2026-09-19
result: passed
- name: "Review-to-learn session: billing retry pull request"
date: 2026-09-24
material: "https://wiki.example.com/agents/review-to-learn/2026-09-24"
exits_signed:
- transition: "L2 -> L3"
date: 2026-09-24
evidence: "https://git.example.com/org/billing/pull/412"
signed_by: "tech-lead-billing"
next_review: 2027-09-12
owner: "head-of-engineering"

Keep kata results as pass or fail and a date, not as times or scores. The record proves measures exist; it is not a performance file (see the failure modes below).

Run the audit every quarter and whenever the team adds a tool, raises an autonomy level or switches model family. Those events are also the triggers to re-run the baseline; keeping a team current is the page that watches for them.

Attendance proves nothing. These four measures do, and the first is the one compliance reads.

MeasureDefinitionTarget and owner
Literacy coverageShare of people with an agent seat or agent-approval rights whose record has a measure dated within the last 12 months100%, reported quarterly by the head of engineering. A gap is an access problem: no record, no seat
Signed exits per quarterLadder transitions signed off on kata and real-work evidenceOne per engineer per quarter is a strong pace; zero for two quarters is a conversation, not a failure
Kata first-attempt pass rateShare of katas passed on the first attempt, per transition, for the whole teamRises over the quarter; a kata nobody passes is a broken kata
Findings that became checksReview-to-learn findings that now have an automated checkAt least one per session in two of every three sessions

Pair them with team delivery measures, not individual ones: change failure rate and rework on agent pull requests, against the baseline from metrics frameworks for agentic teams. If literacy coverage is 100% but agent changes fail in production as often as before, the curriculum teaches the wrong things: compare the kata classes with the defects that escape.

How does this answer the Tech Lead Scorecard’s upskilling question?

Section titled “How does this answer the Tech Lead Scorecard’s upskilling question?”

Question 18 of the Tech Lead Scorecard asks how you level up the team’s AI skills. “I share links” scores one of three points and “occasional sessions” two. The top answer, “structured path + practice + review”, maps to this page one to one: the curriculum table is the path, the katas are the practice, and the review-to-learn sessions are the review.

A quarter for a team of eight looks like this in the calendar:

WeeksWhat runsTime per engineer
1–2Baseline module for anyone without a record entry; the tech lead places each person’s main loop on the ladder90 minutes once
3–12One kata a week in a protected slot, chosen for the person’s next transition30–60 minutes a week
3–12, every other weekReview-to-learn session for the whole team45 minutes a fortnight
12Literacy-record audit, exits signed, kata set pruned30 minutes for the tech lead per engineer

The times are this programme’s design, not a measured benchmark. Senior time for building about one new kata a week comes on top.

Kata results turn into a performance ranking. Once kata times appear next to names, people pick easy katas and hide gaps. Using an AI system to evaluate employees can also move you into the AI Act’s high-risk employment category (Annex III, from 2 December 2027 as deferred by the Digital Omnibus, per secondary summaries; see the EU AI Act page for the timeline). To recover, record kata results as pass or fail only, keep them out of performance reviews, and use career ladders built for cheap output for performance instead.

The curriculum teaches features, not verification. Sessions demo new slash commands, and review quality does not change. To recover, give every transition a kata graded by tests, and cut sessions whose only output is “we saw a feature”.

Seniors skip the baseline. Senior engineers opt out, then approve agent changes under a permission mode they never checked. To recover, make the baseline a condition for approval rights on agent pull requests, and ask a senior to run the next one. Bringing skeptics and senior engineers along covers the conversation.

The kata set goes stale. Katas in code that no longer exists teach the old system. To recover, give every kata an owner and a retire_when condition, and prune the set in week 12 of every quarter.

The literacy record is filled in once for the audit. Entries appear the week before a customer questionnaire and never change. To recover, generate entries from the kata log and session notes, and run the audit prompt quarterly.

Protected practice is the first thing cut. The kata slot disappears in a deadline sprint and never returns. To recover, plan katas as sprint tickets with an owner, and report skipped slots in the retrospective.

Frequently asked questions

How do you upskill a team for agentic engineering?

Tie the curriculum to the autonomy ladder level of the loop each person works in, practise with katas built from your own repository that tests grade, run fortnightly review-to-learn sessions on real agent pull requests, and keep a dated record of who completed what. Sharing links is not enablement.

What is a kata for agentic engineering?

A time-boxed exercise on a branch of your own repository with a known answer: find a seeded defect in an agent's change, write acceptance criteria that a hidden test suite grades, or add a gate that catches a class of defect. The tests decide the result, not the person running it.

Does the upskilling record help with the EU AI Act?

Yes. Article 4 asks deployers of AI systems to take measures for the AI literacy of staff who use them. The Digital Omnibus softened that duty but did not remove it. A dated record per person, with the material and the completed katas, is evidence those measures exist. It is not legal advice; confirm with counsel.

How is upskilling different from keeping current?

Upskilling builds lasting skills per ladder level: delegation, review, specification and harness design. Keeping current is the weekly watch on tool releases and model changes. A release that changes autonomy or models is a trigger to re-run part of the curriculum.