Upskilling a team for agentic engineering
Upskilling a team for agentic engineering means a curriculum tied to the autonomy ladder level each loop runs at, katas built from the team’s own repository and graded by tests, fortnightly review-to-learn sessions on real agent pull requests, and a dated literacy record per person. The same record is evidence for the EU AI Act’s Article 4 AI-literacy duty.
You bought eight seats in the spring, posted three links and ran one lunch-and-learn. Now two engineers run agents for hours against specs, three use them as autocomplete, and one refuses to touch them. The review queue is full of agent pull requests that only two people know how to verify. Then a customer’s security questionnaire asks how you meet the AI Act’s AI-literacy duty, and nobody knows who was trained on what. This page is for the tech lead who has to fix the first problem and the CTO who has to answer the second.
What you get from an upskilling programme built this way
Section titled “What you get from an upskilling programme built this way”- A curriculum table: what people learn at each ladder transition, the kata that proves it, and who signs the exit.
- A 90-minute baseline module for every seat holder.
- A kata manifest and two prompts that turn your own history into katas.
- A fortnightly review-to-learn session with a rule for what it must produce.
- A YAML literacy record that serves the curriculum and Article 4, with an audit prompt.
- Four measures of whether it works, and the failure modes to watch.
What does the evidence say about training people to work with agents?
Section titled “What does the evidence say about training people to work with agents?”There is little direct evidence on training programmes for agent work, so this page builds on two findings, marks the rest as design, and has the programme measure its own effect on your team.
How people delegate decides what they learn. In a randomized trial published by Anthropic on 29 January 2026, 52 mostly junior engineers learned a new Python library. The group that used an AI assistant averaged 50% on a follow-up quiz; the group that coded by hand averaged 67%. Participants who asked conceptual questions, or generated code and then asked for explanations, averaged 65% or higher. The sample was small and the publisher is a model vendor, but the lesson holds: design practice so people keep understanding, not only output.
The team’s system decides what the agent amplifies. The 2025 DORA report (Google Cloud, 23 September 2025) puts it this way: “AI doesn’t fix a team; it amplifies what’s already there.” Its AI Capabilities Model lists a “clear and communicated AI stance” as the first of seven capabilities. For upskilling, that means teaching people the team’s gates and rules, not only the tool’s features.
Build the curriculum around ladder transitions
Section titled “Build the curriculum around ladder transitions”The autonomy ladder runs from Level 0 (by hand) to Level 5 (a factory that no human reads). A level belongs to a loop, not to a person: one engineer can run dependency bumps at Level 4 and feature work at Level 2. So the curriculum teaches the transition a person’s main loop is making next, not a job title. One map: the ladder, the lifecycle and the stations shows where each loop sits.
| Transition | What the person learns | Kata that proves it | Exit evidence (tech lead signs) |
|---|---|---|---|
| Baseline (everyone with a seat) | Permissions and sandboxing in their tool, prompt injection, secrets, hallucinated dependencies, the team’s risk classes, what each CI gate proves | None; a five-question check after the module | Record entry with the date and the check passed |
| L1 → L2: assisted to paired | Delegating a whole task with acceptance criteria; reading the agent’s plan before it edits; stopping and redirecting a run | Delegation kata: finish a scoped ticket through the agent, with the hidden tests passing and no edits to test files | Two delegation katas passed; explain-back on one without the diff open |
| L2 → L3: paired to reviewing diffs | Reviewing an agent pull request against its criteria; risk classes that need a human to read code; shared rules in CLAUDE.md or AGENTS.md | Review kata: find a seeded defect in an agent change and name the check that would catch it | Three review katas passed; one real finding turned into a check |
| L3 → L4: reviewing to writing specs | Writing executable acceptance criteria; oracle strength; stop conditions; reading an evidence bundle instead of a diff | Spec kata: write criteria for a past ticket so the agent’s first run passes the original, hidden test suite | Two spec katas passed on the first run; one real ticket shipped from the person’s spec with its evidence bundle |
| L4 → L5: specs to running the loop | Fitness functions, gates that agents cannot edit, the written case for a loop that runs unattended | Harness kata: add a gate that catches a class of defect the kata set seeds | A gate merged that caught or would have caught a real defect; the case for one unattended loop reviewed by the CTO |
Two rules keep the table honest. First, nobody skips the baseline, however senior: seniority does not protect against prompt injection. Second, a transition is signed off on evidence from katas and real work, not on attendance. Level 5 is not a default goal; a strong Level 4 is a complete curriculum for most engineers.
Run the baseline module before anyone uses agents on production code
Section titled “Run the baseline module before anyone uses agents on production code”The baseline is 90 minutes, run by the tech lead or a senior for each new seat holder and again when the team adopts a new agent. It covers what can go wrong before what the tool can do.
- Permissions and sandboxing (20 minutes). Each person opens their own tool and shows the current mode: Claude Code’s permission modes (on the
latestchannel from v2.1.283, auto mode is the starting mode for most interactive sessions), Codex’s approval policy (-a on-requestor-a never) and sandbox (read-only,workspace-write), or Cursor’s Run modes. Permissions and sandboxing for coding agents is the reference. - Injection and secrets (20 minutes). Walk through one real incident from the agent threat model: text in an issue, a web page or a dependency can steer an agent that has tools and credentials. Show where the team keeps secrets and why an agent session never gets production credentials.
- Dependencies the agent invents (15 minutes). Check that a package the agent adds exists, is the intended one and is pinned. Dependency verification has the checks.
- What the gates prove (20 minutes). Open one merged agent pull request and, for each CI check, say what it proves and what it does not. Name the team’s risk classes (auth, money, schema, migrations) where a human reads the code.
- The check (15 minutes). Five questions, answered in writing, one per step above plus one on the risk classes. Record the date and the result.
Adopt this baseline check as-is and adapt the answers to your repository:
## Agent baseline check (write your answers; 15 minutes)1. Which permission mode or approval policy does your tool start in on this repository, and how do you change it for one session?2. An issue description tells the agent to "also update the deploy token". What stops that from happening in our setup, and what do you do?3. The agent adds a package you have never heard of. Name two checks you run before you accept the change.4. Our CI passed on an agent pull request. Name one thing the gates prove and one thing they do not prove.5. Which change classes on this team need a human to read the code, and who?Build katas from your own repository
Section titled “Build katas from your own repository”A kata is a time-boxed exercise with a known answer. Katas from your own repository teach your system, gates and risk classes. A usable kata has three properties:
- The tests grade it. A hidden test suite or a named, seeded defect decides pass or fail, so nobody argues about the result.
- The answer key stays out of reach. A kata cut from your repository still has its answer in git history, so publish each kata to a separate practice repository with fresh history and keep the key outside it. If that is too heavy, use the honour system: nobody looks at
mainor the original pull request until the debrief. - It has a time box. 30 to 60 minutes, so a kata fits in a protected slot every week.
Keep the kata set in a manifest so it can be reviewed, reused and retired like code:
- id: review-03-retry-idempotency transition: "L2 -> L3" type: review # delegation | review | spec | harness branch: kata/review-03 timebox_min: 45 task: "Review the agent's change to the payment retry worker." pass: "Names the missing idempotency check and the test that would catch it." answer_key: "private: senior's notes, kept outside the practice repository" owner: "a.nowak" created: 2026-09-26 retire_when: "the retry worker is rewritten or the seeded defect class gets a gate"The senior builds each kata with the agent, then checks the result by hand before anyone runs it. The two prompts below cover the kata types that take the most time to build.
The kata branch still carries the original change in its history, so the senior copies it into a fresh practice repository, with the pre-change tree on main and the kata on its branch:
# Run from the root of your working repository. ../kata-review-03 becomes the practice repository.base=$(git merge-base main kata/review-03)git init -b main ../kata-review-03git archive "$base" | tar -x -C ../kata-review-03git -C ../kata-review-03 add -A && git -C ../kata-review-03 commit -m "Base"git -C ../kata-review-03 switch -c kata/review-03git -C ../kata-review-03 rm -rq . && git archive kata/review-03 | tar -x -C ../kata-review-03git -C ../kata-review-03 add -A && git -C ../kata-review-03 commit -m "Kata change"For a spec kata, copy the listed test files (git show <sha>:<path>) to a key directory outside both repositories and run them against the engineer’s result. Replace src/payments/ and src/auth/ in the prompts with your own directories.
How does each tool run a kata?
Section titled “How does each tool run a kata?”The kata and its grading are the same in every tool; only session isolation and the review-agent comparison differ.
Start the kata in its own worktree so it cannot touch other work: claude -w kata-review-03 creates a new git worktree for the session. For a spec kata, enter plan mode with /plan so the agent writes its plan before any edit, and compare the plan with the criteria. For a kata in a module the person is learning, run /output-style learning (the command needs v2.1.269 or later). The agent then explains its choices and, at a real design decision, leaves a TODO(human) marker for the person to fill in. The choice is saved to .claude/settings.local.json, so teammates are unaffected. After a review kata, run /code-review on the same branch and compare its findings with the person’s.
Start the kata with codex --worktree, which runs the session in a new managed git worktree. Use /plan for spec katas. We found no built-in learning mode in codex --help or codex features list (Codex CLI 0.157.1, checked 26 September 2026), so for a learning kata the person pastes a learning contract at the start of the session. After a review kata, run codex review --base main from the kata branch and compare its findings with the person’s.
Run the kata in a Cursor worktree so the Agent works in an isolated Git checkout, then start an Agent chat. Use Plan Mode, which “creates detailed implementation plans before writing any code” (Cursor docs, checked 28 August 2026), for spec katas. For a learning kata, the person pastes a learning contract into the chat, not into the team’s shared Rules. After a review kata, if your team uses Bugbot, open the kata branch as a draft pull request and compare Bugbot’s comments with the person’s findings.
Your craft and career when agents write the code has a learning contract to paste in Codex and Cursor, and growing junior developers has a debugging-kata prompt for juniors.
Run review-to-learn sessions for the whole team
Section titled “Run review-to-learn sessions for the whole team”Katas teach on known answers; review-to-learn sessions teach on real work. Run one every two weeks for 45 minutes with the whole team, rotating the facilitator.
- The facilitator picks one merged agent pull request from the last two weeks that touched error handling, state or a boundary. Avoid renames and dependency bumps.
- Predict (5 minutes). From the ticket and acceptance criteria only, everyone writes two lines: what the change must touch and what could go wrong.
- Review (15 minutes). Everyone reviews the diff and the evidence bundle and writes findings, classed as logic, security, performance, design or test gap.
- Compare with the review agent (10 minutes). Show what
/code-review,codex reviewor Bugbot found on the same change. List what people found that the agent missed, and the reverse. - Decide one durable artifact (15 minutes). End with one change: a new check, a rule in
CLAUDE.mdorAGENTS.md, a new kata, or an explicit “nothing”. Name an owner and a date.
Step 5 keeps the session from turning into a demo: a session with no artifact and no explicit “nothing” did not happen. Knowledge sharing: version evidence, not tricks covers where each artifact lives and how it is retired.
Keep a literacy record that also serves Article 4
Section titled “Keep a literacy record that also serves Article 4”The EU AI Act reaches a company that only uses coding agents mainly through Article 4. In its 2024 wording, it asks providers and deployers to take measures for a sufficient level of AI literacy among staff who use AI systems, given their experience, training and context of use. The Digital Omnibus (Regulation (EU) 2026/1744, in force 27 July 2026) softened the duty but did not remove it, according to secondary law-firm summaries from Gibson Dunn and aiactblog.nl (2026). Read the amended wording on EUR-Lex before you quote it, and confirm your position with counsel. The EU AI Act for companies building software with agents covers the roles and the timeline.
A curriculum already produces what Article 4 asks for. This format extends the AI Act page’s record with the ladder transition and katas, so one file serves the tech lead and compliance:
- person: "j.kowalski" role: developer systems: [claude-code, codex] autonomy: "local, reversible changes; no production credentials" ladder: main_loop: "feature work, billing service" current_level: L3 working_towards: L4 measures: - name: "Agent baseline: permissions, injection, secrets, dependencies, gates" date: 2026-09-12 material: "https://wiki.example.com/agents/baseline" check: passed - name: "Review kata review-03-retry-idempotency" date: 2026-09-19 result: passed - name: "Review-to-learn session: billing retry pull request" date: 2026-09-24 material: "https://wiki.example.com/agents/review-to-learn/2026-09-24" exits_signed: - transition: "L2 -> L3" date: 2026-09-24 evidence: "https://git.example.com/org/billing/pull/412" signed_by: "tech-lead-billing" next_review: 2027-09-12 owner: "head-of-engineering"Keep kata results as pass or fail and a date, not as times or scores. The record proves measures exist; it is not a performance file (see the failure modes below).
Run the audit every quarter and whenever the team adds a tool, raises an autonomy level or switches model family. Those events are also the triggers to re-run the baseline; keeping a team current is the page that watches for them.
How do you verify the programme works?
Section titled “How do you verify the programme works?”Attendance proves nothing. These four measures do, and the first is the one compliance reads.
| Measure | Definition | Target and owner |
|---|---|---|
| Literacy coverage | Share of people with an agent seat or agent-approval rights whose record has a measure dated within the last 12 months | 100%, reported quarterly by the head of engineering. A gap is an access problem: no record, no seat |
| Signed exits per quarter | Ladder transitions signed off on kata and real-work evidence | One per engineer per quarter is a strong pace; zero for two quarters is a conversation, not a failure |
| Kata first-attempt pass rate | Share of katas passed on the first attempt, per transition, for the whole team | Rises over the quarter; a kata nobody passes is a broken kata |
| Findings that became checks | Review-to-learn findings that now have an automated check | At least one per session in two of every three sessions |
Pair them with team delivery measures, not individual ones: change failure rate and rework on agent pull requests, against the baseline from metrics frameworks for agentic teams. If literacy coverage is 100% but agent changes fail in production as often as before, the curriculum teaches the wrong things: compare the kata classes with the defects that escape.
How does this answer the Tech Lead Scorecard’s upskilling question?
Section titled “How does this answer the Tech Lead Scorecard’s upskilling question?”Question 18 of the Tech Lead Scorecard asks how you level up the team’s AI skills. “I share links” scores one of three points and “occasional sessions” two. The top answer, “structured path + practice + review”, maps to this page one to one: the curriculum table is the path, the katas are the practice, and the review-to-learn sessions are the review.
A quarter for a team of eight looks like this in the calendar:
| Weeks | What runs | Time per engineer |
|---|---|---|
| 1–2 | Baseline module for anyone without a record entry; the tech lead places each person’s main loop on the ladder | 90 minutes once |
| 3–12 | One kata a week in a protected slot, chosen for the person’s next transition | 30–60 minutes a week |
| 3–12, every other week | Review-to-learn session for the whole team | 45 minutes a fortnight |
| 12 | Literacy-record audit, exits signed, kata set pruned | 30 minutes for the tech lead per engineer |
The times are this programme’s design, not a measured benchmark. Senior time for building about one new kata a week comes on top.
When an upskilling programme goes wrong
Section titled “When an upskilling programme goes wrong”Kata results turn into a performance ranking. Once kata times appear next to names, people pick easy katas and hide gaps. Using an AI system to evaluate employees can also move you into the AI Act’s high-risk employment category (Annex III, from 2 December 2027 as deferred by the Digital Omnibus, per secondary summaries; see the EU AI Act page for the timeline). To recover, record kata results as pass or fail only, keep them out of performance reviews, and use career ladders built for cheap output for performance instead.
The curriculum teaches features, not verification. Sessions demo new slash commands, and review quality does not change. To recover, give every transition a kata graded by tests, and cut sessions whose only output is “we saw a feature”.
Seniors skip the baseline. Senior engineers opt out, then approve agent changes under a permission mode they never checked. To recover, make the baseline a condition for approval rights on agent pull requests, and ask a senior to run the next one. Bringing skeptics and senior engineers along covers the conversation.
The kata set goes stale. Katas in code that no longer exists teach the old system. To recover, give every kata an owner and a retire_when condition, and prune the set in week 12 of every quarter.
The literacy record is filled in once for the audit. Entries appear the week before a customer questionnaire and never change. To recover, generate entries from the kata log and session notes, and run the audit prompt quarterly.
Protected practice is the first thing cut. The kata slot disappears in a deadline sprint and never returns. To recover, plan katas as sprint tickets with an owner, and report skipped slots in the retrospective.
Where to go next with upskilling
Section titled “Where to go next with upskilling”Frequently asked questions
How do you upskill a team for agentic engineering?
Tie the curriculum to the autonomy ladder level of the loop each person works in, practise with katas built from your own repository that tests grade, run fortnightly review-to-learn sessions on real agent pull requests, and keep a dated record of who completed what. Sharing links is not enablement.
What is a kata for agentic engineering?
A time-boxed exercise on a branch of your own repository with a known answer: find a seeded defect in an agent's change, write acceptance criteria that a hidden test suite grades, or add a gate that catches a class of defect. The tests decide the result, not the person running it.
Does the upskilling record help with the EU AI Act?
Yes. Article 4 asks deployers of AI systems to take measures for the AI literacy of staff who use them. The Digital Omnibus softened that duty but did not remove it. A dated record per person, with the material and the completed katas, is evidence those measures exist. It is not legal advice; confirm with counsel.
How is upskilling different from keeping current?
Upskilling builds lasting skills per ladder level: delegation, review, specification and harness design. Keeping current is the weekly watch on tool releases and model changes. A release that changes autonomy or models is a trigger to re-run part of the curriculum.