Skip to content

Start here: pick your path through agentic engineering

AI Developer Toolkit teaches developers, tech leads, CTOs, and executives to build software with coding agents (Cursor, Claude Code, and Codex) while keeping quality provable. The site has two entry points: a quick start for each tool, ending in a first real task, and four role tracks, each opening with a free scorecard that sets a baseline.

You have three agents to evaluate, several hundred pages, and a real job to get back to. This page gets you to the right first page.

Then pick the next page by where you are coming from:

If you are…Read next
New to agentsQuick wins
Coming from GitHub CopilotCopilot compared with the three agents
Adding a second toolCursor vs Claude Code or Codex vs the others
Already fluent in one toolPower-user tips, then the Cookbook
A platform or DevOps engineerCI/CD with Claude Code, Codex automation, and MCP servers

Each track is an ordered list of pages with an exit criterion per step. Take the scorecard first (about eight minutes): your answers show which steps to skip.

Developer

You want agents to ship work you can trust. The track runs from a first real task to reviewing evidence instead of every diff.

Developer Scorecard · Developer track

Tech lead

You need one workflow for the whole team and a review queue that is not the bottleneck: a shared harness, verification design, and review capacity.

Tech Lead Scorecard · Tech lead track

CTO or VP Engineering

You decide the operating model, the guardrails, the vendors, and how to prove any of it works. Ten decisions, each ending in an artifact with an owner.

C-Level Scorecard · CTO track

Executive

You need the effect on cost, speed, risk, and headcount, with a source on every number, and the decisions leadership owns.

C-Level Scorecard (shared with CTOs) · Executive track

The site is organized around one map: Dan Shapiro’s autonomy ladder, Level 0 to Level 5, published 23 January 2026. At Level 3 you review every diff; at Level 4 you write the specification and check the evidence; at Level 5 a pipeline ships while you run the factory. Shapiro puts 90% of self-described “AI-native” developers at Level 2 and says almost everyone tops out at Level 3. The useful next page is the one that moves you up one rung.

  1. Take your role’s scorecard and keep the result as your baseline.
  2. Finish your tool’s quick start on a real repository, not a toy project.
  3. Run the two prompts below in that repository: the first gives the agent its context, the second shows which checks stand between it and a bad merge.
  4. Commit the instructions file, and put the missing check from prompt 2 on your backlog.

A faster agent without automated checks only moves work into review; evidence, not diffs explains what a reviewer checks instead of every line.

The sidebar is ordered by use; every group but Start here starts collapsed.

SectionWhat it holdsStart at
Start hereThis page, the four tracks, quick wins, and why it mattersWhy agentic engineering
Autonomy LadderEach level and the human’s job at itThe ladder
Cursor · Claude Code · CodexOne tree per tool, cut by task; the same task in another tool is a one-segment URL swapCursor · Claude Code · Codex
Workflows & methodThe tool-independent lifecycle and cross-tool workflowsMethod · Workflows
Ecosystem & comparisonsChoosing a tool, models, MCP servers, skills, and frameworksEcosystem · Comparison
CookbookPrompts by language, framework, and stackCookbook
Lead a teamRollout, the shared harness, review, and peopleLead a team
Engineering organizationOperating model, controls, security, measurement, and spendEngineering organization
Strategy & businessEconomics and evidence for executivesStrategy
Reference & helpCommands, configuration, glossary, and troubleshootingReference · Help

Why readers stall, and how to get moving again

Section titled “Why readers stall, and how to get moving again”
  • Learning three tools at once. Nothing sticks. Finish one quick start, and add a second tool only for work the first does badly.
  • Reading without doing. Nothing changes. Run each prompt on your real repository the day you read it, and keep what works in the instructions file.
  • No instructions file. The agent repeats the same mistakes every session. Run the first prompt above and commit the result.
  • No verification gate. Review becomes the bottleneck, or bugs reach production. Run the second prompt, add the missing check it names, and re-take your scorecard after four weeks to see which answers moved.