Developer
You want agents to ship work you can trust. The track runs from a first real task to reviewing evidence instead of every diff.
AI Developer Toolkit teaches developers, tech leads, CTOs, and executives to build software with coding agents (Cursor, Claude Code, and Codex) while keeping quality provable. The site has two entry points: a quick start for each tool, ending in a first real task, and four role tracks, each opening with a free scorecard that sets a baseline.
You have three agents to evaluate, several hundred pages, and a real job to get back to. This page gets you to the right first page.
Then pick the next page by where you are coming from:
| If you are… | Read next |
|---|---|
| New to agents | Quick wins |
| Coming from GitHub Copilot | Copilot compared with the three agents |
| Adding a second tool | Cursor vs Claude Code or Codex vs the others |
| Already fluent in one tool | Power-user tips, then the Cookbook |
| A platform or DevOps engineer | CI/CD with Claude Code, Codex automation, and MCP servers |
Each track is an ordered list of pages with an exit criterion per step. Take the scorecard first (about eight minutes): your answers show which steps to skip.
Developer
You want agents to ship work you can trust. The track runs from a first real task to reviewing evidence instead of every diff.
Tech lead
You need one workflow for the whole team and a review queue that is not the bottleneck: a shared harness, verification design, and review capacity.
CTO or VP Engineering
You decide the operating model, the guardrails, the vendors, and how to prove any of it works. Ten decisions, each ending in an artifact with an owner.
Executive
You need the effect on cost, speed, risk, and headcount, with a source on every number, and the decisions leadership owns.
C-Level Scorecard (shared with CTOs) · Executive track
The site is organized around one map: Dan Shapiro’s autonomy ladder, Level 0 to Level 5, published 23 January 2026. At Level 3 you review every diff; at Level 4 you write the specification and check the evidence; at Level 5 a pipeline ships while you run the factory. Shapiro puts 90% of self-described “AI-native” developers at Level 2 and says almost everyone tops out at Level 3. The useful next page is the one that moves you up one rung.
A faster agent without automated checks only moves work into review; evidence, not diffs explains what a reviewer checks instead of every line.
The sidebar is ordered by use; every group but Start here starts collapsed.
| Section | What it holds | Start at |
|---|---|---|
| Start here | This page, the four tracks, quick wins, and why it matters | Why agentic engineering |
| Autonomy Ladder | Each level and the human’s job at it | The ladder |
| Cursor · Claude Code · Codex | One tree per tool, cut by task; the same task in another tool is a one-segment URL swap | Cursor · Claude Code · Codex |
| Workflows & method | The tool-independent lifecycle and cross-tool workflows | Method · Workflows |
| Ecosystem & comparisons | Choosing a tool, models, MCP servers, skills, and frameworks | Ecosystem · Comparison |
| Cookbook | Prompts by language, framework, and stack | Cookbook |
| Lead a team | Rollout, the shared harness, review, and people | Lead a team |
| Engineering organization | Operating model, controls, security, measurement, and spend | Engineering organization |
| Strategy & business | Economics and evidence for executives | Strategy |
| Reference & help | Commands, configuration, glossary, and troubleshooting | Reference · Help |