Developer track: a reading path to agentic engineering
The developer track is the ordered reading path for developers who use Cursor, Claude Code or Codex. It starts with a free baseline, the Developer Scorecard, then climbs the autonomy ladder from a first real task to reviewing evidence instead of every diff. Each step ends in an artifact that proves it is done.
You already use an agent. Some days it saves an afternoon; other days you spend the afternoon reading what it wrote, because nothing tells you whether to trust it. The path below puts the habits that fix that in the order they pay off.
What the developer track gives you
Section titled “What the developer track gives you”- An order that compounds. Each step assumes the one before it: an instructions file makes the feedback loop reliable, and the feedback loop makes parallel agents safe.
- A “Done when” criterion per step. You move on when an artifact exists, not when the reading feels finished.
- One branch per tool where the tools differ, and one shared step where they do not.
- A clear free boundary. Every step is marked Free or Subscribers.
How do you use the developer track?
Section titled “How do you use the developer track?”- Take the Developer Scorecard first. Your answers show which steps you can skip, and the answer key maps each question to the page that moves it.
- Read steps 1 to 3 in order. Everything later assumes an installed tool and a reviewed instructions file.
- After step 3, skip any step whose “Done when” artifact you already have.
- Re-take the scorecard after four weeks and compare the sections.
Steps that are still being written are hidden and appear here when they are published. The list grows toward the end of the path, where it moves from reviewing diffs to reviewing evidence.
-
Why now: Place your current workflow on the ladder, so every later step is a move up one rung.
Done when: You can name your level and the one habit that keeps you on it.
-
Why now: Get one tool installed, configured and through a first real task in your repository.
Done when: A merged change the agent wrote, with the tests it ran.
-
Why now: The instructions file is the cheapest lever on output quality; write it before you scale use.
Done when: A reviewed CLAUDE.md or AGENTS.md under 150 lines with the real test commands.
-
Why now: Decide what the agent may run without asking before you let it run longer.
Done when: A committed permission policy and a sandbox for anything with network or secrets.
-
Why now: Intent, spec and plan are what you review instead of every line the agent writes.
Done when: One change shipped from an accepted intent.md, spec.md and plan.md.
-
Why now: Acceptance criteria written as failing tests are the contract the agent works to.
Done when: A story whose criteria exist as tests before the agent starts.
-
Why now: A single-command feedback loop lets the agent prove its own work before you look.
Done when: The agent runs the loop and reports exit status on every task.
-
Why now: Tests the agent can edit are not evidence; lock the checks it must not weaken.
Done when: A hook or CI rule that fails when protected tests change.
-
Why now: Parallel agents need isolated checkouts, ports and state before they need more tokens.
Done when: Two agents on two tasks in two worktrees without a collision.
-
Why now: Choose between a skill and an MCP server by the job, not the hype.
Done when: One capability added in the right layer, with a reason you can state.
-
Why now: Start on the default model, tune effort, and switch only when your own evals say so.
Done when: A routing rule for your repository backed by a before-and-after run.
-
Why now: See which named frameworks exist and what each prescribes.
Done when: You know whether any framework fits your next change.
-
Why now: The turn of the whole path: what you read instead of the diff, and when you still read code.
Done when: A change approved on its evidence bundle without reading every line.
-
Why now: Review an agent pull request by risk class and evidence, with a short escalation list.
Done when: A review checklist the review agent runs and you only escalate from.
-
Why now: Once the gates hold, let a well-shaped issue become a pull request without you.
Done when: One issue closed by a pipeline run you did not start by hand.
-
Why now: What your craft becomes when agents write the code, and how to show it.
Done when: A written plan for the skills you will keep sharp and the ones you hand over.
Which tool branch should you take?
Section titled “Which tool branch should you take?”Step 2 branches by tool; the rest of the path is shared, and each linked page shows all three tools where they differ. If you have not chosen a tool, read the tool comparison first.
| Cursor | Claude Code | Codex | |
|---|---|---|---|
| Quick start | Cursor quick start | Claude Code quick start | Codex quick start |
| Instructions file (step 3) | Rules | CLAUDE.md | AGENTS.md |
| Read-only planning | Plan Mode | claude --permission-mode plan or /plan | /plan |
| Parallel work | Worktrees | claude -w <name> | codex --worktree |
| Headless checks | Cursor CLI print mode (-p) | claude -p | codex exec |
Claude Code and Codex commands were checked on 26 September 2026 against Claude Code 2.1.283 and Codex CLI 0.157.1. Cursor features were last verified on 28 August 2026. Claude Code reads AGENTS.md when a project has no CLAUDE.md, from v2.1.277 on the latest release channel only.
How do you prove a step is done without reading every line?
Section titled “How do you prove a step is done without reading every line?”The exit criteria are built so that a check, not your reading, answers “is it done”. Three rules keep them honest:
- One command decides. From the feedback-loop step on, every task ends with the agent running the repository’s single test, lint and type-check command and reporting its exit status. A zero exit on a command you trust is the evidence; the agent’s summary is not.
- The agent cannot weaken the checks. Tests the agent can edit prove nothing. Lock the checks it must not weaken with a hook or a CI rule (the path’s protect-the-oracle step, which appears when published), and let CI re-run the same command on every pull request.
- You read the contract, not the diff. From the artifact-chain step on, you review the intent, spec and plan before the agent writes code, and after it finishes you check the evidence against them. You still read code line by line where a change is high-risk or the evidence is missing.
When does the developer path stall?
Section titled “When does the developer path stall?”Four blockers repeat, and each has a fast recovery:
- The agent reads the wrong files. Build output and generated files drown the context in large repositories. Name the files and directories that matter in your prompt, and record the ones to ignore in your instructions file.
- An MCP server will not connect. A missing environment variable is the usual cause. Run the server on its own first, confirm the token is exported, then restart the agent; in Claude Code,
claude --debug "mcp"prints the connection failure, and/mcpshows each server’s status. - Quality drops suddenly. Check the active model and effort before anything else:
/modelin Claude Code and Codex, the model picker in Cursor. Start on the tool’s default model, raise effort before you switch models, and switch only when your own evals say so. The models hub keeps the current defaults. - Headless runs hang. A run in CI waits on an approval prompt nobody will answer. In Claude Code, pre-approve the tools the job needs with
--allowedToolsor the project settings. In Codex, pass the approval policy before the subcommand, as incodex -a never exec …, inside a sandboxed runner.
For the full catalog of error messages and fixes, see the troubleshooting guide.
Where to go from the developer track
Section titled “Where to go from the developer track”Frequently asked questions
What is the developer track?
An ordered reading path for developers who use Cursor, Claude Code or Codex. It opens with the free Developer Scorecard as a baseline, then runs from the autonomy ladder through a tool quick start, the instructions file, the artifact chain and the test feedback loop to parallel worktrees, the ecosystem and model routing. Every step says why it comes at that point and names the evidence that shows you can move on.
Do I need to read the steps in order?
Read steps 1 to 3 in order, because each later step assumes an installed tool and a reviewed instructions file. After that, skip any step whose exit criterion you already meet, and come back to one when the scorecard or your own work shows a gap.
How do I know a step is done?
Each step has a "Done when" criterion that is an artifact, not a feeling: a merged change with the tests it ran, a reviewed instructions file, a single command whose exit status the agent reports. If you cannot point at the artifact, the step is not done.
Which steps are free?
The ladder pages and everything in Start here are free. The tool guides, workflows and scorecard answer pages need a subscription; each step in the path is marked Free or Subscribers.
What should I do when the path stalls?
Four blockers repeat: an agent that reads the wrong files, an MCP server that fails to connect because an environment variable is missing, a quality drop after a model or effort change, and headless runs that hang on an approval prompt. Each has a short recovery on this page.