Developer onboarding — prove one safe workflow
Developer onboarding time, as CTO Scorecard Q19 scores it, is how long a new engineer takes from a clean machine to a verified sample change in one approved agent workflow. The maximum score needs a repeatable onboarding fixture that passes within one working day, with no shared secrets.
This page is for the tech lead who owns onboarding and the CTO who answers Q19. The situation: a new engineer joined on Monday, the install script finished in ten minutes, and on Thursday they are still asking in chat why the agent ignores the repository conventions and which token the integration tests need. The install worked. The workflow did not.
What you get from the onboarding fixture
Section titled “What you get from the onboarding fixture”- The Q19 scoring table, and how it relates to question 4 of the Tech Lead Scorecard.
- A fixture card template that defines the task, the evidence and the clock, ready to fill in.
- A six-step run with the exact bootstrap checks for Claude Code and Codex, and what to check by hand in Cursor.
- A timing sheet that separates hands-on time from waiting time, so every delay has an owner.
- Three copy-paste prompts, a way to verify the run without watching over the engineer’s shoulder, and the failure modes that make onboarding look solved when it is not.
How does CTO Scorecard Q19 score onboarding time?
Section titled “How does CTO Scorecard Q19 score onboarding time?”The scorecard asks “How long is ‘time to productive AI workflow’ for a new engineer?” and scores four answers.
| Points | Answer | Evidence that earns it |
|---|---|---|
| 0 | We don’t measure, each engineer figures it out | None. |
| 1 | A week or more with setup docs | A wiki page and a buddy. |
| 2 | One to two days with a bootstrap script and docs | A script that installs tools, plus docs that describe the rest. |
| 3 | A representative onboarding fixture passes within one working day: safe bootstrap, repository context, a sample task, required checks, and no shared secrets | A fixture card, timing sheets from the last two runs, and the evidence bundle each run produced. |
The step from 2 to 3 is the one that matters. A bootstrap script proves that tools install. A fixture proves that a new engineer can plan, change, verify and hand off work with an agent, which is what “productive” means on this page.
The Tech Lead Scorecard asks the same question from the team side (“When a new dev joins, how fast are they productive with your AI setup?”) and gives its top score to a one-command bootstrap plus an onboarding path. A passing fixture satisfies both: the bootstrap is its first half, the sample task its second.
Choose a fixture task that exercises the real path
Section titled “Choose a fixture task that exercises the real path”The fixture is one task that every new engineer runs in their first day. It has to touch the same repository, gates and approval boundaries as real work while staying safe to repeat.
| Choose a task that… | Avoid a task that… |
|---|---|
| Changes one small, documented behavior in a production repository | Lives in a toy repository nobody ships from |
| Needs a new or changed test that fails before the change | Can pass with no test at all |
| Runs the same lint, type and test gates as every pull request | Skips CI “because it is onboarding” |
| Ends in a draft pull request with an evidence bundle | Ends with “it works on my machine” |
| Touches no production data, deploy or secret beyond the engineer’s own sign-in | Needs a shared service token or production access |
| Stays valid for months (a stable module, a known edge case) | Depends on this sprint’s backlog |
Keep a pool of three to five such tasks, reset by a script, so two engineers starting the same week do not copy each other’s answer. Run the fixture in the tool your tooling policy approves; do not require all three tools when your team standardizes on one.
Write the fixture card before the first run
Section titled “Write the fixture card before the first run”The card fixes what counts as a pass, so the timing means the same thing every run. Store it in the repository beside the onboarding docs.
ONBOARDING FIXTURE CARD — payments-apiOwner: Tech lead, payments teamApproved tools: Claude Code, Codex or Cursor (per tooling policy v3)Preconditions: Managed individual account, repository access and secret-store access provisioned BEFORE day one (lead time recorded separately)Bootstrap: ./scripts/onboard.sh — pinned tool versions, idempotent, exits non-zero with a fix-it message on every failed checkContext loaded: AGENTS.md (or CLAUDE.md), architecture map docs/architecture.md, protected paths list, the approval boundaries in docs/policy/Sample task pool: docs/onboarding/tasks/ (5 tasks, reset by scripts/reset-fixture.sh)Pass criteria: 1. Plan accepted by the buddy before any edit 2. Diff limited to the files named in the plan 3. New test fails before the change and passes after 4. Lint, type check and test suite green in CI 5. Draft PR carries the evidence bundle and names the human gates 6. No credential in any file, commit, prompt or chat messageThe clock: Starts when the engineer opens the provisioned laptop; stops when the buddy accepts the draft PR against the pass criteriaTarget: Within one working day, hands-on plus waitingRecorded: Hands-on time, waiting time by cause, interventions, failed stepsRetest: After every change to onboard.sh, AGENTS.md or tool versions, and at least once a quarter by someone who did not write itThe clock starts after provisioning on purpose. Identity, repository access and secret-store access are prerequisites that should be ready before the first day; record their lead time on the timing sheet as its own line, so a slow access request is visible without hiding inside the fixture time.
Run the onboarding fixture in six steps
Section titled “Run the onboarding fixture in six steps”-
Provision identity before day one. Give the engineer a managed individual account for the approved tool, least-privilege repository access and their own entry in the secret store. Never a shared login or a team API key. For the account and secret model, see agent identity and secrets.
-
Run the bootstrap and its checks. The bootstrap installs pinned versions and then checks itself: tool on
PATH, signed in as the engineer, MCP servers connected, repository instructions present. Every failed check prints the fix. The per-tool commands are in the next section. -
Load and confirm repository context. The engineer asks the agent to summarize the repository rules (first prompt below) and compares the answer with the docs. An agent that cannot quote the protected paths or the required checks has not loaded the instructions, whatever the install log says. For what those instructions should contain, see shared agent rules.
-
Plan the sample task and get the plan accepted. The agent names files, acceptance criteria, tests, risks and stop conditions; the buddy accepts or corrects the plan before any edit. This is the first human gate the engineer learns.
-
Execute, then verify from a clean state. The agent writes the failing test first, makes the change, and runs the same gates CI runs. The engineer opens a draft pull request with the evidence bundle: commands, results, and each pass criterion mapped to its evidence. For the manifest format, see the evidence bundle.
-
Record the timing sheet and debrief for 15 minutes. Log hands-on time, waiting time by cause, every intervention, and every step the docs did not cover. Each gap becomes a ticket with an owner and a retest date.
Check the bootstrap in each tool
Section titled “Check the bootstrap in each tool”The bootstrap script should end by running these checks and failing on the first one that does not pass. They cost seconds and catch the failures a new engineer cannot diagnose alone.
Checked against Claude Code 2.1.283.
claude --version # compare with the version the fixture card pinsclaude doctor # installation and settings healthclaude auth status --text # signed in with the engineer's own accountclaude mcp list # project .mcp.json servers: connected, or pending approvalTwo traps. First, claude mcp list shows servers from the project’s .mcp.json as “Pending approval” until the engineer approves them; a pending server is not connected, so the sample task fails later with a missing tool. Second, Claude Code reads AGENTS.md only when the project has no CLAUDE.md, and only from v2.1.277 on the latest channel (v2.1.281 on Bedrock, Google Cloud, Microsoft Foundry, LLM gateways or with telemetry off); the stable channel was on 2.1.274 on 2026-09-26. If your repository keeps its rules only in AGENTS.md, keep a CLAUDE.md as well, or install the native build on the latest channel (claude install latest, with autoUpdatesChannel: "latest", the default for native installs; the Homebrew cask claude-code tracks stable, so use claude-code@latest), and let the context prompt in step 3 confirm the rules loaded.
Checked against Codex CLI 0.157.1.
codex --version # compare with the version the fixture card pinscodex doctor --json # installation, config, auth and runtime health (redacted)codex login status # signed in with the engineer's own accountcodex mcp list # external MCP servers Codex will connect toThe trap. Since Codex 0.150.0, an untrusted project does not supply its project-level AGENTS.md, and Codex hooks need persisted trust before they run. A new engineer who declines the trust prompt gets an agent with none of the repository’s rules and none of its guardrail hooks, and nothing fails loudly. Make trusting the repository an explicit bootstrap step, and let the context prompt in step 3 prove it worked. Keep the --json report from codex doctor in the evidence bundle.
Cursor’s documentation could not be reached to verify current setup details on 2026-09-26, so this tab names no command. Check by hand in the IDE that the engineer is signed in with their managed account, that the project’s Rules and MCP servers appear as loaded, and that the context prompt in step 3 quotes the repository rules correctly. That last check is tool-independent and is the one that counts.
A bootstrap check that only confirms the tool is installed is not enough. The failures that cost new engineers days are the silent ones: a server waiting for approval, instructions not loaded, a project not trusted.
Time the fixture so every delay has an owner
Section titled “Time the fixture so every delay has an owner”A single “took 2 days” number cannot be acted on. Split it, and each part points at a different owner.
| Line on the timing sheet | What it measures | Usual owner of a slow result |
|---|---|---|
| Provisioning lead time | Access request to account, repository and secret-store access ready | IT or platform team |
| Bootstrap time | Laptop opened to all bootstrap checks green | Platform or developer-experience owner |
| Context time | Bootstrap green to a correct summary of repository rules | Owner of AGENTS.md / CLAUDE.md |
| Task hands-on time | Plan written, change made, checks run | The engineer, with the fixture task’s author |
| Waiting time, by cause | Plan acceptance, CI queue, review, access tickets | Whoever owns that queue |
| Interventions | Each time someone helped outside the documented path | The docs or script that should have covered it |
Report the median of the last runs per repository, never a per-person ranking. Q19 evidence is the fixture card, the last two timing sheets and their evidence bundles.
Copy-paste prompts for the onboarding fixture
Section titled “Copy-paste prompts for the onboarding fixture”Run these in the approved tool, from the repository root, on the engineer’s first day. They work the same in Claude Code, Codex and Cursor.
Replace the task file name with one from your own pool. The last prompt does two jobs: it produces the evidence bundle, and it turns the engineer’s confusion into documentation tickets.
How do you verify the fixture without watching the engineer?
Section titled “How do you verify the fixture without watching the engineer?”The buddy should not sit beside the new engineer or read every line of the sample diff. The run is verified from its artifacts.
- The pass criteria are checked by systems. CI runs the same gates as every pull request. The failing-before, passing-after test proves the change does what the plan said.
- The plan is the human’s review point. The buddy reviews the plan against the task, then reviews the evidence bundle against the plan, not the diff line by line.
- Secret hygiene is scanned, not trusted. The repository’s secret scanner and pre-commit hooks run on the fixture branch; a hit fails the run, whatever else passed.
- Context is proven by the answer. The context prompt’s output, with quoted source lines, goes into the evidence bundle. If it misquotes a rule, the context step failed.
- A second person re-runs it. Once a quarter, someone who did not write the fixture runs it from a fresh machine or container, which catches knowledge the authors no longer notice they supply.
- A named owner signs off. The tech lead accepts each run against the card and owns the tickets its debrief produced.
What goes wrong when you measure onboarding time?
Section titled “What goes wrong when you measure onboarding time?”The bootstrap counts as the proof. The installer is green in ten minutes, and the score goes to 3 while engineers still cannot tell a sandbox from production. Recovery: stop the clock at the accepted draft pull request, not at the end of the script.
A shared token lives in the onboarding doc. It made day one faster, and it now sits in a wiki, a dotfile and several chat histories. Recovery: rotate the token today, give each engineer their own access through the secret store, and add a scanner rule for that token’s format.
The agent runs without the repository’s rules. A pending MCP server, an untrusted Codex project, or a Claude Code install on a channel that does not read AGENTS.md leaves the agent guessing, and the new engineer blames themselves. Recovery: make the context prompt a pass criterion and add the matching check to the bootstrap.
The buddy passes the fixture. Help in chat does not show up anywhere, and the run looks fast. Recovery: log every intervention on the timing sheet, and treat any run with an unlogged intervention as a failed run.
The fixture rots. Tool versions move weekly, the task’s module is refactored, and the next engineer hits errors the card does not mention. Recovery: retest after every change to the bootstrap, the instructions or pinned versions, and let the quarterly outside run catch the rest. For keeping pace with tool releases, see keeping a team current.
The task is too easy to teach anything. A typo fix passes in 20 minutes and proves nothing about planning or verification. Recovery: require a failing test and a plan acceptance in every task in the pool.
Waiting time disappears into the total. Three hours in a CI queue look like a slow engineer. Recovery: report waiting time by cause, and send each cause to the owner of that queue.
Where to go next with developer onboarding
Section titled “Where to go next with developer onboarding”Once new engineers pass the fixture, measure whether they keep using the workflow, and turn the gaps each run finds into shared knowledge. The prerequisite for a fast fixture is a repository an agent can work in; see making a codebase agent-ready.
For every other question on the scorecard, start from the CTO Scorecard answer key.