Skip to content

Developer onboarding — prove one safe workflow

Developer onboarding time, as CTO Scorecard Q19 scores it, is how long a new engineer takes from a clean machine to a verified sample change in one approved agent workflow. The maximum score needs a repeatable onboarding fixture that passes within one working day, with no shared secrets.

This page is for the tech lead who owns onboarding and the CTO who answers Q19. The situation: a new engineer joined on Monday, the install script finished in ten minutes, and on Thursday they are still asking in chat why the agent ignores the repository conventions and which token the integration tests need. The install worked. The workflow did not.

  • The Q19 scoring table, and how it relates to question 4 of the Tech Lead Scorecard.
  • A fixture card template that defines the task, the evidence and the clock, ready to fill in.
  • A six-step run with the exact bootstrap checks for Claude Code and Codex, and what to check by hand in Cursor.
  • A timing sheet that separates hands-on time from waiting time, so every delay has an owner.
  • Three copy-paste prompts, a way to verify the run without watching over the engineer’s shoulder, and the failure modes that make onboarding look solved when it is not.

How does CTO Scorecard Q19 score onboarding time?

Section titled “How does CTO Scorecard Q19 score onboarding time?”

The scorecard asks “How long is ‘time to productive AI workflow’ for a new engineer?” and scores four answers.

PointsAnswerEvidence that earns it
0We don’t measure, each engineer figures it outNone.
1A week or more with setup docsA wiki page and a buddy.
2One to two days with a bootstrap script and docsA script that installs tools, plus docs that describe the rest.
3A representative onboarding fixture passes within one working day: safe bootstrap, repository context, a sample task, required checks, and no shared secretsA fixture card, timing sheets from the last two runs, and the evidence bundle each run produced.

The step from 2 to 3 is the one that matters. A bootstrap script proves that tools install. A fixture proves that a new engineer can plan, change, verify and hand off work with an agent, which is what “productive” means on this page.

The Tech Lead Scorecard asks the same question from the team side (“When a new dev joins, how fast are they productive with your AI setup?”) and gives its top score to a one-command bootstrap plus an onboarding path. A passing fixture satisfies both: the bootstrap is its first half, the sample task its second.

Choose a fixture task that exercises the real path

Section titled “Choose a fixture task that exercises the real path”

The fixture is one task that every new engineer runs in their first day. It has to touch the same repository, gates and approval boundaries as real work while staying safe to repeat.

Choose a task that…Avoid a task that…
Changes one small, documented behavior in a production repositoryLives in a toy repository nobody ships from
Needs a new or changed test that fails before the changeCan pass with no test at all
Runs the same lint, type and test gates as every pull requestSkips CI “because it is onboarding”
Ends in a draft pull request with an evidence bundleEnds with “it works on my machine”
Touches no production data, deploy or secret beyond the engineer’s own sign-inNeeds a shared service token or production access
Stays valid for months (a stable module, a known edge case)Depends on this sprint’s backlog

Keep a pool of three to five such tasks, reset by a script, so two engineers starting the same week do not copy each other’s answer. Run the fixture in the tool your tooling policy approves; do not require all three tools when your team standardizes on one.

Write the fixture card before the first run

Section titled “Write the fixture card before the first run”

The card fixes what counts as a pass, so the timing means the same thing every run. Store it in the repository beside the onboarding docs.

ONBOARDING FIXTURE CARD — payments-api
Owner: Tech lead, payments team
Approved tools: Claude Code, Codex or Cursor (per tooling policy v3)
Preconditions: Managed individual account, repository access and secret-store
access provisioned BEFORE day one (lead time recorded separately)
Bootstrap: ./scripts/onboard.sh — pinned tool versions, idempotent, exits
non-zero with a fix-it message on every failed check
Context loaded: AGENTS.md (or CLAUDE.md), architecture map docs/architecture.md,
protected paths list, the approval boundaries in docs/policy/
Sample task pool: docs/onboarding/tasks/ (5 tasks, reset by scripts/reset-fixture.sh)
Pass criteria: 1. Plan accepted by the buddy before any edit
2. Diff limited to the files named in the plan
3. New test fails before the change and passes after
4. Lint, type check and test suite green in CI
5. Draft PR carries the evidence bundle and names the human gates
6. No credential in any file, commit, prompt or chat message
The clock: Starts when the engineer opens the provisioned laptop; stops when
the buddy accepts the draft PR against the pass criteria
Target: Within one working day, hands-on plus waiting
Recorded: Hands-on time, waiting time by cause, interventions, failed steps
Retest: After every change to onboard.sh, AGENTS.md or tool versions,
and at least once a quarter by someone who did not write it

The clock starts after provisioning on purpose. Identity, repository access and secret-store access are prerequisites that should be ready before the first day; record their lead time on the timing sheet as its own line, so a slow access request is visible without hiding inside the fixture time.

  1. Provision identity before day one. Give the engineer a managed individual account for the approved tool, least-privilege repository access and their own entry in the secret store. Never a shared login or a team API key. For the account and secret model, see agent identity and secrets.

  2. Run the bootstrap and its checks. The bootstrap installs pinned versions and then checks itself: tool on PATH, signed in as the engineer, MCP servers connected, repository instructions present. Every failed check prints the fix. The per-tool commands are in the next section.

  3. Load and confirm repository context. The engineer asks the agent to summarize the repository rules (first prompt below) and compares the answer with the docs. An agent that cannot quote the protected paths or the required checks has not loaded the instructions, whatever the install log says. For what those instructions should contain, see shared agent rules.

  4. Plan the sample task and get the plan accepted. The agent names files, acceptance criteria, tests, risks and stop conditions; the buddy accepts or corrects the plan before any edit. This is the first human gate the engineer learns.

  5. Execute, then verify from a clean state. The agent writes the failing test first, makes the change, and runs the same gates CI runs. The engineer opens a draft pull request with the evidence bundle: commands, results, and each pass criterion mapped to its evidence. For the manifest format, see the evidence bundle.

  6. Record the timing sheet and debrief for 15 minutes. Log hands-on time, waiting time by cause, every intervention, and every step the docs did not cover. Each gap becomes a ticket with an owner and a retest date.

The bootstrap script should end by running these checks and failing on the first one that does not pass. They cost seconds and catch the failures a new engineer cannot diagnose alone.

Checked against Claude Code 2.1.283.

Terminal window
claude --version # compare with the version the fixture card pins
claude doctor # installation and settings health
claude auth status --text # signed in with the engineer's own account
claude mcp list # project .mcp.json servers: connected, or pending approval

Two traps. First, claude mcp list shows servers from the project’s .mcp.json as “Pending approval” until the engineer approves them; a pending server is not connected, so the sample task fails later with a missing tool. Second, Claude Code reads AGENTS.md only when the project has no CLAUDE.md, and only from v2.1.277 on the latest channel (v2.1.281 on Bedrock, Google Cloud, Microsoft Foundry, LLM gateways or with telemetry off); the stable channel was on 2.1.274 on 2026-09-26. If your repository keeps its rules only in AGENTS.md, keep a CLAUDE.md as well, or install the native build on the latest channel (claude install latest, with autoUpdatesChannel: "latest", the default for native installs; the Homebrew cask claude-code tracks stable, so use claude-code@latest), and let the context prompt in step 3 confirm the rules loaded.

A bootstrap check that only confirms the tool is installed is not enough. The failures that cost new engineers days are the silent ones: a server waiting for approval, instructions not loaded, a project not trusted.

Time the fixture so every delay has an owner

Section titled “Time the fixture so every delay has an owner”

A single “took 2 days” number cannot be acted on. Split it, and each part points at a different owner.

Line on the timing sheetWhat it measuresUsual owner of a slow result
Provisioning lead timeAccess request to account, repository and secret-store access readyIT or platform team
Bootstrap timeLaptop opened to all bootstrap checks greenPlatform or developer-experience owner
Context timeBootstrap green to a correct summary of repository rulesOwner of AGENTS.md / CLAUDE.md
Task hands-on timePlan written, change made, checks runThe engineer, with the fixture task’s author
Waiting time, by causePlan acceptance, CI queue, review, access ticketsWhoever owns that queue
InterventionsEach time someone helped outside the documented pathThe docs or script that should have covered it

Report the median of the last runs per repository, never a per-person ranking. Q19 evidence is the fixture card, the last two timing sheets and their evidence bundles.

Copy-paste prompts for the onboarding fixture

Section titled “Copy-paste prompts for the onboarding fixture”

Run these in the approved tool, from the repository root, on the engineer’s first day. They work the same in Claude Code, Codex and Cursor.

Replace the task file name with one from your own pool. The last prompt does two jobs: it produces the evidence bundle, and it turns the engineer’s confusion into documentation tickets.

How do you verify the fixture without watching the engineer?

Section titled “How do you verify the fixture without watching the engineer?”

The buddy should not sit beside the new engineer or read every line of the sample diff. The run is verified from its artifacts.

  • The pass criteria are checked by systems. CI runs the same gates as every pull request. The failing-before, passing-after test proves the change does what the plan said.
  • The plan is the human’s review point. The buddy reviews the plan against the task, then reviews the evidence bundle against the plan, not the diff line by line.
  • Secret hygiene is scanned, not trusted. The repository’s secret scanner and pre-commit hooks run on the fixture branch; a hit fails the run, whatever else passed.
  • Context is proven by the answer. The context prompt’s output, with quoted source lines, goes into the evidence bundle. If it misquotes a rule, the context step failed.
  • A second person re-runs it. Once a quarter, someone who did not write the fixture runs it from a fresh machine or container, which catches knowledge the authors no longer notice they supply.
  • A named owner signs off. The tech lead accepts each run against the card and owns the tickets its debrief produced.

What goes wrong when you measure onboarding time?

Section titled “What goes wrong when you measure onboarding time?”

The bootstrap counts as the proof. The installer is green in ten minutes, and the score goes to 3 while engineers still cannot tell a sandbox from production. Recovery: stop the clock at the accepted draft pull request, not at the end of the script.

A shared token lives in the onboarding doc. It made day one faster, and it now sits in a wiki, a dotfile and several chat histories. Recovery: rotate the token today, give each engineer their own access through the secret store, and add a scanner rule for that token’s format.

The agent runs without the repository’s rules. A pending MCP server, an untrusted Codex project, or a Claude Code install on a channel that does not read AGENTS.md leaves the agent guessing, and the new engineer blames themselves. Recovery: make the context prompt a pass criterion and add the matching check to the bootstrap.

The buddy passes the fixture. Help in chat does not show up anywhere, and the run looks fast. Recovery: log every intervention on the timing sheet, and treat any run with an unlogged intervention as a failed run.

The fixture rots. Tool versions move weekly, the task’s module is refactored, and the next engineer hits errors the card does not mention. Recovery: retest after every change to the bootstrap, the instructions or pinned versions, and let the quarterly outside run catch the rest. For keeping pace with tool releases, see keeping a team current.

The task is too easy to teach anything. A typo fix passes in 20 minutes and proves nothing about planning or verification. Recovery: require a failing test and a plan acceptance in every task in the pool.

Waiting time disappears into the total. Three hours in a CI queue look like a slow engineer. Recovery: report waiting time by cause, and send each cause to the owner of that queue.

Where to go next with developer onboarding

Section titled “Where to go next with developer onboarding”

Once new engineers pass the fixture, measure whether they keep using the workflow, and turn the gaps each run finds into shared knowledge. The prerequisite for a fast fixture is a repository an agent can work in; see making a codebase agent-ready.

For every other question on the scorecard, start from the CTO Scorecard answer key.