Developer Scorecard answer key
The Developer Scorecard answer key maps each of the 25 questions in the Developer Scorecard to the page that moves that answer, and names the repository evidence a 3-point answer needs. It is written for a developer who has a result and wants to know which three gaps to close first, and how to prove each fix.
Your result says Level 2 · Paired, 31 points out of 75. You already run Claude Code or Codex every day, yet Stage 4 and Stage 6 are close to zero, and 25 links on the results page do not tell you where to start. This key picks the first three questions, sends each one to its canonical page, and tells you which file, job, or run earns the higher answer.
What does each Level 1–4 band mean for a developer?
Section titled “What does each Level 1–4 band mean for a developer?”The bands borrow their names from the autonomy ladder, but they score something narrower: the harness around your own work, as you report it. The ladder describes how a loop actually ships merged change. Check your band against the ladder level page that matches it, and answer for the repository you work in most, not your best side project.
| Band | Points | What it describes | First move |
|---|---|---|---|
| Level 1 · Assisted | 0–18 | AI is a faster search engine: pasted snippets, an agent that never sees the repository. | Stage 1: pick one agentic harness, commit an instruction file, and give it one verify command. See Levels 1–2. |
| Level 2 · Paired | 19–37 | One agent, one task, and you watch every step. Instruction files and a few MCP servers exist. | Stage 4: make the session prove its own work before you look at it. |
| Level 3 · Review manager | 38–56 | Agents write, you review diffs. Automated PRs and extensions are in place. | Stages 2 and 5: move your judgment into the spec and the review contract. See Level 3. |
| Level 4 · Spec manager | 57–75 | You write the spec and the tests, then check what passed. | Close the Stage 6 loop and teach the setup to your team. The scorecard has no Level 5 band; running the software factory is a per-loop decision. |
How do you pick the first three questions to fix?
Section titled “How do you pick the first three questions to fix?”- Divide each section’s points by its maximum. A 7 out of 18 in Stage 3 (39%) is a bigger gap than 5 out of 9 in Stage 2 (56%), although Stage 3 has more points.
- Take the two sections with the lowest share. In each, pick the question where your answer is furthest from 3 points.
- If Stage 4 is below half and is not already one of the two, add its weakest question. Otherwise take the weakest question of the third-lowest section. Worktrees (question 13) and PR loops (question 19) multiply the output you have to check; without a Stage 4 feedback loop, that output lands on you as diffs to read.
- For each question, write the target answer and the artifact that will prove it in the plan file below.
- Re-take the scorecard after 30 days and compare section shares, not the band. One question moving from 1 to 3 points rarely changes the band, but it shows in its section.
Commit the plan next to the code it describes, so the agent can read it and your team can review it:
# docs/scorecard-plan.yaml — one entry per question you are working on- question: 14 # How does the agent verify its own work before human review? current_points: 1 target_points: 3 target_answer: "Quantifiable targets, an automated fix loop, and a verifier subagent in fresh context" evidence: "scripts/verify.sh named in AGENTS.md; three merged PRs whose description links a green verify run" check_on: 2026-10-26Which page answers each Developer Scorecard question?
Section titled “Which page answers each Developer Scorecard question?”Each table gives the question, the page to read first and a second one, and the evidence a 3-point answer needs. Questions 4, 5, 8, 10, 14, 15, and 23 link their topic page directly, because their per-question pages are being merged into it.
Stage 1 — Plan: intent and foundations (questions 1–4, 12 points)
Section titled “Stage 1 — Plan: intent and foundations (questions 1–4, 12 points)”| # | Question | Read | Evidence for 3 points |
|---|---|---|---|
| 1 | How does your primary AI coding setup handle repository work? | Choose a primary harness · Tool map | Versioned instructions, plan and artifact gates, bounded tool permissions, and a reproducible verify command, all in the repository |
| 2 | Which plan best matches your actual workload? | Size an AI plan · Plan prices and billing | A quarterly review against representative tasks, concurrency, interruptions, and cost per completed task |
| 3 | How do you choose a model and runtime for each task? | Model routing by evidence · Models hub | Versioned evals per task class that record quality, latency, failures, and completed-work cost |
| 4 | How do you capture problem statements before any code or plan (intent.md)? | Plan: capture intent.md · The artifact chain | A committed intent.md template, and each feature’s intent.md reviewed by its product owner before design starts |
Stage 2 — Design: specs and policy skills (questions 5–7, 9 points)
Section titled “Stage 2 — Design: specs and policy skills (questions 5–7, 9 points)”| # | Question | Read | Evidence for 3 points |
|---|---|---|---|
| 5 | How do you turn intent into a specification (spec.md)? | Design: write spec.md · Spec-driven development | spec.md produced in one session under the policy skills, with every flagged concern resolved before build |
| 6 | How do you encode and enforce policies (brand, security, compliance, UX)? | Agent skills as policy · Building your own skills | Versioned skill folders in the repository, each with a named owner and test fixtures |
| 7 | How do you prototype and validate UI before implementing code? | Visual acceptance · Design tools MCP | An inspectable mock from an approved design tool, iterated and accepted before code generation |
Stage 3 — Build: plans, rules, and parallelism (questions 8–13, 18 points)
Section titled “Stage 3 — Build: plans, rules, and parallelism (questions 8–13, 18 points)”| # | Question | Read | Evidence for 3 points |
|---|---|---|---|
| 8 | How often do you start in plan mode and commit plan.md? | Build: plan.md, then implement · The artifact chain | plan.md committed before implementation on every non-trivial change, and updated when the work departs from it |
| 9 | How is codebase knowledge configured (CLAUDE.md, AGENTS.md, rules)? | AGENTS.md and CLAUDE.md · Pruning instruction files | A short root file plus path-scoped rules, and a record that a mistake repeated twice became a rule |
| 10 | Which MCP servers do you use in a typical week? | The MCP starter stack · MCP security | Three servers in weekly use, each tied to a recurring task and scoped to least privilege |
| 11 | How do you use scoped subagents with restricted tools? | Scoped subagents · The harness | Verifier, simplifier, and explorer definitions with explicit tool restrictions and their own context |
| 12 | How do you use hooks as deterministic guardrails? | Hooks as guardrails · Claude Code hooks | Hooks that block protected paths and secret leaks and run formatting and policy checks, each with a test |
| 13 | How do you run parallel sessions without file collisions? | Parallel agents in worktrees · Parallel agent tools | A script that runs 2–4 worktrees with an explicit base revision, ports, local state, verification, and cleanup |
Stage 4 — Test: feedback loops and continuous evals (questions 14–17, 12 points)
Section titled “Stage 4 — Test: feedback loops and continuous evals (questions 14–17, 12 points)”| # | Question | Read | Evidence for 3 points |
|---|---|---|---|
| 14 | How does the agent verify its own work before human review? | Test: give the session a feedback loop · How strong is your oracle? | One verify command with quantifiable targets, run in a loop, plus a verifier subagent in fresh context |
| 15 | How do you fix bugs with agents? | Test-driven development with AI · Protecting the oracle | A failing test committed before the fix, and a hook that blocks test-file edits while the fix runs |
| 16 | How do you verify visual and end-to-end browser behaviour? | Agent-driven browser verification · End-to-end test automation | A browser run against the acceptance criteria with retained artifacts, and critical paths promoted to CI tests |
| 17 | How do you regression-test your agent configuration? | Continuous evals · Evals for coding agents | A CI eval job that runs on every PR touching instructions, skills, or hooks, with production incidents added as cases |
Stage 5 — Deploy: review loops and production gates (questions 18–21, 12 points)
Section titled “Stage 5 — Deploy: review loops and production gates (questions 18–21, 12 points)”| # | Question | Read | Evidence for 3 points |
|---|---|---|---|
| 18 | How are pull requests reviewed before merge (REVIEW.md)? | Layered PR review · Reviewing an agent’s PR | A REVIEW.md with separate bug, security, and spec-compliance passes, findings graded by severity, and human sign-off |
| 19 | How are review comments and failed CI checks resolved? | Bounded PR review-fix loop · The evidence bundle | A loop with an attempt limit that fixes verified causes, reruns required checks, and stops for a named approver |
| 20 | What stops an agent from deploying unvetted changes to production? | Production approval gate · Permissions and sandboxes | No production credentials in the agent’s environment, direct push blocked, and an explicit human release approval |
| 21 | How are rollbacks managed after a regression? | Rehearsed rollback · Progressive delivery | A single-command rollback job with a dated rehearsal record and a deterministic recovery check |
Stage 6 — Maintain: closed loop and incident response (questions 22–25, 12 points)
Section titled “Stage 6 — Maintain: closed loop and incident response (questions 22–25, 12 points)”| # | Question | Read | Evidence for 3 points |
|---|---|---|---|
| 22 | How do you monitor production and detect actionable anomalies? | Anomaly detection and diagnosis · Monitoring and observability | Version-controlled thresholds, read-only diagnosis, retained evidence, and routing by approval tier |
| 23 | What happens when a production alert fires? | Maintain: close the loop · Failure taxonomy | The alert starts a headless agent that files an evidence-backed intent.md into the triage queue |
| 24 | How do you scan for deep vulnerabilities and architectural drift? | Scheduled security scans · Security scanning | Scheduled scanner and agent runs that retain evidence, route patches through PR gates, and add regression evals |
| 25 | How does AI take part in incident response and on-call? | Scoped AI incident triage · Incident response | A scoped agent gathers evidence and drafts the timeline; a named human approves actions, recovery, and the post-mortem |
How do you prove a higher answer without reading every diff?
Section titled “How do you prove a higher answer without reading every diff?”Every 3-point answer above names an artifact, not a habit: a committed file, a CI job, a hook with a test, a retained browser run, a rehearsal record. Claim the higher answer only when you can link that artifact. For the build and test questions, the strongest proof is a check that fails when the code is wrong. Reading evidence instead of code explains why that replaces line-by-line review, and a teammate or your tech lead confirms the link before you re-take the scorecard.
An agent can collect the evidence for you. Start it in the mode each tool provides for read-only work, from the repository you are scoring, rather than trusting the prompt alone to keep it read-only.
Start the session in plan mode, which explores without editing files: claude --permission-mode plan (checked in claude --help, v2.1.285). Paste the prompt below.
Run the prompt non-interactively in the read-only sandbox: codex exec --sandbox read-only - < scorecard-evidence.txt (checked in codex exec --help, v0.157.1). The prompt spans several lines and contains double quotes, so save it to scorecard-evidence.txt first: pasted inside quotes as an argument, it breaks shell quoting. The - tells Codex to read the instructions from stdin.
Switch the agent to Plan Mode, which produces a plan before writing any code, then paste the prompt below.
The prompt text is the same in all three tools.
A verify command that never fails proves nothing. Before you claim 3 points on question 14 or 15, test the oracle in a throwaway worktree. In a terminal, run git worktree add ../scorecard-probe HEAD, start your agent inside ../scorecard-probe with ordinary edit permissions, and paste the second prompt. A fresh worktree has none of your untracked files, such as node_modules or .env, so the prompt installs dependencies first; without that, step 1 can fail for reasons unrelated to the oracle. Remove the worktree afterwards with git worktree remove --force ../scorecard-probe.
Every uncaught bug is a gap in your oracle, and the fix belongs on how strong is your oracle?, not in your score.
When does the Developer Scorecard result mislead?
Section titled “When does the Developer Scorecard result mislead?”- You scored habits, not artifacts. “I usually plan first” earns the same answer as a committed
plan.mdif you let it. Recovery: re-score every answer you cannot back with a linked file, job, or run one point lower. That lower score is your real starting point. - Your setup lives on your laptop. Skills, hooks, and MCP servers in your home directory help you, but a fresh clone gets none of them. Recovery: answer for what a teammate’s clone of the repository gets, and move what you rely on into the repository.
- Question 10 became a server count. Three connected MCP servers you rarely call still score 3 points, and each one loads its tool definitions into the context. Recovery: count only servers you called this week for a recurring task, and remove the rest, as reducing MCP token cost shows.
- Stage 3 parallelism rose before Stage 4. Four worktrees without a feedback loop produce four times the diffs to read. Recovery: cap parallel sessions at one until questions 14 and 15 reach 3 points.
- The band became the goal. A jump from Level 2 to Level 3 in one month usually means answers moved faster than artifacts. Recovery: re-check each changed answer against its evidence before you tell anyone about the new band.