Skip to content

Career ladders and performance reviews when output is cheap

Career ladders for engineers who work with coding agents reward outcomes, specification, verification and harness contributions instead of code volume. Pull request counts, lines changed and token spend turn into Goodhart targets as soon as they reach a review form, so they stay on team dashboards. Promotion evidence is a packet of linked artifacts the engineer owned.

It is calibration week. One engineer merged 212 pull requests this half and tops the vendor’s usage leaderboard. Another merged 19, but wrote the acceptance suite and holdout scenarios that let the payments team stop reading every diff. Your ladder still says “delivers a high volume of quality code”, and a manager has pasted the leaderboard export into a promotion packet.

This page is for the CTO or VP Engineering who owns the ladder and the tech lead or manager who writes reviews. It gives you a competency matrix, the metrics that stay out of reviews, an evidence packet, a telemetry charter with the settings that enforce it, and a one-cycle rollout. Hiring-stage role profiles are in hiring and interviewing for agentic engineering.

Why do career ladders break when agents write the code?

Section titled “Why do career ladders break when agents write the code?”

Most ladders were written when output was expensive, so “delivers a large volume of high-quality code” was a fair proxy for impact. Agents removed the cost of the first half of that sentence and made the second half harder to see. The evidence points the same way:

  • Output rose faster than delivery. Faros AI’s AI Engineering Report 2026 (April 2026, telemetry from 22,000 developers) reports merge rate per developer up 16.2%, with median time in review up 441.5%, incidents per pull request up 242.7% and code churn up 861%. Rewarding the first number rewards the cause of the other three.
  • The work moved to review. In Anthropic’s study of its own engineers (2025-12-02, a vendor’s self-report from 132 survey respondents and 53 interviews), some describe their work shifting “70%+ to being a code reviewer/reviser rather than a net-new code writer.” A commit count does not see that work.
  • AI share describes the industry, not the person. DX reports an average of 51.9% AI-authored code across more than 400 companies in Q2 2026 (Justin Reock, 2026-06-17, self-reported). An individual’s share shows tool habits, not judgment.
  • Self-assessment of AI speed is unreliable. In METR’s 2025 randomized trial (2025-07-10, early-2025 tools), experienced open-source developers took 19% longer with AI and still believed it had sped them up by 20%. “I shipped much more with agents” is a hypothesis, not evidence.

The fix is not a new metric but a ladder that names the work that now separates engineers, and a review that asks for evidence of it. Some old signals need a replacement, not only a removal: “writes clean code” becomes designing checks that keep code clean without reading it (fitness functions, lint and type gates), and depth shown in hard code becomes depth shown in specs, oracle design and incident diagnosis. Volume, speed and review counts are covered by the Goodhart table below.

The competency matrix for agentic engineering

Section titled “The competency matrix for agentic engineering”

Use the matrix to replace your ladder’s technical rows. Map “Engineer” to your mid-level band, “Senior” to senior, and “Staff and above” to staff, principal and equivalent. For entry-level expectations, see growing junior developers when agents write the code.

DimensionEngineerSenior engineerStaff engineer and above
OutcomesShips scoped changes that meet their acceptance criteria and names the outcome each servedOwns an outcome metric for a feature or service (lead time, error rate, adoption) and moves itDecides which outcomes a group pursues, and stops agent work that does not move them
SpecificationWrites acceptance criteria before the agent runs; agents rarely misread themWrites specs and short ADRs that other engineers’ agents implement; spec-caused rework is rare and tracedSets the spec and ADR standard for a group and resolves cross-team ambiguity before it reaches an agent
VerificationProves each change with tests and an evidence bundle; can explain the change without the diff openDesigns oracles (acceptance tests, property tests, holdout scenarios) and raises oracle strength for a module; catches defects in review before mergeDecides which loops may move to evidence-only review, owns the verification strategy, and rolls trust back when a gate misses
Harness and leverageReports and fixes harness friction: a flaky check, a misleading rule, a missing commandShips skills, hooks, checks or context files that other engineers adopt and that measurably cut reworkBuilds harness products used across teams (evals, policy, environments) and retires what does not help
Operation and ownershipDebugs their own changes in production without an agent when neededLeads incidents involving agent-written code and writes reviews that change a gate, not only a lineIs accountable for a system’s reliability and for the autonomy level its agents run at
People and craftKeeps unassisted skills current (debugging, reading unfamiliar code) and asks for review-to-learn sessionsRuns review-to-learn and spec reviews for juniors; interviews with the agent-allowed loopGrows seniors, shapes the hiring loop and the ladder itself

The matrix shares its vocabulary with the interview rubric in hiring for agentic engineering, so people are hired and promoted on the same evidence.

Two rules keep the matrix honest:

  1. Accountability does not transfer to the agent. The engineer owns what their agents shipped, including the parts nobody read. An engineer who cannot explain or roll back their agent’s change has not met the Verification row at any level.
  2. Leverage counts only when others use it. A skill, hook or check earns credit when another team adopts it and the rework it targets goes down.

Every claim in a review needs an artifact a calibration committee can open; the counts in the next section never qualify.

DimensionEvidence that counts
OutcomesThe team’s outcome metric before and after, with the changes that moved it
SpecificationSpecs, acceptance criteria and ADRs, plus the clarifications they needed after hand-off
VerificationOracles added with mutation score or holdout pass rate before and after; defects caught before merge; escaped defects traced to the engineer’s checks
Harness and leverageThe artifact, who adopted it, and the rework or review time it removed
Operation and ownershipIncident reviews led, rollbacks run, gates added after an incident
People and craftJuniors’ progress on their own plan, review-to-learn notes, interview feedback quality

Goodhart’s law, in Marilyn Strathern’s phrasing: “When a measure becomes a target, it ceases to be a good measure” (Marilyn Strathern, “Improving ratings: audit in the British University system”, European Review 5(3), 1997). Agents shorten the time from target to gamed metric, because inflating a count now costs a prompt.

MetricHow agents make it easy to inflateWhat it is still good forIn an individual review?
Merged pull requests, commitsSplit work into many small agent pull requestsTeam throughput, read beside change fail rateNever
Lines added or changed, accepted linesMore lines is often worseNothing on its ownNever
Tokens, spend, cost per personA target rewards loops left running; a cap punishes expensive verificationTeam cost per accepted change; budget managementNever
”% AI-written”, sessions, active timeRoute every task through an agentTeam-level utilization and frictionNever
Vendor leaderboard rankCombines the counts aboveFinding people who can share workflowsNever
Test count, line coverageTests that assert little or mirror the implementationA CI floorOnly as a linked oracle with its mutation score
Reviews performed, approvalsApprove faster, read lessBalancing reviewer loadNever; defects caught before merge count instead

The vendor leaderboard is not hypothetical. Claude Code’s Team and Enterprise analytics dashboard ranks the top 10 users by pull requests or lines of code and offers an Export all users CSV (Claude Code analytics docs, checked 2026-09-26). Anthropic presents it as a way to find power users who can help others; use it for exactly that, never in a promotion packet.

The team-level definitions (accepted change rate, cost per accepted change, review load) are in DORA, SPACE, DX Core 4 and AI measurement, which shares this page’s rule: no metric is reported per engineer.

What goes into a review instead: the evidence packet

Section titled “What goes into a review instead: the evidence packet”

Ask every engineer for one evidence packet per review cycle (not to be confused with a PR’s evidence bundle, which proves a single change). It replaces the self-review’s list of things built. The engineer writes it; the manager checks every link.

# Evidence packet: <name>, <cycle, e.g. 2026-H2>
## Outcomes (1–3)
- Outcome: <metric the team tracks, e.g. checkout p95 latency>
Before → after: <values, dates, source dashboard link>
My part: <spec / oracle / change links>
## Specification
- <link to spec, acceptance criteria or ADR> — implemented by: <who or which agent loop>
Clarifications needed after hand-off: <count, links>
## Verification
- Oracle added or strengthened: <link> — mutation score or holdout pass rate before → after
- Defects caught before merge: <PR comment links>
- Escaped defects in code I verified: <incident or bug links, with what the gate missed>
## Harness and leverage
- <skill / hook / check / context file link> — adopted by: <teams or people>
Effect: <rework, review time or failure rate before → after, with source>
## Operation and ownership
- Incidents led or rollbacks run: <links>; gate changed afterwards: <link>
## People and craft
- Review-to-learn or spec reviews run: <links or notes>; who grew and how
- Unassisted practice this cycle: <what, and what I learned>
## What I would do differently
- <one or two sentences, with a link if one exists>

Keep it to two pages; a staff packet is mostly outcomes, verification strategy and leverage.

How do you keep telemetry from becoming surveillance?

Section titled “How do you keep telemetry from becoming surveillance?”

You need agent telemetry for cost, harness friction and team metrics. Cut per person, it becomes a monitoring system, and engineers who suspect it feeds their review stop reporting failed runs.

There is also a legal side. A model that scores, ranks or allocates work to engineers from agent telemetry is an Annex III point 4(b) use under the EU AI Act (monitoring and evaluating the performance and behaviour of workers), high-risk from 2 December 2027 under the 2026 omnibus deferral (secondary source; confirm on EUR-Lex); see the EU AI Act for companies building with agents. Employee monitoring also raises GDPR questions today, and German works councils co-determine technical systems suited to monitoring employees’ performance (BetrVG §87(1) no. 6). Take these to your counsel; this is not legal advice. Adopt and publish the charter below before you switch telemetry on.

# Agent telemetry charter — <company>, effective <date>
1. Purpose. We collect coding-agent telemetry to manage cost, find harness friction,
and measure delivery at team level. We do not collect it to evaluate individuals.
2. Aggregation. Every dashboard and report shows teams of five or more people.
A group smaller than five is merged into its parent group.
3. Identity. Personal identifiers (email, account IDs, user and session IDs) are removed
in the collector before storage, except in the licence-administration store (rule 5).
4. Content. Prompt text, assistant responses, tool arguments and tool output are
not collected.
5. Per-person access. An engineer can see their own usage. Licence administrators
can see per-seat usage to manage seats and spend limits, and nothing else.
6. Reviews. No telemetry, vendor leaderboard or analytics export enters a
performance review, promotion packet, calibration or performance plan.
7. No automated judgement. No model or rule scores, ranks or assigns work to
individual engineers from this data.
8. Retention. Raw events are kept 90 days; team aggregates are kept 24 months.
9. Changes. Any change to this charter is announced 30 days ahead and agreed with
<employee representatives / works council, where one exists>.
10. Audit. <Role> checks rules 2–7 each quarter and publishes the result internally.

The charter’s numbers are defaults, not legal thresholds; set them with your counsel.

The charter only holds if the configuration matches it.

What it sends. For a signed-in user, Claude Code attaches user.email, user.account_uuid, user.account_id and organization.id to every metric and event it exports to your OpenTelemetry endpoint. Every event also carries user.id and session.id; on a Claude apps gateway session user.id is the IdP subject and user.groups lists the person’s IdP groups. Prompt text, assistant responses, tool arguments, tool output and full API bodies are off by default and gated by OTEL_LOG_USER_PROMPTS, OTEL_LOG_ASSISTANT_RESPONSES, OTEL_LOG_TOOL_DETAILS, OTEL_LOG_TOOL_CONTENT and OTEL_LOG_RAW_API_BODIES (Claude Code monitoring docs, checked 2026-09-26 against v2.1.283). Charter rule 4 needs all five unset in managed settings. OTEL_LOG_ASSISTANT_RESPONSES inherits OTEL_LOG_USER_PROMPTS when unset, so leave both unset (or set both to 0 explicitly).

Configure it centrally. Claude Code ignores the exporter variables in a repository’s .claude/settings.json, so set them in managed settings, tagged by team, with account IDs dropped from metrics:

{
"env": {
"CLAUDE_CODE_ENABLE_TELEMETRY": "1",
"OTEL_METRICS_EXPORTER": "otlp",
"OTEL_LOGS_EXPORTER": "otlp",
"OTEL_EXPORTER_OTLP_PROTOCOL": "grpc",
"OTEL_EXPORTER_OTLP_ENDPOINT": "http://collector.example.com:4317",
"OTEL_RESOURCE_ATTRIBUTES": "team.id=payments,cost_center=eng-123",
"OTEL_METRICS_INCLUDE_ACCOUNT_UUID": "false",
"OTEL_METRICS_INCLUDE_SESSION_ID": "false"
}
}

Distribute a different team.id per team. user.email and user.id are still attached, and events still carry session.id, so remove them in the collector (below).

The analytics dashboard. Every Admin and Owner can view the leaderboard and the Export all users CSV. Keep those role lists short, and keep the export and the per-user spend report out of review tooling.

Whichever tools you run, strip identity in the OpenTelemetry Collector before the data reaches storage. The attributes processor removes the keys from metrics, logs and traces. It ships in the OpenTelemetry Collector contrib distribution, not the core otelcol build, so run this with otelcol-contrib or a distribution that includes the processor. It also leaves resource attributes alone: if a tool puts identity in the resource block, add a resource processor with the same delete actions. This is a complete collector configuration; replace the otlphttp endpoint with your storage backend:

receivers:
otlp:
protocols:
grpc: {}
http: {}
processors:
attributes/drop-identity:
actions:
- key: user.email
action: delete
- key: user.account_uuid
action: delete
- key: user.account_id
action: delete
- key: user.id
action: delete
- key: user.groups
action: delete
- key: session.id
action: delete
- key: conversation.id
action: delete
- key: turn.id
action: delete
batch: {}
exporters:
otlphttp:
endpoint: https://storage.example.com
debug: {}
service:
pipelines:
metrics:
receivers: [otlp]
processors: [attributes/drop-identity, batch]
exporters: [otlphttp]
logs:
receivers: [otlp]
processors: [attributes/drop-identity, batch]
exporters: [otlphttp]

Delete user.id even though it holds no email: it is stable for an installation, or the IdP subject on a gateway session, so it is a per-person pseudonym that lets anyone rebuild a leaderboard. Codex’s conversation.id is the same kind of per-session key, and its turn.id is a finer one (one value per turn of one person’s session), so both go: charter rule 3 removes session IDs, and a turn ID is a session ID with a counter. If you need session counts per team, use action: hash for session.id instead of delete.

To check it, add debug to the exporters list of both pipelines, send one session through, and confirm that no user.*, session.id, conversation.id, turn.id or account attribute survives, and that team.id does. Repeat the check after every agent upgrade: new versions add attributes.

Roll out the new ladder in one review cycle

Section titled “Roll out the new ladder in one review cycle”

Remove the bad incentive first, then give people the new evidence format, then calibrate against it.

  1. Announce what stops now. Before the next cycle opens, state in writing that pull request counts, lines, tokens, AI share and leaderboard positions will not appear in any review, packet or calibration.

  2. Publish and enforce the telemetry charter. Agree it with employee representatives where you have them, and ship the collector and managed-settings changes.

  3. Rewrite the technical rows of the ladder. Replace them with the matrix, adjusted to your levels; the ladder-audit prompt below flags every volume-based phrase.

  4. Pilot the evidence packet with two teams, one that runs agents heavily and one that does not. What engineers found hard to evidence shows where the process hides work.

  5. Calibrate with the checklist below. The calibration chair reads packets before the meeting and stops any discussion that falls back on counts.

  6. Audit the cycle. Run the checks in the next section, publish the aggregate results, and fix any dimension that produced no evidence.

The chair reads this aloud at the start of every calibration meeting.

  • Every rating cites at least one linked artifact per dimension discussed.
  • No one quotes a count of pull requests, commits, lines, tokens, sessions or a leaderboard position.
  • For each “exceeds” rating, the committee names the outcome and who else benefited.
  • Review and verification work is discussed for every senior and staff engineer.
  • Escaped defects are discussed as gate failures first: “which check should have caught this?”
  • Harness contributions cite adoption by someone else.
  • Light and heavy agent users are rated on the same dimensions and evidence.

How do you know the new ladder rewards the right work?

Section titled “How do you know the new ladder rewards the right work?”

A ladder’s output is who gets promoted. Check that output every cycle rather than trusting the wording.

CheckHowHealthy resultOwner
Evidence coverageSample 10 packets and count claims without a working linkFewer than one unlinked claim per packetCalibration chair
Volume leakagePeople analytics correlates ratings with merged pull request counts on pseudonymised data and reports only the coefficientWeak or none; a strong positive one means volume is still rewardedHead of people analytics
Dimension balanceTally which dimensions each promotion case citedVerification and harness work appear in senior and staff promotions, not only outcomesVP Engineering
Team outcomesChange fail rate, time in review and accepted change rate from the metrics panelStable or betterTech leads
TrustAnonymous survey item: “I can report a failed agent run or a reverted change without it hurting my review” (five-point scale)Most answers 4 or 5, and not fallingEngineering managers
Charter complianceCollector check for identity attributes; dashboard access reviewNo personal identifiers outside licence administrationPlatform or security team

The VP Engineering signs off the ladder and the cycle audit; the people partner and any employee representative body co-sign the charter.

What goes wrong when you rewrite a career ladder for agents?

Section titled “What goes wrong when you rewrite a career ladder for agents?”

The leaderboard reaches calibration anyway. A manager brings the vendor export “for context”. Recovery: the chair stops the discussion, the export leaves the packet, and charter rule 6 goes into the next invitation.

Harness busywork replaces code volume. Skills nobody installs appear; credit harness work only with adoption and a measured effect.

Test counts become the new line counts. Recovery: accept only oracle strength (mutation score, holdout pass rate) or defects caught, and protect the oracle from agent edits.

Senior reviewers are rated down for “low output”. Recovery: put defects caught before merge in the packet and show reviewer load from running the review queue.

Engineers switch telemetry off or route around it. Recovery: publish the charter, the collector configuration and the audit result. One breach of the charter undoes it.

Specs get longer, not better. Judge specs by what happened after hand-off, never by length.

Juniors cannot show evidence at their level. Recovery: give entry levels a smaller packet built around their development plan from growing junior developers.

A model starts writing the ratings. A manager asks an assistant to rank the team from telemetry. Recovery: prohibit it in the AI usage policy; it is the automated worker evaluation the EU AI Act classifies as high-risk. The review-check prompt above checks a review; it never scores one.

Change the hiring loop next, so new engineers are selected on the same dimensions they will be promoted on.

Frequently asked questions

Should pull request counts or token usage appear in engineers' performance reviews?

No. Once agents write the code, both can be raised without delivering anything: more, smaller agent pull requests, or agent loops left running. Keep them at team level, next to a stability metric, and judge individuals on linked evidence of outcomes, specs, verification and harness work.

What should a career ladder reward when agents write most of the code?

Outcomes the engineer owned, specifications other people's agents could implement without misreading, verification that lets a loop run on evidence instead of line-by-line reading, harness contributions other engineers adopt, and ownership of incidents and rollbacks.

How do you use agent telemetry without it becoming surveillance?

Write a telemetry charter before you collect anything: team-level aggregates only, a minimum group size, prompt content off, per-person data visible only to the engineer and to licence administration, and a written rule that no telemetry enters a review, promotion or performance plan.