Career ladders and performance reviews when output is cheap
Career ladders for engineers who work with coding agents reward outcomes, specification, verification and harness contributions instead of code volume. Pull request counts, lines changed and token spend turn into Goodhart targets as soon as they reach a review form, so they stay on team dashboards. Promotion evidence is a packet of linked artifacts the engineer owned.
It is calibration week. One engineer merged 212 pull requests this half and tops the vendor’s usage leaderboard. Another merged 19, but wrote the acceptance suite and holdout scenarios that let the payments team stop reading every diff. Your ladder still says “delivers a high volume of quality code”, and a manager has pasted the leaderboard export into a promotion packet.
This page is for the CTO or VP Engineering who owns the ladder and the tech lead or manager who writes reviews. It gives you a competency matrix, the metrics that stay out of reviews, an evidence packet, a telemetry charter with the settings that enforce it, and a one-cycle rollout. Hiring-stage role profiles are in hiring and interviewing for agentic engineering.
Why do career ladders break when agents write the code?
Section titled “Why do career ladders break when agents write the code?”Most ladders were written when output was expensive, so “delivers a large volume of high-quality code” was a fair proxy for impact. Agents removed the cost of the first half of that sentence and made the second half harder to see. The evidence points the same way:
- Output rose faster than delivery. Faros AI’s AI Engineering Report 2026 (April 2026, telemetry from 22,000 developers) reports merge rate per developer up 16.2%, with median time in review up 441.5%, incidents per pull request up 242.7% and code churn up 861%. Rewarding the first number rewards the cause of the other three.
- The work moved to review. In Anthropic’s study of its own engineers (2025-12-02, a vendor’s self-report from 132 survey respondents and 53 interviews), some describe their work shifting “70%+ to being a code reviewer/reviser rather than a net-new code writer.” A commit count does not see that work.
- AI share describes the industry, not the person. DX reports an average of 51.9% AI-authored code across more than 400 companies in Q2 2026 (Justin Reock, 2026-06-17, self-reported). An individual’s share shows tool habits, not judgment.
- Self-assessment of AI speed is unreliable. In METR’s 2025 randomized trial (2025-07-10, early-2025 tools), experienced open-source developers took 19% longer with AI and still believed it had sped them up by 20%. “I shipped much more with agents” is a hypothesis, not evidence.
The fix is not a new metric but a ladder that names the work that now separates engineers, and a review that asks for evidence of it. Some old signals need a replacement, not only a removal: “writes clean code” becomes designing checks that keep code clean without reading it (fitness functions, lint and type gates), and depth shown in hard code becomes depth shown in specs, oracle design and incident diagnosis. Volume, speed and review counts are covered by the Goodhart table below.
The competency matrix for agentic engineering
Section titled “The competency matrix for agentic engineering”Use the matrix to replace your ladder’s technical rows. Map “Engineer” to your mid-level band, “Senior” to senior, and “Staff and above” to staff, principal and equivalent. For entry-level expectations, see growing junior developers when agents write the code.
| Dimension | Engineer | Senior engineer | Staff engineer and above |
|---|---|---|---|
| Outcomes | Ships scoped changes that meet their acceptance criteria and names the outcome each served | Owns an outcome metric for a feature or service (lead time, error rate, adoption) and moves it | Decides which outcomes a group pursues, and stops agent work that does not move them |
| Specification | Writes acceptance criteria before the agent runs; agents rarely misread them | Writes specs and short ADRs that other engineers’ agents implement; spec-caused rework is rare and traced | Sets the spec and ADR standard for a group and resolves cross-team ambiguity before it reaches an agent |
| Verification | Proves each change with tests and an evidence bundle; can explain the change without the diff open | Designs oracles (acceptance tests, property tests, holdout scenarios) and raises oracle strength for a module; catches defects in review before merge | Decides which loops may move to evidence-only review, owns the verification strategy, and rolls trust back when a gate misses |
| Harness and leverage | Reports and fixes harness friction: a flaky check, a misleading rule, a missing command | Ships skills, hooks, checks or context files that other engineers adopt and that measurably cut rework | Builds harness products used across teams (evals, policy, environments) and retires what does not help |
| Operation and ownership | Debugs their own changes in production without an agent when needed | Leads incidents involving agent-written code and writes reviews that change a gate, not only a line | Is accountable for a system’s reliability and for the autonomy level its agents run at |
| People and craft | Keeps unassisted skills current (debugging, reading unfamiliar code) and asks for review-to-learn sessions | Runs review-to-learn and spec reviews for juniors; interviews with the agent-allowed loop | Grows seniors, shapes the hiring loop and the ladder itself |
The matrix shares its vocabulary with the interview rubric in hiring for agentic engineering, so people are hired and promoted on the same evidence.
Two rules keep the matrix honest:
- Accountability does not transfer to the agent. The engineer owns what their agents shipped, including the parts nobody read. An engineer who cannot explain or roll back their agent’s change has not met the Verification row at any level.
- Leverage counts only when others use it. A skill, hook or check earns credit when another team adopts it and the rework it targets goes down.
What evidence counts for each dimension?
Section titled “What evidence counts for each dimension?”Every claim in a review needs an artifact a calibration committee can open; the counts in the next section never qualify.
| Dimension | Evidence that counts |
|---|---|
| Outcomes | The team’s outcome metric before and after, with the changes that moved it |
| Specification | Specs, acceptance criteria and ADRs, plus the clarifications they needed after hand-off |
| Verification | Oracles added with mutation score or holdout pass rate before and after; defects caught before merge; escaped defects traced to the engineer’s checks |
| Harness and leverage | The artifact, who adopted it, and the rework or review time it removed |
| Operation and ownership | Incident reviews led, rollbacks run, gates added after an incident |
| People and craft | Juniors’ progress on their own plan, review-to-learn notes, interview feedback quality |
Which metrics become Goodhart targets?
Section titled “Which metrics become Goodhart targets?”Goodhart’s law, in Marilyn Strathern’s phrasing: “When a measure becomes a target, it ceases to be a good measure” (Marilyn Strathern, “Improving ratings: audit in the British University system”, European Review 5(3), 1997). Agents shorten the time from target to gamed metric, because inflating a count now costs a prompt.
| Metric | How agents make it easy to inflate | What it is still good for | In an individual review? |
|---|---|---|---|
| Merged pull requests, commits | Split work into many small agent pull requests | Team throughput, read beside change fail rate | Never |
| Lines added or changed, accepted lines | More lines is often worse | Nothing on its own | Never |
| Tokens, spend, cost per person | A target rewards loops left running; a cap punishes expensive verification | Team cost per accepted change; budget management | Never |
| ”% AI-written”, sessions, active time | Route every task through an agent | Team-level utilization and friction | Never |
| Vendor leaderboard rank | Combines the counts above | Finding people who can share workflows | Never |
| Test count, line coverage | Tests that assert little or mirror the implementation | A CI floor | Only as a linked oracle with its mutation score |
| Reviews performed, approvals | Approve faster, read less | Balancing reviewer load | Never; defects caught before merge count instead |
The vendor leaderboard is not hypothetical. Claude Code’s Team and Enterprise analytics dashboard ranks the top 10 users by pull requests or lines of code and offers an Export all users CSV (Claude Code analytics docs, checked 2026-09-26). Anthropic presents it as a way to find power users who can help others; use it for exactly that, never in a promotion packet.
The team-level definitions (accepted change rate, cost per accepted change, review load) are in DORA, SPACE, DX Core 4 and AI measurement, which shares this page’s rule: no metric is reported per engineer.
What goes into a review instead: the evidence packet
Section titled “What goes into a review instead: the evidence packet”Ask every engineer for one evidence packet per review cycle (not to be confused with a PR’s evidence bundle, which proves a single change). It replaces the self-review’s list of things built. The engineer writes it; the manager checks every link.
# Evidence packet: <name>, <cycle, e.g. 2026-H2>
## Outcomes (1–3)- Outcome: <metric the team tracks, e.g. checkout p95 latency> Before → after: <values, dates, source dashboard link> My part: <spec / oracle / change links>
## Specification- <link to spec, acceptance criteria or ADR> — implemented by: <who or which agent loop> Clarifications needed after hand-off: <count, links>
## Verification- Oracle added or strengthened: <link> — mutation score or holdout pass rate before → after- Defects caught before merge: <PR comment links>- Escaped defects in code I verified: <incident or bug links, with what the gate missed>
## Harness and leverage- <skill / hook / check / context file link> — adopted by: <teams or people> Effect: <rework, review time or failure rate before → after, with source>
## Operation and ownership- Incidents led or rollbacks run: <links>; gate changed afterwards: <link>
## People and craft- Review-to-learn or spec reviews run: <links or notes>; who grew and how- Unassisted practice this cycle: <what, and what I learned>
## What I would do differently- <one or two sentences, with a link if one exists>Keep it to two pages; a staff packet is mostly outcomes, verification strategy and leverage.
How do you keep telemetry from becoming surveillance?
Section titled “How do you keep telemetry from becoming surveillance?”You need agent telemetry for cost, harness friction and team metrics. Cut per person, it becomes a monitoring system, and engineers who suspect it feeds their review stop reporting failed runs.
There is also a legal side. A model that scores, ranks or allocates work to engineers from agent telemetry is an Annex III point 4(b) use under the EU AI Act (monitoring and evaluating the performance and behaviour of workers), high-risk from 2 December 2027 under the 2026 omnibus deferral (secondary source; confirm on EUR-Lex); see the EU AI Act for companies building with agents. Employee monitoring also raises GDPR questions today, and German works councils co-determine technical systems suited to monitoring employees’ performance (BetrVG §87(1) no. 6). Take these to your counsel; this is not legal advice. Adopt and publish the charter below before you switch telemetry on.
# Agent telemetry charter — <company>, effective <date>
1. Purpose. We collect coding-agent telemetry to manage cost, find harness friction, and measure delivery at team level. We do not collect it to evaluate individuals.2. Aggregation. Every dashboard and report shows teams of five or more people. A group smaller than five is merged into its parent group.3. Identity. Personal identifiers (email, account IDs, user and session IDs) are removed in the collector before storage, except in the licence-administration store (rule 5).4. Content. Prompt text, assistant responses, tool arguments and tool output are not collected.5. Per-person access. An engineer can see their own usage. Licence administrators can see per-seat usage to manage seats and spend limits, and nothing else.6. Reviews. No telemetry, vendor leaderboard or analytics export enters a performance review, promotion packet, calibration or performance plan.7. No automated judgement. No model or rule scores, ranks or assigns work to individual engineers from this data.8. Retention. Raw events are kept 90 days; team aggregates are kept 24 months.9. Changes. Any change to this charter is announced 30 days ahead and agreed with <employee representatives / works council, where one exists>.10. Audit. <Role> checks rules 2–7 each quarter and publishes the result internally.The charter’s numbers are defaults, not legal thresholds; set them with your counsel.
Enforce the charter in each tool
Section titled “Enforce the charter in each tool”The charter only holds if the configuration matches it.
What it sends. For a signed-in user, Claude Code attaches user.email, user.account_uuid, user.account_id and organization.id to every metric and event it exports to your OpenTelemetry endpoint. Every event also carries user.id and session.id; on a Claude apps gateway session user.id is the IdP subject and user.groups lists the person’s IdP groups. Prompt text, assistant responses, tool arguments, tool output and full API bodies are off by default and gated by OTEL_LOG_USER_PROMPTS, OTEL_LOG_ASSISTANT_RESPONSES, OTEL_LOG_TOOL_DETAILS, OTEL_LOG_TOOL_CONTENT and OTEL_LOG_RAW_API_BODIES (Claude Code monitoring docs, checked 2026-09-26 against v2.1.283). Charter rule 4 needs all five unset in managed settings. OTEL_LOG_ASSISTANT_RESPONSES inherits OTEL_LOG_USER_PROMPTS when unset, so leave both unset (or set both to 0 explicitly).
Configure it centrally. Claude Code ignores the exporter variables in a repository’s .claude/settings.json, so set them in managed settings, tagged by team, with account IDs dropped from metrics:
{ "env": { "CLAUDE_CODE_ENABLE_TELEMETRY": "1", "OTEL_METRICS_EXPORTER": "otlp", "OTEL_LOGS_EXPORTER": "otlp", "OTEL_EXPORTER_OTLP_PROTOCOL": "grpc", "OTEL_EXPORTER_OTLP_ENDPOINT": "http://collector.example.com:4317", "OTEL_RESOURCE_ATTRIBUTES": "team.id=payments,cost_center=eng-123", "OTEL_METRICS_INCLUDE_ACCOUNT_UUID": "false", "OTEL_METRICS_INCLUDE_SESSION_ID": "false" }}Distribute a different team.id per team. user.email and user.id are still attached, and events still carry session.id, so remove them in the collector (below).
The analytics dashboard. Every Admin and Owner can view the leaderboard and the Export all users CSV. Keep those role lists short, and keep the export and the per-user spend report out of review tooling.
What it sends. In the openai/codex source at rust-v0.157.1 (checked 2026-10-02 against CLI 0.157.1), the log and trace exporters are off by default and log_user_prompt is false, but metrics_exporter defaults to statsig, OpenAI’s own metrics pipeline (codex-rs/config/src/types.rs). Every metric carries the session tags auth_mode, session_source, originator, service_name, model and app.version (codex-rs/otel/src/metrics/tags.rs), and individual metrics add their own: codex.turn.cost_microusd carries conversation.id, a per-session key, and turn.id, and codex.tool.call and codex.tool.call.duration_ms carry the tool name (codex-rs/otel/src/events/session_telemetry.rs). codex-rs/otel/src/metrics/config.rs keeps those metrics out of Statsig so that custom OTLP exporters still receive them, which means a metrics-only exporter still sends a per-session and a per-turn key to your collector; the collector below deletes both. Log events carry more. Every one carries user.email, user.account_id and conversation.id (codex-rs/otel/src/events/shared.rs), and the codex.tool_result event records the tool’s full arguments and up to [otel.tool_result] max_bytes of its output, 2048 bytes by default (codex-rs/otel/src/tool_result.rs). No flag switches that event off, and max_bytes does not shorten the arguments, so a Codex log exporter breaks charter rule 4. Send metrics only:
[otel]environment = "prod"log_user_prompt = false# No `exporter` (logs) and no `trace_exporter`: Codex log events carry tool# arguments, tool output and user.email. Metrics still carry conversation.id# and turn.id; the collector deletes both.metrics_exporter = { otlp-http = { endpoint = "https://otel.example.com/v1/metrics", protocol = "binary" } }Tag it by team. Codex’s metrics resource sets only service.name, service.version, env, os and os_version (codex-rs/otel/src/metrics/client.rs), and config.toml has no key that adds a resource attribute to metrics. Codex builds that resource with Resource::builder() from opentelemetry_sdk 0.31, which also reads OTEL_RESOURCE_ATTRIBUTES: in a test on 2026-10-02, CLI 0.157.1 started with OTEL_RESOURCE_ATTRIBUTES=team.id=payments exported team.id=payments on the resource of every metric batch. Set that variable per team in the shell or launch environment your device management controls, since config.toml cannot carry it. A session started without it arrives with no team.id; attach a team to it in the pseudonymisation load or drop it.
The main branch adds log_agent_responses and log_guardian_assessments (both off by default); CLI 0.157.1 ignores them. Run the collector check below on the first Codex session anyway: a later version can add tags. The admin constraints file requirements.toml has no key for the [otel] table (config_requirements.rs, checked 2026-09-26), so distribute config.toml through device management and rely on that check; see enforcing one policy across every coding agent.
Cursor keeps per-member usage in its team admin analytics; this page could not check the current fields on 2026-09-26. Restrict the admin role to licence administration (charter rule 5). Treat any Cursor export as raw telemetry: load it through the same pseudonymisation job as your collector data, which drops every column that names or keys a person and keeps the team, and run the same identity check below on the first load before anyone builds a report on it.
Whichever tools you run, strip identity in the OpenTelemetry Collector before the data reaches storage. The attributes processor removes the keys from metrics, logs and traces. It ships in the OpenTelemetry Collector contrib distribution, not the core otelcol build, so run this with otelcol-contrib or a distribution that includes the processor. It also leaves resource attributes alone: if a tool puts identity in the resource block, add a resource processor with the same delete actions. This is a complete collector configuration; replace the otlphttp endpoint with your storage backend:
receivers: otlp: protocols: grpc: {} http: {}
processors: attributes/drop-identity: actions: - key: user.email action: delete - key: user.account_uuid action: delete - key: user.account_id action: delete - key: user.id action: delete - key: user.groups action: delete - key: session.id action: delete - key: conversation.id action: delete - key: turn.id action: delete batch: {}
exporters: otlphttp: endpoint: https://storage.example.com debug: {}
service: pipelines: metrics: receivers: [otlp] processors: [attributes/drop-identity, batch] exporters: [otlphttp] logs: receivers: [otlp] processors: [attributes/drop-identity, batch] exporters: [otlphttp]Delete user.id even though it holds no email: it is stable for an installation, or the IdP subject on a gateway session, so it is a per-person pseudonym that lets anyone rebuild a leaderboard. Codex’s conversation.id is the same kind of per-session key, and its turn.id is a finer one (one value per turn of one person’s session), so both go: charter rule 3 removes session IDs, and a turn ID is a session ID with a counter. If you need session counts per team, use action: hash for session.id instead of delete.
To check it, add debug to the exporters list of both pipelines, send one session through, and confirm that no user.*, session.id, conversation.id, turn.id or account attribute survives, and that team.id does. Repeat the check after every agent upgrade: new versions add attributes.
Roll out the new ladder in one review cycle
Section titled “Roll out the new ladder in one review cycle”Remove the bad incentive first, then give people the new evidence format, then calibrate against it.
-
Announce what stops now. Before the next cycle opens, state in writing that pull request counts, lines, tokens, AI share and leaderboard positions will not appear in any review, packet or calibration.
-
Publish and enforce the telemetry charter. Agree it with employee representatives where you have them, and ship the collector and managed-settings changes.
-
Rewrite the technical rows of the ladder. Replace them with the matrix, adjusted to your levels; the ladder-audit prompt below flags every volume-based phrase.
-
Pilot the evidence packet with two teams, one that runs agents heavily and one that does not. What engineers found hard to evidence shows where the process hides work.
-
Calibrate with the checklist below. The calibration chair reads packets before the meeting and stops any discussion that falls back on counts.
-
Audit the cycle. Run the checks in the next section, publish the aggregate results, and fix any dimension that produced no evidence.
Calibration checklist
Section titled “Calibration checklist”The chair reads this aloud at the start of every calibration meeting.
- Every rating cites at least one linked artifact per dimension discussed.
- No one quotes a count of pull requests, commits, lines, tokens, sessions or a leaderboard position.
- For each “exceeds” rating, the committee names the outcome and who else benefited.
- Review and verification work is discussed for every senior and staff engineer.
- Escaped defects are discussed as gate failures first: “which check should have caught this?”
- Harness contributions cite adoption by someone else.
- Light and heavy agent users are rated on the same dimensions and evidence.
How do you know the new ladder rewards the right work?
Section titled “How do you know the new ladder rewards the right work?”A ladder’s output is who gets promoted. Check that output every cycle rather than trusting the wording.
| Check | How | Healthy result | Owner |
|---|---|---|---|
| Evidence coverage | Sample 10 packets and count claims without a working link | Fewer than one unlinked claim per packet | Calibration chair |
| Volume leakage | People analytics correlates ratings with merged pull request counts on pseudonymised data and reports only the coefficient | Weak or none; a strong positive one means volume is still rewarded | Head of people analytics |
| Dimension balance | Tally which dimensions each promotion case cited | Verification and harness work appear in senior and staff promotions, not only outcomes | VP Engineering |
| Team outcomes | Change fail rate, time in review and accepted change rate from the metrics panel | Stable or better | Tech leads |
| Trust | Anonymous survey item: “I can report a failed agent run or a reverted change without it hurting my review” (five-point scale) | Most answers 4 or 5, and not falling | Engineering managers |
| Charter compliance | Collector check for identity attributes; dashboard access review | No personal identifiers outside licence administration | Platform or security team |
The VP Engineering signs off the ladder and the cycle audit; the people partner and any employee representative body co-sign the charter.
What goes wrong when you rewrite a career ladder for agents?
Section titled “What goes wrong when you rewrite a career ladder for agents?”The leaderboard reaches calibration anyway. A manager brings the vendor export “for context”. Recovery: the chair stops the discussion, the export leaves the packet, and charter rule 6 goes into the next invitation.
Harness busywork replaces code volume. Skills nobody installs appear; credit harness work only with adoption and a measured effect.
Test counts become the new line counts. Recovery: accept only oracle strength (mutation score, holdout pass rate) or defects caught, and protect the oracle from agent edits.
Senior reviewers are rated down for “low output”. Recovery: put defects caught before merge in the packet and show reviewer load from running the review queue.
Engineers switch telemetry off or route around it. Recovery: publish the charter, the collector configuration and the audit result. One breach of the charter undoes it.
Specs get longer, not better. Judge specs by what happened after hand-off, never by length.
Juniors cannot show evidence at their level. Recovery: give entry levels a smaller packet built around their development plan from growing junior developers.
A model starts writing the ratings. A manager asks an assistant to rank the team from telemetry. Recovery: prohibit it in the AI usage policy; it is the automated worker evaluation the EU AI Act classifies as high-risk. The review-check prompt above checks a review; it never scores one.
Where to go next with career ladders
Section titled “Where to go next with career ladders”Change the hiring loop next, so new engineers are selected on the same dimensions they will be promoted on.
Frequently asked questions
Should pull request counts or token usage appear in engineers' performance reviews?
No. Once agents write the code, both can be raised without delivering anything: more, smaller agent pull requests, or agent loops left running. Keep them at team level, next to a stability metric, and judge individuals on linked evidence of outcomes, specs, verification and harness work.
What should a career ladder reward when agents write most of the code?
Outcomes the engineer owned, specifications other people's agents could implement without misreading, verification that lets a loop run on evidence instead of line-by-line reading, harness contributions other engineers adopt, and ownership of incidents and rollbacks.
How do you use agent telemetry without it becoming surveillance?
Write a telemetry charter before you collect anything: team-level aggregates only, a minimum group size, prompt content off, per-person data visible only to the engineer and to licence administration, and a written rule that no telemetry enters a review, promotion or performance plan.