Skip to content

Hiring and interviewing for agentic engineering

Hiring for agentic engineering means running an interview loop in which candidates use a coding agent, and scoring what the agent cannot supply: a spec another engineer could build from, a verification plan that would catch a wrong result, and review judgment on an agent-written pull request. Anchored rubrics, scored independently, keep the loop comparable across candidates.

Your interview loop still bans AI tools. Last quarter it hired the candidate who solved the whiteboard problem cleanly. Three months later that hire merges agent pull requests faster than anyone, and approved one whose only test asserted the value its own mock returned. This page is for the CTO who sets the hiring policy and the tech lead who runs the loop.

  • A role profile: what changes, and what stays, when agents write most of the code
  • A four-stage loop with agents allowed: four hours of stages (4.5 hours with breaks), with one short no-agent segment
  • Three work samples you can build from your own codebase: a spec exercise, an agent work sample and a seeded-defect review
  • A five-dimension rubric with behavioural anchors, and hire bars per level
  • Environment setup for Claude Code, Codex and Cursor, and three prompts to build the exercises
  • Metrics that tell you whether the loop predicts on-the-job performance

The loop assumes your team already works this way: acceptance criteria before the agent runs and evidence instead of line-by-line review. If it does not, fix that first.

Why the old interview loop measures the wrong thing

Section titled “Why the old interview loop measures the wrong thing”

A classic loop scores how well a candidate writes correct code from memory, under time pressure, alone. The evidence on how agentic work actually runs points somewhere else:

SourceWhat it found
Anthropic workplace study, 2025-12-02 (132 engineers and researchers surveyed, 53 interviewed)Some engineers said their work had shifted “70%+ to being a code reviewer/reviser rather than a net-new code writer”. The report names a “paradox of supervision”: overseeing agents needs the skills that over-delegation erodes.
Anthropic skill-formation trial, 2026-01-29 (randomized, 52 mostly junior engineers)The AI group averaged 50% on a follow-up quiz against 67% for the hand-coding group, with the largest gap on debugging. Participants who asked conceptual questions or asked for explanations of generated code averaged 65% or higher.
Faros AI telemetry report, April 2026 (22,000 developers)Median time in review rose 441.5% (last verified 2026-08-28).
DORA report, 2025-09-23“90% of survey respondents report using AI at work”.

Review and supervision are now the core of the job. A ban tests a setup candidates no longer work in; allowing agents under the old scoring (“did the code work in 60 minutes?”) measures the agent. Allow agents and score the human’s decisions: what to build, how to prove it works, and what to reject.

The job does not shrink to “prompting”: it moves from producing code to specifying it and deciding whether it is correct, as the humans’ job describes. Rewrite the role profile before the interview.

DimensionWhat the new loop scoresWhat stays mandatory
SpecificationTurns a vague request into testable criteria; names missing decisionsAsking the product question first
Verification designChooses evidence that fails on a wrong implementation; spots weak testsUnderstanding what a test asserts
Agent directionGives context, splits work, redirects the agent, verifies before “done”Knowing when to write code by hand
Review judgmentFinds and ranks consequential defects in an agent-written pull requestCareful reading where risk is high
FundamentalsDebugging from a hypothesis, explaining code without the agentBoth, at every level

Put this in your job descriptions so candidates know what the loop tests:

## How we work and how we interview
Our engineers use coding agents (Claude Code, Codex or Cursor) for most
implementation work. The job is deciding what to build, proving it works and
reviewing what agents produce. Our interviews allow the agent of your choice
in every stage except one 15-minute debugging segment. We score your
acceptance criteria, your verification plan and your review findings, not
how fast the code appears.

The four-stage interview loop with agents allowed

Section titled “The four-stage interview loop with agents allowed”

Every candidate at a level gets the same exercises, tool options and time.

  1. Spec exercise, 45 minutes, agent allowed. Turn a product request into acceptance criteria, out-of-scope items and a verification plan; can run as a time-boxed take-home. Tests: specification, verification design.

  2. Agent work sample, 90 minutes, live, the candidate’s own tool. Implement a small change in a prepared repository while an interviewer observes. Tests: agent direction, verification design.

  3. Review exercise, 45 minutes, agent allowed. Review an agent-written pull request with seeded defects and write a verdict. Tests: review judgment.

  4. Verification design and fundamentals, 60 minutes. 45 minutes on making a risky change safe to ship without reading every line, then 15 minutes of debugging with no agent. Tests: verification design, fundamentals.

Send the candidate the format a week ahead: which stages allow an agent, which tools the machine offers, and that the transcript is kept for the interview record. For juniors, give 30 minutes of stage 4 to debugging; for staff, replace half of stage 2 with a review of the harness itself.

Use a realistic request with two decisions deliberately missing. For a subscription product:

“Customers on annual plans should be able to pause their subscription for up to three months. They shouldn’t be charged while paused. Support keeps getting this request, so we’d like it this quarter.”

The hidden decisions: how often a customer may pause, and what happens to the renewal date. A strong candidate asks or states each as an assumption.

What the candidate hands in:

  • Given/When/Then acceptance criteria with concrete values (“Given an annual plan renewing on 2027-03-01, when the customer pauses for 90 days on 2026-11-01, then the renewal moves to 2027-05-30”)
  • What is out of scope
  • The test that proves each criterion, and which tests to protect from agent edits

Levels differ on the edges, not the count: the boundary (a pause of exactly 90 days, and 91), concurrency (a pause during a renewal charge), failure (the payment provider times out mid-pause) and time zones. An agent drafts happy-path criteria in seconds, so reward the edges and the verification plan.

Build the exercise repository once and reuse it for a year. It needs:

  • A small, real-looking service. A few thousand lines, a test suite and one command that runs the checks, such as npm test.
  • A task an agent can get plausibly wrong. Put the trap in the domain: a date boundary, idempotency on retry or an authorisation rule. The agent’s first attempt should pass the existing tests and still be wrong.
  • One weak existing test that passes whatever the code does.
  • No instruction file. Leave out CLAUDE.md and AGENTS.md: how the candidate gives the agent context is part of the signal.

The interviewer answers only product questions and takes notes with this template:

# Stage 2 notes: <candidate id>, <date>, tool: <Claude Code | Codex | Cursor>
## Before the first agent run (timestamp)
Context given to the agent: files, constraints, acceptance criteria?
Plan requested or written before code?
## During
Times the candidate stopped or redirected the agent, and why:
Agent claims accepted without checking (quote them):
Agent claims checked, and how:
## Before "done"
Evidence the candidate produced: tests run, new tests, manual check
Did they find the weak existing test? Did they find the trap in the task?
## Rubric evidence (one line per dimension, no score yet)
Agent direction:
Verification design:

Two candidates both finish in 70 minutes. One accepts the agent’s “all tests pass”. The other asks why a retry cannot double-charge, sees that nothing tests it, writes that test and watches it fail. Only the second did the job.

Set up the interview environment in Claude Code, Codex and Cursor

Section titled “Set up the interview environment in Claude Code, Codex and Cursor”

Use a disposable machine or container per candidate: no company credentials, no internal repositories, no personal configuration, and outbound network limited to the model provider and your package registry. Pin the model so candidates meet the same agent; the models hub lists current defaults. The disposable machine is what makes the candidate’s usual permission mode safe.

Create a fresh OS user with an empty ~/.claude directory, clone the exercise repository, and start the session with a known ID so you can find its transcript afterwards (checked against Claude Code 2.1.283):

Terminal window
# Terminal on the disposable interview machine, inside the exercise repository
SESSION_ID=$(uuidgen)
claude --model claude-opus-5-5 --session-id "$SESSION_ID" -n "interview-cand-17-stage2"

Use the full model ID, not the opus alias, which resolves differently per provider and moves when Anthropic changes it. claude-opus-5-5 needs v2.1.280 or later, which on 2026-09-26 is the latest release channel only; stable (2.1.274) does not offer it yet. The transcript is a JSONL file under ~/.claude/projects/: copy it into the interview record, or reopen it with claude --resume "$SESSION_ID". For the review exercise, calibrate against /code-review inside a session.

Build one agent-written pull request with six defects of different kinds and a confident description that claims all tests pass. A candidate who finds only style issues and misses the authorisation hole fails the stage.

Seeded defectClassSeverity
Any authenticated user can pause any subscription by IDSecurityCritical
A webhook retry pauses the subscription twiceLogicCritical
Pause end computed in server local time, one day off for customers in UTC+10LogicHigh
A unit test asserts the value its own mock returnsTest gapHigh
A caught error returns HTTP 200 with an empty bodyError handlingMedium
An unrelated rename across 12 logging filesScopeMedium

Score the verdict, not only the list. Strong candidates return the pull request, rank the two critical defects first, and propose a check that would catch each class automatically, such as an authorisation test per endpoint or an idempotency test on the webhook handler.

Calibrate first with the review prompt further down: if your review agent finds all six defects, make them subtler. The human bar is what the machine reviewer missed, in order of consequence.

Stage 4: questions for the verification design conversation

Section titled “Stage 4: questions for the verification design conversation”

Start from a change in the candidate’s own work, then move to one of yours:

  • “Your agent opened a 900-line pull request that touches billing. What evidence do you need before you approve it without reading every line?”
  • “Which tests in that change must the agent never edit, and how would you enforce that?”
  • “The agent says the migration is backwards-compatible. How do you check the claim?”
  • “When would you stop an agent and write the code yourself?”

Strong answers name concrete oracles (property tests, contract tests, a canary with a rollback trigger, protected acceptance tests) and who signs off. Weak answers: “I’d review it carefully” or “the agent writes tests”.

Then run the 15-minute debugging segment without an agent: a failing test, a stack trace and a log excerpt. Score the method, not the fix: does the candidate form a hypothesis, test it with the smallest experiment and revise it?

The scoring rubric for agentic engineering interviews

Section titled “The scoring rubric for agentic engineering interviews”

Each interviewer scores only their stage’s dimensions, with one evidence line per score, before the debrief. Scores without evidence do not count. For inter-rater agreement, have a second interviewer shadow and score one stage per candidate, independently.

Dimension1: absent2: partial3: solid4: exemplary
SpecificationRestates the request; no testable criteriaHappy-path criteria with concrete values; misses the hidden decisionsNames both hidden decisions; covers boundary and failure casesAlso covers concurrency and time; states what is out of scope and why
Verification design“Run the tests”Adds tests, but they would pass on a wrong implementationEvidence that fails on the likely wrong implementation; finds the weak existing testProtects the key tests from agent edits; names the check that replaces a human read, per risk
Agent directionPastes the task and accepts the resultGives some context; accepts unchecked claimsPlans first, gives constraints, stops and redirects on a wrong turn, verifies before “done”Splits the work so each agent step is checkable; knows when to write it by hand
Review judgmentStyle comments onlyFinds some defects, misses a critical oneFinds both critical defects and ranks them first; returns the pull requestAlso proposes a check per defect class and asks for the scope split
FundamentalsCannot explain the code or form a hypothesisFixes by trial and errorDebugs from a hypothesis and explains the code without the agentExplains the root cause and the test that should have caught it

Set the hire bars per level in writing before the first candidate, and never move them for one candidate:

LevelMinimum on every dimensionAlso required
Junior2Fundamentals 3, and stage 2 notes show the candidate asking conceptual questions of the agent
Mid23 on specification, verification design and review judgment
Senior34 on verification design or review judgment
Staff34 on verification design and review judgment

The junior bar follows the skill-formation trial: conceptual questions kept understanding high.

Copy-paste prompts for building the exercises

Section titled “Copy-paste prompts for building the exercises”

These prompts are for interviewers. Run them in the exercise repository, never on candidate material.

The tech lead who runs the loop collects these each quarter; the CTO owns the rubric and signs off on any change to the hire bars.

MetricDefinitionTarget to start fromAction when it misses
Calibration on current staffBefore the first candidate, run the loop on three engineers you rate as strong at the target levelAll three clear the barFix the exercise or the bar
Inter-rater agreementShare of scores where two interviewers of the same stage are within one pointAt least 80%Retrain on anchors with recorded examples
Seeded-defect catch distributionPer candidate: critical defects found out of two, and total found out of sixSpread across the range, not all 6/6Defects too easy, or the exercise leaked; rotate it
Stage pass ratesShare of candidates passing each stageNo stage passes nearly everyoneThat stage is not discriminating; cut or redesign it
Six-month outcomeAt the end of probation, the manager’s rating of the hire’s review and spec work against the interview scoreMost hires match or exceedRevise the anchor that misled
Candidate timeTotal hours asked of the candidate, including take-homeAt most 4.5 hours, breaks includedCut, do not extend

Rotate the exercise repository and the seeded pull request at least once a year, and immediately if a candidate has seen them.

What goes wrong with agent-allowed interviews, and how to recover

Section titled “What goes wrong with agent-allowed interviews, and how to recover”

The loop scores the agent, and speed wins. Everyone passes stage 2, the notes say “working solution”, and interviewers reward whoever finished first. Recovery: remove time-to-finish from the debrief and rescore only the rubric dimensions. Retrain observers whose notes hold no verification evidence.

The review exercise is too easy. The candidate pastes the diff into an agent and scores 6/6. Recovery: calibrate, and score the verdict, ranking and proposed checks.

Fundamentals disappear from the loop. Dropping the debugging segment hires people who cannot supervise what they cannot explain. Recovery: keep the no-agent segment at every level.

The company stops hiring juniors because agents are cheaper. A year later no one is ready to become a reviewer. Recovery: keep a junior track with its own bar, paired with growing junior developers.

An AI tool starts screening résumés or scoring transcripts, usually through a vendor feature. Recovery: switch it off for decisions, record it in your AI register, and bring counsel in before 2 December 2027.

The exercise leaks. Pass rates jump and answers look identical. Recovery: switch to a second version you keep ready.

Related practice: the oracle strength of your tests and how seniors take ownership of verification.

Frequently asked questions

Should candidates be allowed to use AI coding agents in interviews?

Yes, in most stages. The job now includes directing agents, so a loop that bans them tests work the hire will not do. Change what you score instead: the spec, the verification plan and the review judgment, which the agent cannot supply. Keep one short segment without an agent for debugging and explaining code.

What should an agent-allowed coding interview measure?

Five dimensions: turning intent into testable acceptance criteria, designing verification that would catch a wrong result, directing the agent, finding consequential defects in an agent-written pull request, and fundamentals such as debugging by hypothesis without an agent.

Can we use AI to score candidates' interview transcripts?

Not for the decision. Evaluating candidates is an Annex III area under the EU AI Act, with high-risk obligations from 2 December 2027 according to Gibson Dunn's summary (a secondary source) of the 2026 amendment, and it raises GDPR questions today. An agent may prepare exercises and extract timestamps; trained humans score against the rubric.