Hiring and interviewing for agentic engineering
Hiring for agentic engineering means running an interview loop in which candidates use a coding agent, and scoring what the agent cannot supply: a spec another engineer could build from, a verification plan that would catch a wrong result, and review judgment on an agent-written pull request. Anchored rubrics, scored independently, keep the loop comparable across candidates.
Your interview loop still bans AI tools. Last quarter it hired the candidate who solved the whiteboard problem cleanly. Three months later that hire merges agent pull requests faster than anyone, and approved one whose only test asserted the value its own mock returned. This page is for the CTO who sets the hiring policy and the tech lead who runs the loop.
What this hiring loop gives you
Section titled “What this hiring loop gives you”- A role profile: what changes, and what stays, when agents write most of the code
- A four-stage loop with agents allowed: four hours of stages (4.5 hours with breaks), with one short no-agent segment
- Three work samples you can build from your own codebase: a spec exercise, an agent work sample and a seeded-defect review
- A five-dimension rubric with behavioural anchors, and hire bars per level
- Environment setup for Claude Code, Codex and Cursor, and three prompts to build the exercises
- Metrics that tell you whether the loop predicts on-the-job performance
The loop assumes your team already works this way: acceptance criteria before the agent runs and evidence instead of line-by-line review. If it does not, fix that first.
Why the old interview loop measures the wrong thing
Section titled “Why the old interview loop measures the wrong thing”A classic loop scores how well a candidate writes correct code from memory, under time pressure, alone. The evidence on how agentic work actually runs points somewhere else:
| Source | What it found |
|---|---|
| Anthropic workplace study, 2025-12-02 (132 engineers and researchers surveyed, 53 interviewed) | Some engineers said their work had shifted “70%+ to being a code reviewer/reviser rather than a net-new code writer”. The report names a “paradox of supervision”: overseeing agents needs the skills that over-delegation erodes. |
| Anthropic skill-formation trial, 2026-01-29 (randomized, 52 mostly junior engineers) | The AI group averaged 50% on a follow-up quiz against 67% for the hand-coding group, with the largest gap on debugging. Participants who asked conceptual questions or asked for explanations of generated code averaged 65% or higher. |
| Faros AI telemetry report, April 2026 (22,000 developers) | Median time in review rose 441.5% (last verified 2026-08-28). |
| DORA report, 2025-09-23 | “90% of survey respondents report using AI at work”. |
Review and supervision are now the core of the job. A ban tests a setup candidates no longer work in; allowing agents under the old scoring (“did the code work in 60 minutes?”) measures the agent. Allow agents and score the human’s decisions: what to build, how to prove it works, and what to reject.
How the engineering role profile changes
Section titled “How the engineering role profile changes”The job does not shrink to “prompting”: it moves from producing code to specifying it and deciding whether it is correct, as the humans’ job describes. Rewrite the role profile before the interview.
| Dimension | What the new loop scores | What stays mandatory |
|---|---|---|
| Specification | Turns a vague request into testable criteria; names missing decisions | Asking the product question first |
| Verification design | Chooses evidence that fails on a wrong implementation; spots weak tests | Understanding what a test asserts |
| Agent direction | Gives context, splits work, redirects the agent, verifies before “done” | Knowing when to write code by hand |
| Review judgment | Finds and ranks consequential defects in an agent-written pull request | Careful reading where risk is high |
| Fundamentals | Debugging from a hypothesis, explaining code without the agent | Both, at every level |
Put this in your job descriptions so candidates know what the loop tests:
## How we work and how we interviewOur engineers use coding agents (Claude Code, Codex or Cursor) for mostimplementation work. The job is deciding what to build, proving it works andreviewing what agents produce. Our interviews allow the agent of your choicein every stage except one 15-minute debugging segment. We score youracceptance criteria, your verification plan and your review findings, nothow fast the code appears.The four-stage interview loop with agents allowed
Section titled “The four-stage interview loop with agents allowed”Every candidate at a level gets the same exercises, tool options and time.
-
Spec exercise, 45 minutes, agent allowed. Turn a product request into acceptance criteria, out-of-scope items and a verification plan; can run as a time-boxed take-home. Tests: specification, verification design.
-
Agent work sample, 90 minutes, live, the candidate’s own tool. Implement a small change in a prepared repository while an interviewer observes. Tests: agent direction, verification design.
-
Review exercise, 45 minutes, agent allowed. Review an agent-written pull request with seeded defects and write a verdict. Tests: review judgment.
-
Verification design and fundamentals, 60 minutes. 45 minutes on making a risky change safe to ship without reading every line, then 15 minutes of debugging with no agent. Tests: verification design, fundamentals.
Send the candidate the format a week ahead: which stages allow an agent, which tools the machine offers, and that the transcript is kept for the interview record. For juniors, give 30 minutes of stage 4 to debugging; for staff, replace half of stage 2 with a review of the harness itself.
Stage 1: design the spec exercise
Section titled “Stage 1: design the spec exercise”Use a realistic request with two decisions deliberately missing. For a subscription product:
“Customers on annual plans should be able to pause their subscription for up to three months. They shouldn’t be charged while paused. Support keeps getting this request, so we’d like it this quarter.”
The hidden decisions: how often a customer may pause, and what happens to the renewal date. A strong candidate asks or states each as an assumption.
What the candidate hands in:
- Given/When/Then acceptance criteria with concrete values (“Given an annual plan renewing on 2027-03-01, when the customer pauses for 90 days on 2026-11-01, then the renewal moves to 2027-05-30”)
- What is out of scope
- The test that proves each criterion, and which tests to protect from agent edits
Levels differ on the edges, not the count: the boundary (a pause of exactly 90 days, and 91), concurrency (a pause during a renewal charge), failure (the payment provider times out mid-pause) and time zones. An agent drafts happy-path criteria in seconds, so reward the edges and the verification plan.
Stage 2: build the agent work sample
Section titled “Stage 2: build the agent work sample”Build the exercise repository once and reuse it for a year. It needs:
- A small, real-looking service. A few thousand lines, a test suite and one command that runs the checks, such as
npm test. - A task an agent can get plausibly wrong. Put the trap in the domain: a date boundary, idempotency on retry or an authorisation rule. The agent’s first attempt should pass the existing tests and still be wrong.
- One weak existing test that passes whatever the code does.
- No instruction file. Leave out
CLAUDE.mdandAGENTS.md: how the candidate gives the agent context is part of the signal.
The interviewer answers only product questions and takes notes with this template:
# Stage 2 notes: <candidate id>, <date>, tool: <Claude Code | Codex | Cursor>
## Before the first agent run (timestamp)Context given to the agent: files, constraints, acceptance criteria?Plan requested or written before code?
## DuringTimes the candidate stopped or redirected the agent, and why:Agent claims accepted without checking (quote them):Agent claims checked, and how:
## Before "done"Evidence the candidate produced: tests run, new tests, manual checkDid they find the weak existing test? Did they find the trap in the task?
## Rubric evidence (one line per dimension, no score yet)Agent direction:Verification design:Two candidates both finish in 70 minutes. One accepts the agent’s “all tests pass”. The other asks why a retry cannot double-charge, sees that nothing tests it, writes that test and watches it fail. Only the second did the job.
Set up the interview environment in Claude Code, Codex and Cursor
Section titled “Set up the interview environment in Claude Code, Codex and Cursor”Use a disposable machine or container per candidate: no company credentials, no internal repositories, no personal configuration, and outbound network limited to the model provider and your package registry. Pin the model so candidates meet the same agent; the models hub lists current defaults. The disposable machine is what makes the candidate’s usual permission mode safe.
Create a fresh OS user with an empty ~/.claude directory, clone the exercise repository, and start the session with a known ID so you can find its transcript afterwards (checked against Claude Code 2.1.283):
# Terminal on the disposable interview machine, inside the exercise repositorySESSION_ID=$(uuidgen)claude --model claude-opus-5-5 --session-id "$SESSION_ID" -n "interview-cand-17-stage2"Use the full model ID, not the opus alias, which resolves differently per provider and moves when Anthropic changes it. claude-opus-5-5 needs v2.1.280 or later, which on 2026-09-26 is the latest release channel only; stable (2.1.274) does not offer it yet. The transcript is a JSONL file under ~/.claude/projects/: copy it into the interview record, or reopen it with claude --resume "$SESSION_ID". For the review exercise, calibrate against /code-review inside a session.
Use a fresh OS user with an empty ~/.codex directory. Pin the model with -m and give the agent a writable workspace through a permission profile, which is beta in Codex 0.157.1. Do not combine the profile with --sandbox or --approve-for-me: OpenAI says the profile and legacy sandbox systems “do not compose”, and --approve-for-me uses the legacy workspace-write sandbox. Start the session like this:
# Terminal on the disposable interview machinecodex -m gpt-6-astra -C ~/exercise -c default_permissions=":workspace"GPT-6 Astra is Codex’s bundled default since CLI 0.153.4; pinning it keeps a default change from splitting your candidate pool. After the interview, codex resume opens a picker of the locally recorded sessions. For the review exercise, calibrate against codex review --base main run on the seeded branch.
Create a dedicated interview account on your Cursor team with no access to company repositories, and open the exercise repository in a clean profile.
We could not verify Cursor’s chat export on 2026-09-26 (cursor.com was unreachable from our environment), so keep the record as a screen recording, with the candidate’s written consent. If your team uses Bugbot, run it on the seeded pull request to calibrate the review exercise.
Stage 3: seed the review exercise
Section titled “Stage 3: seed the review exercise”Build one agent-written pull request with six defects of different kinds and a confident description that claims all tests pass. A candidate who finds only style issues and misses the authorisation hole fails the stage.
| Seeded defect | Class | Severity |
|---|---|---|
| Any authenticated user can pause any subscription by ID | Security | Critical |
| A webhook retry pauses the subscription twice | Logic | Critical |
| Pause end computed in server local time, one day off for customers in UTC+10 | Logic | High |
| A unit test asserts the value its own mock returns | Test gap | High |
| A caught error returns HTTP 200 with an empty body | Error handling | Medium |
| An unrelated rename across 12 logging files | Scope | Medium |
Score the verdict, not only the list. Strong candidates return the pull request, rank the two critical defects first, and propose a check that would catch each class automatically, such as an authorisation test per endpoint or an idempotency test on the webhook handler.
Calibrate first with the review prompt further down: if your review agent finds all six defects, make them subtler. The human bar is what the machine reviewer missed, in order of consequence.
Stage 4: questions for the verification design conversation
Section titled “Stage 4: questions for the verification design conversation”Start from a change in the candidate’s own work, then move to one of yours:
- “Your agent opened a 900-line pull request that touches billing. What evidence do you need before you approve it without reading every line?”
- “Which tests in that change must the agent never edit, and how would you enforce that?”
- “The agent says the migration is backwards-compatible. How do you check the claim?”
- “When would you stop an agent and write the code yourself?”
Strong answers name concrete oracles (property tests, contract tests, a canary with a rollback trigger, protected acceptance tests) and who signs off. Weak answers: “I’d review it carefully” or “the agent writes tests”.
Then run the 15-minute debugging segment without an agent: a failing test, a stack trace and a log excerpt. Score the method, not the fix: does the candidate form a hypothesis, test it with the smallest experiment and revise it?
The scoring rubric for agentic engineering interviews
Section titled “The scoring rubric for agentic engineering interviews”Each interviewer scores only their stage’s dimensions, with one evidence line per score, before the debrief. Scores without evidence do not count. For inter-rater agreement, have a second interviewer shadow and score one stage per candidate, independently.
| Dimension | 1: absent | 2: partial | 3: solid | 4: exemplary |
|---|---|---|---|---|
| Specification | Restates the request; no testable criteria | Happy-path criteria with concrete values; misses the hidden decisions | Names both hidden decisions; covers boundary and failure cases | Also covers concurrency and time; states what is out of scope and why |
| Verification design | “Run the tests” | Adds tests, but they would pass on a wrong implementation | Evidence that fails on the likely wrong implementation; finds the weak existing test | Protects the key tests from agent edits; names the check that replaces a human read, per risk |
| Agent direction | Pastes the task and accepts the result | Gives some context; accepts unchecked claims | Plans first, gives constraints, stops and redirects on a wrong turn, verifies before “done” | Splits the work so each agent step is checkable; knows when to write it by hand |
| Review judgment | Style comments only | Finds some defects, misses a critical one | Finds both critical defects and ranks them first; returns the pull request | Also proposes a check per defect class and asks for the scope split |
| Fundamentals | Cannot explain the code or form a hypothesis | Fixes by trial and error | Debugs from a hypothesis and explains the code without the agent | Explains the root cause and the test that should have caught it |
Set the hire bars per level in writing before the first candidate, and never move them for one candidate:
| Level | Minimum on every dimension | Also required |
|---|---|---|
| Junior | 2 | Fundamentals 3, and stage 2 notes show the candidate asking conceptual questions of the agent |
| Mid | 2 | 3 on specification, verification design and review judgment |
| Senior | 3 | 4 on verification design or review judgment |
| Staff | 3 | 4 on verification design and review judgment |
The junior bar follows the skill-formation trial: conceptual questions kept understanding high.
Copy-paste prompts for building the exercises
Section titled “Copy-paste prompts for building the exercises”These prompts are for interviewers. Run them in the exercise repository, never on candidate material.
How you know the hiring loop works
Section titled “How you know the hiring loop works”The tech lead who runs the loop collects these each quarter; the CTO owns the rubric and signs off on any change to the hire bars.
| Metric | Definition | Target to start from | Action when it misses |
|---|---|---|---|
| Calibration on current staff | Before the first candidate, run the loop on three engineers you rate as strong at the target level | All three clear the bar | Fix the exercise or the bar |
| Inter-rater agreement | Share of scores where two interviewers of the same stage are within one point | At least 80% | Retrain on anchors with recorded examples |
| Seeded-defect catch distribution | Per candidate: critical defects found out of two, and total found out of six | Spread across the range, not all 6/6 | Defects too easy, or the exercise leaked; rotate it |
| Stage pass rates | Share of candidates passing each stage | No stage passes nearly everyone | That stage is not discriminating; cut or redesign it |
| Six-month outcome | At the end of probation, the manager’s rating of the hire’s review and spec work against the interview score | Most hires match or exceed | Revise the anchor that misled |
| Candidate time | Total hours asked of the candidate, including take-home | At most 4.5 hours, breaks included | Cut, do not extend |
Rotate the exercise repository and the seeded pull request at least once a year, and immediately if a candidate has seen them.
What goes wrong with agent-allowed interviews, and how to recover
Section titled “What goes wrong with agent-allowed interviews, and how to recover”The loop scores the agent, and speed wins. Everyone passes stage 2, the notes say “working solution”, and interviewers reward whoever finished first. Recovery: remove time-to-finish from the debrief and rescore only the rubric dimensions. Retrain observers whose notes hold no verification evidence.
The review exercise is too easy. The candidate pastes the diff into an agent and scores 6/6. Recovery: calibrate, and score the verdict, ranking and proposed checks.
Fundamentals disappear from the loop. Dropping the debugging segment hires people who cannot supervise what they cannot explain. Recovery: keep the no-agent segment at every level.
The company stops hiring juniors because agents are cheaper. A year later no one is ready to become a reviewer. Recovery: keep a junior track with its own bar, paired with growing junior developers.
An AI tool starts screening résumés or scoring transcripts, usually through a vendor feature. Recovery: switch it off for decisions, record it in your AI register, and bring counsel in before 2 December 2027.
The exercise leaks. Pass rates jump and answers look identical. Recovery: switch to a second version you keep ready.
Where to go next with hiring
Section titled “Where to go next with hiring”Related practice: the oracle strength of your tests and how seniors take ownership of verification.
Frequently asked questions
Should candidates be allowed to use AI coding agents in interviews?
Yes, in most stages. The job now includes directing agents, so a loop that bans them tests work the hire will not do. Change what you score instead: the spec, the verification plan and the review judgment, which the agent cannot supply. Keep one short segment without an agent for debugging and explaining code.
What should an agent-allowed coding interview measure?
Five dimensions: turning intent into testable acceptance criteria, designing verification that would catch a wrong result, directing the agent, finding consequential defects in an agent-written pull request, and fundamentals such as debugging by hypothesis without an agent.
Can we use AI to score candidates' interview transcripts?
Not for the decision. Evaluating candidates is an Annex III area under the EU AI Act, with high-risk obligations from 2 December 2027 according to Gibson Dunn's summary (a secondary source) of the 2026 amendment, and it raises GDPR questions today. An agent may prepare exercises and extract timestamps; trained humans score against the rubric.