The Claude Agent SDK: building your own agents on Claude Code's loop
The Claude Agent SDK is a Python and TypeScript library that runs the Claude Code binary as a subprocess and exposes its agent loop to your code: built-in tools, permission rules, hooks, subagents, MCP servers and resumable sessions. It fits bots that need custom in-process tools, code-level guardrails or typed results that one claude -p call cannot express.
Your team’s CI goes red twenty times a week, and each time somebody opens a 4,000-line log, scrolls past the cascade of follow-on failures, and decides whether it is a real regression, a flaky test or a broken runner. You tried a claude -p step, and it produced prose that nobody could act on, occasionally cited a file that does not exist, and once asked for a permission nobody was there to grant. This page builds the version that does not do those things.
What you’ll build with the Claude Agent SDK
Section titled “What you’ll build with the Claude Agent SDK”- A CI triage bot in about 130 lines of Python (with a TypeScript equivalent) that classifies a failed GitHub Actions run and cites its evidence
- A permission setup that makes the bot read-only by construction, not by instruction: no
Bash, noEdit, reads confined to the checkout - One custom in-process tool, one
PreToolUsehook, one subagent and a JSON-schema contract, each doing a job you can name - A deterministic check that rejects any verdict whose evidence is not in the log or the repository, and an eval loop that measures the bot against labelled past failures
- A decision rule for when the SDK is worth it and when
claude -por a routine is enough
When should you use the Agent SDK instead of claude -p?
Section titled “When should you use the Agent SDK instead of claude -p?”Start with the smallest thing that works. Most CI bots never need the SDK, and the SDK brings a dependency, a subprocess and an API key into your pipeline.
| You need | Use | Why |
|---|---|---|
| One prompt, one JSON answer, in a shell step | claude -p with --output-format json, --json-schema and --max-budget-usd | No code to maintain. See task automation with headless mode |
A PR review or @claude responder on GitHub | anthropics/claude-code-action@v1 | The action owns checkout, comments and tokens. See CI/CD with Claude Code |
| A scheduled or webhook-triggered run on Anthropic’s cloud | Routines | No runner to operate |
| Custom tools written in your language, code-level guardrails, multi-turn control, typed results, streaming into your own UI | Agent SDK | Everything below this table |
| An agent Anthropic hosts, with no process for you to run | Claude Managed Agents | Anthropic runs the loop and sandbox; you pay model tokens plus $0.08 per session-hour |
| The Claude API with your own tool loop and no Claude Code tools | The Claude API client SDK | The Agent SDK is Claude Code as a library; the client SDK is the raw API |
The SDK is Python and TypeScript only. From any other language, the Agent SDK overview says to run the CLI as a subprocess with -p and --output-format json. For the same job written with the Codex SDK and the Cursor SDK side by side, see driving agents from code with the three SDKs.
Install the Agent SDK and authenticate it
Section titled “Install the Agent SDK and authenticate it”# Python 3.10+. The wheel bundles the Claude Code CLI; nothing else to install.pip install claude-agent-sdk==0.2.160To use a system-wide claude instead of the bundled one, pass cli_path="/path/to/claude" in ClaudeAgentOptions.
# Node 18+. zod is a peer dependency, needed for tool().npm install @anthropic-ai/claude-agent-sdk@0.3.283 zodThe Claude Code binary arrives as a platform-specific optional dependency. If your CI installs with --omit=optional, the binary is missing; install it and pass pathToClaudeCodeExecutable.
Authenticate with an API key from the Claude Console, exported as ANTHROPIC_API_KEY in the process that runs the agent. The SDK does not read .env files for you. For cloud providers, set CLAUDE_CODE_USE_BEDROCK=1, CLAUDE_CODE_USE_VERTEX=1 (Google Cloud’s Agent Platform) or CLAUDE_CODE_USE_FOUNDRY=1 with that provider’s credentials. The Agent SDK overview also states that third-party developers may not offer claude.ai login or its rate limits in products built on the SDK without prior approval, so a product you ship to others uses API-key authentication.
How the SDK drives Claude Code’s loop
Section titled “How the SDK drives Claude Code’s loop”Every SDK call starts the claude binary, sends your prompt over a JSON stream, and yields typed messages back: SystemMessage, AssistantMessage (text and tool-use blocks), UserMessage (tool results), and one ResultMessage at the end of each turn. The ResultMessage is the part your code acts on. It carries subtype (success, error_max_turns, error_max_budget_usd, error_max_structured_output_retries or error_during_execution), is_error, result, structured_output, total_cost_usd, num_turns, session_id and permission_denials.
Two entry points exist in both languages:
query()is one call: prompt in, message stream out. Use it for batch jobs where you know the whole input up front.ClaudeSDKClient(Python) keeps a connection open, so you can send follow-ups,interrupt()a turn and change the permission mode mid-session. The Python README presents it as the entry point for custom tools and hooks, so the example below uses it. In TypeScript,query()returns aQueryobject that does the same when you pass an async iterable as the prompt.
Three defaults surprise people who arrive from the CLI:
- No Claude Code system prompt unless you ask for it. When
system_promptis omitted, both SDKs start the CLI with an empty system prompt (checked in the 0.2.160 and 0.3.283 source). For a coding agent, pass{"type": "preset", "preset": "claude_code", "append": "..."}. - Every settings file loads unless you narrow it.
setting_sourcesleft unset loads user, project and local settings, so a CI runner’s own~/.claude/settings.jsonleaks into the run. Pass["project"]only when the code undercwdis trusted; it must include"project"forCLAUDE.mdto load. Whencwdis a contributor’s commit, pass[]: project settings can define hooks, and hooks run shell commands with the job’s secrets in the environment. - SDK sessions do not start in auto mode. From v2.1.283 (the
latestchannel), auto mode is the starting mode for interactive sessions, butclaude -pand the Agent SDK still start in the default, manual mode. Setpermission_modeexplicitly.
Build the CI triage bot with the Agent SDK
Section titled “Build the CI triage bot with the Agent SDK”The bot runs after a failed CI workflow, reads the failed-step log and the checked-out code, and writes verdict.json. A later workflow step turns that file into a PR comment. The agent itself can neither write files nor run commands.
-
Write the contract before the prompt. The verdict is a JSON schema: a
categoryfrom a fixed list, aculprit_file, a non-emptyevidencearray, aconfidenceand asummary. Passing it asoutput_formatmakes the CLI validate the final answer and retry until it conforms. If it cannot, the turn ends witherror_max_structured_output_retries, not with free text that a later step has to parse. -
Remove the tools the bot must never have.
tools=["Read", "Grep", "Glob", "Agent"]is the complete built-in toolset for this session.Bash,EditandWriteare not restricted; they do not exist for the model.allowed_toolsis a different thing: it only auto-approves, and names nothing about availability. -
Approve by mode, not by list.
permission_mode="dontAsk"denies everything that no rule approves, and never waits for a prompt. Claude Code approves file reads andGrepinside the working directory without any rule, soReadstays out ofallowed_tools. The same holds forAgent, which never asks before running, so the only allow rule is the custom log tool. A bareReadallow rule would approve reads anywhere on the runner, including files outside the checkout. -
Give the agent a narrow tool instead of a shell.
get_failed_logis a Python function exposed as an in-process MCP tool. The agent can fetch the failed-step log for this run and nothing else. It cannot runghwith other arguments, because it has noBash. -
Enforce the rule the model might skip with a hook. Reads inside the checkout are auto-approved, so a
PreToolUsehook denies anyRead,GreporGlobwhose path, glob filter orGlobpattern looks like.env*orsecrets/. Hooks run first in the permission evaluation, and a hook’s deny holds in every mode. The hook is a backstop, not the protection: aGrepwith no path searches the wholecwdand has nothing to match. The real protection is a checkout that holds no secrets. -
Keep the log out of the main context with a subagent. The
log-readersubagent, on the cheaperhaikualias, reads up to 3,000 log lines and returns only the first real error with 20 lines of context. The main agent spends its context on the code, not on the cascade. -
Cap the run and check the answer.
max_turns=30andmax_budget_usd=2.0bound the cost of a confused run. After the result arrives, plain Python checks every evidence item against the log and the checkout, and downgrades the verdict tounknownif any item is not grounded.
# .github/scripts/triage.py: explains why a CI run failed. Read-only by construction.import jsonimport osimport reimport sysfrom pathlib import Path
import anyiofrom claude_agent_sdk import ( AgentDefinition, ClaudeAgentOptions, ClaudeSDKClient, HookMatcher, ResultMessage, create_sdk_mcp_server, tool,)
RUN_ID = os.environ["RUN_ID"]REPO = Path(os.environ.get("TARGET_DIR", ".")).resolve() # the untrusted checkout: data, never code
async def failed_log() -> str: proc = await anyio.run_process(["gh", "run", "view", RUN_ID, "--log-failed"]) return proc.stdout.decode(errors="replace")
# 1. A custom tool: the agent gets the failed-step log, not a shell.@tool("get_failed_log", "Log of the failed steps in the CI run under triage", {"max_lines": int})async def get_failed_log(args): lines = (await failed_log()).splitlines() return {"content": [{"type": "text", "text": "\n".join(lines[-args["max_lines"]:])}]}
# 2. A hook: runs on every matching call, whatever the model decides.async def deny_secret_reads(input_data, tool_use_id, context): tool_input = input_data["tool_input"] keys = ["file_path", "path", "glob"] + (["pattern"] if input_data["tool_name"] == "Glob" else []) # A Grep over the whole cwd has no path to match: the checkout itself must hold no secrets. if any(re.search(r"(^|/)(\.env[^/]*|secrets?)(/|$)", str(tool_input.get(k) or "")) for k in keys): return { "hookSpecificOutput": { "hookEventName": "PreToolUse", "permissionDecision": "deny", "permissionDecisionReason": "Triage never reads secret files.", } } return {}
# 3. The verdict contract.VERDICT = { "type": "object", "properties": { "category": {"enum": ["test-regression", "flaky-test", "build-config", "infrastructure", "unknown"]}, "culprit_file": {"type": ["string", "null"]}, "evidence": {"type": "array", "items": {"type": "string"}, "minItems": 1}, "confidence": {"enum": ["high", "medium", "low"]}, "summary": {"type": "string"}, }, "required": ["category", "culprit_file", "evidence", "confidence", "summary"], "additionalProperties": False,}
options = ClaudeAgentOptions( system_prompt={"type": "preset", "preset": "claude_code", "append": "You are a CI triage bot. You never change files."}, cwd=REPO, tools=["Read", "Grep", "Glob", "Agent"], # the entire built-in toolset: no Bash, Edit or Write allowed_tools=["mcp__ci__get_failed_log"], # reads inside cwd and Agent calls need no rule permission_mode="dontAsk", # anything not approved is denied, never prompted setting_sources=[], # load NO settings from the untrusted checkout: its hooks would run with your API key mcp_servers={"ci": create_sdk_mcp_server(name="ci", version="1.0.0", tools=[get_failed_log])}, strict_mcp_config=True, # ignore any .mcp.json the checked-out branch brings along agents={ "log-reader": AgentDefinition( description="Reads the CI failure log and returns only the first real error.", prompt=( "Call get_failed_log with max_lines=3000. Return the first error that is not a " "consequence of an earlier one, quoted verbatim with 20 lines of context, and the " "test or step name. Do not speculate about causes." ), tools=["mcp__ci__get_failed_log"], model="haiku", ) }, hooks={"PreToolUse": [HookMatcher(matcher="Read|Grep|Glob", hooks=[deny_secret_reads])]}, output_format={"type": "json_schema", "schema": VERDICT}, max_turns=30, max_budget_usd=2.0,)
PROMPT = ( "CI run failed. Use the log-reader subagent to get the first real error, then find the code it " "points at. Classify the failure. Every evidence item must be a verbatim log line or a " "path:line you opened. If you cannot tie the error to a file, answer unknown; do not guess.")
# 4. Deterministic check: every evidence item must exist in the log or in the checkout.MIN_EVIDENCE = 20 # "Error" or "FAIL" appear in every failed log and prove nothing
def grounded(item: str, log: str) -> bool: item = item.strip() if len(item) >= MIN_EVIDENCE and item in log: return True m = re.fullmatch(r"([\w./-]+):(\d+)", item) if not m: return False path = (REPO / m[1]).resolve() if not (path.is_file() and path.is_relative_to(REPO)): return False return int(m[2]) <= len(path.read_text(errors="replace").splitlines())
async def main() -> int: async with ClaudeSDKClient(options=options) as client: await client.query(PROMPT) async for msg in client.receive_response(): if not isinstance(msg, ResultMessage): continue print(f"subtype={msg.subtype} cost_usd={msg.total_cost_usd} turns={msg.num_turns}", file=sys.stderr) if msg.is_error or msg.structured_output is None: return 1 # no verdict, no comment verdict = msg.structured_output log = await failed_log() if not all(grounded(e, log) for e in verdict["evidence"]): verdict |= {"category": "unknown", "confidence": "low"} Path("verdict.json").write_text(json.dumps(verdict, indent=2)) return 0 return 1
if __name__ == "__main__": sys.exit(anyio.run(main))// .github/scripts/triage.mts: the same bot, including the grounding check.import { execFile } from 'node:child_process';import { existsSync, readFileSync, realpathSync, statSync } from 'node:fs';import { writeFile } from 'node:fs/promises';import { resolve, sep } from 'node:path';import { promisify } from 'node:util';import { createSdkMcpServer, query, tool, type HookCallback } from '@anthropic-ai/claude-agent-sdk';import { z } from 'zod';
const run = promisify(execFile);const RUN_ID = process.env.RUN_ID!;const REPO = realpathSync(resolve(process.env.TARGET_DIR ?? '.')); // the untrusted checkout: data, never code
async function failedLog(): Promise<string> { const { stdout } = await run('gh', ['run', 'view', RUN_ID, '--log-failed'], { maxBuffer: 64 * 1024 * 1024 }); return stdout;}
const getFailedLog = tool( 'get_failed_log', 'Log of the failed steps in the CI run under triage', { max_lines: z.number().int().positive() }, async ({ max_lines }) => ({ content: [{ type: 'text', text: (await failedLog()).split('\n').slice(-max_lines).join('\n') }], }),);
const denySecretReads: HookCallback = async (input) => { if (input.hook_event_name !== 'PreToolUse') return {}; const { file_path, path, glob, pattern } = input.tool_input as Record<string, string | undefined>; const targets = [file_path, path, glob, input.tool_name === 'Glob' ? pattern : undefined]; // A Grep over the whole cwd has no path to match: the checkout itself must hold no secrets. if (targets.some((t) => /(^|\/)(\.env[^/]*|secrets?)(\/|$)/.test(t ?? ''))) { return { hookSpecificOutput: { hookEventName: 'PreToolUse', permissionDecision: 'deny', permissionDecisionReason: 'Triage never reads secret files.', }, }; } return {};};
const verdictSchema = { type: 'object', properties: { category: { enum: ['test-regression', 'flaky-test', 'build-config', 'infrastructure', 'unknown'] }, culprit_file: { type: ['string', 'null'] }, evidence: { type: 'array', items: { type: 'string' }, minItems: 1 }, confidence: { enum: ['high', 'medium', 'low'] }, summary: { type: 'string' }, }, required: ['category', 'culprit_file', 'evidence', 'confidence', 'summary'], additionalProperties: false,};
// Deterministic check: every evidence item must exist in the log or in the checkout.const MIN_EVIDENCE = 20; // "Error" or "FAIL" appear in every failed log and prove nothing
function grounded(item: string, log: string): boolean { const s = item.trim(); if (s.length >= MIN_EVIDENCE && log.includes(s)) return true; const m = /^([\w./-]+):(\d+)$/.exec(s); if (!m || !existsSync(resolve(REPO, m[1]))) return false; const path = realpathSync(resolve(REPO, m[1])); if (!path.startsWith(REPO + sep) || !statSync(path).isFile()) return false; return Number(m[2]) <= readFileSync(path, 'utf8').replace(/\n$/, '').split('\n').length;}
for await (const msg of query({ prompt: 'CI run failed. Use the log-reader subagent to get the first real error, then find the code it points at. ' + 'Classify the failure. Every evidence item must be a verbatim log line or a path:line you opened. ' + 'If you cannot tie the error to a file, answer unknown; do not guess.', options: { systemPrompt: { type: 'preset', preset: 'claude_code', append: 'You are a CI triage bot. You never change files.' }, cwd: REPO, tools: ['Read', 'Grep', 'Glob', 'Agent'], allowedTools: ['mcp__ci__get_failed_log'], // reads inside cwd and Agent calls need no rule permissionMode: 'dontAsk', settingSources: [], // load NO settings from the untrusted checkout: its hooks would run with your API key mcpServers: { ci: createSdkMcpServer({ name: 'ci', version: '1.0.0', tools: [getFailedLog] }) }, strictMcpConfig: true, agents: { 'log-reader': { description: 'Reads the CI failure log and returns only the first real error.', prompt: 'Call get_failed_log with max_lines=3000. Return the first error that is not a consequence of an ' + 'earlier one, quoted verbatim with 20 lines of context, and the test or step name. Do not speculate.', tools: ['mcp__ci__get_failed_log'], model: 'haiku', }, }, hooks: { PreToolUse: [{ matcher: 'Read|Grep|Glob', hooks: [denySecretReads] }] }, outputFormat: { type: 'json_schema', schema: verdictSchema }, maxTurns: 30, maxBudgetUsd: 2, },})) { if (msg.type === 'result') { console.error(`cost=$${msg.total_cost_usd} turns=${msg.num_turns} subtype=${msg.subtype}`); if (msg.subtype !== 'success' || msg.is_error || msg.structured_output == null) process.exit(1); const verdict = msg.structured_output as { category: string; confidence: string; evidence: string[] }; const log = await failedLog(); if (!verdict.evidence.every((e) => grounded(e, log))) Object.assign(verdict, { category: 'unknown', confidence: 'low' }); await writeFile('verdict.json', JSON.stringify(verdict, null, 2)); }}This version type-checks under tsc --strict against 0.3.283.
The MCP tool name follows the pattern mcp__<server>__<tool>, where <server> is the key in mcp_servers (ci), not the name passed to create_sdk_mcp_server. Get the key wrong and allowed_tools approves a tool that does not exist, which dontAsk then silently denies.
Run the triage bot in GitHub Actions
Section titled “Run the triage bot in GitHub Actions”The workflow fires when your CI workflow completes with a failure. It runs the triage script from the default branch, checks out the failing commit into a separate untrusted/ directory that the agent only reads, and posts the verdict only when the script produced one.
name: ci-triageon: workflow_run: workflows: [CI] types: [completed]permissions: actions: read contents: read pull-requests: writejobs: triage: if: github.event.workflow_run.conclusion == 'failure' runs-on: ubuntu-latest timeout-minutes: 15 env: GH_TOKEN: ${{ github.token }} GH_REPO: ${{ github.repository }} steps: # Trusted code: the triage script comes from the default branch, never from the failing commit. - uses: actions/checkout@v7 with: ref: ${{ github.event.repository.default_branch }} persist-credentials: false # Untrusted data: the failing commit, checked out beside it and only ever read by the agent. - uses: actions/checkout@v7 with: ref: ${{ github.event.workflow_run.head_sha }} path: untrusted persist-credentials: false - uses: actions/setup-python@v7 with: python-version: '3.12' - run: pip install claude-agent-sdk==0.2.160 - name: Triage env: ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }} RUN_ID: ${{ github.event.workflow_run.id }} TARGET_DIR: untrusted run: python .github/scripts/triage.py - name: Comment on the pull request if: github.event.workflow_run.pull_requests[0] != null run: | jq -r '"**CI triage:** \(.category) (\(.confidence) confidence)\n\n\(.summary)\n\n" + (.evidence | map("- `" + . + "`") | join("\n"))' verdict.json > comment.md gh pr comment ${{ github.event.workflow_run.pull_requests[0].number }} --body-file comment.mdFor pull requests from forks, workflow_run.pull_requests is empty, so the step skips; look up the PR with gh pr list --search <head_sha> --state open --json number if you need to comment there.
workflow_run runs with the default branch’s workflow file and its secrets, while the failing commit can come from any contributor. So nothing from that commit ever executes: the script comes from the default branch, setting_sources=[] keeps the commit’s .claude/settings.json hooks and .mcp.json from loading, and the agent gets no Bash, no write tools and no credentials in .git/config. Running a script or loading settings from the failing commit with ANTHROPIC_API_KEY in the environment would let any pull request exfiltrate the key. Read the agent threat model before you widen any of those.
How do the Agent SDK’s permission options combine?
Section titled “How do the Agent SDK’s permission options combine?”The options overlap, and most permission bugs come from treating one as another. Claude Code evaluates each tool call in this order: hooks, deny rules, ask rules, the permission mode, allow rules, then your can_use_tool callback.
| Option (Python / TypeScript) | What it does | Common mistake |
|---|---|---|
tools / tools | Sets the base built-in toolset. [] disables all built-in tools | Leaving it unset and trying to restrict with allowed_tools |
allowed_tools / allowedTools | Allow rules: listed calls are auto-approved | Believing it hides unlisted tools. The README: it “does not remove tools from Claude’s toolset” |
disallowed_tools / disallowedTools | Deny rules. A bare name (Bash) removes the tool; a scoped rule (Bash(rm *)) blocks matching calls in every mode | Expecting Bash(rm *) to catch /bin/rm; it matches the command as written |
permission_mode / permissionMode | default, acceptEdits, plan, dontAsk, auto, bypassPermissions | Using bypassPermissions in CI. TypeScript also requires allowDangerouslySkipPermissions: true for it |
hooks / hooks | Your function runs before the rules. A deny holds in every mode; an allow does not skip later deny or ask rules | Assuming a hook’s allow overrides a deny rule |
can_use_tool / canUseTool | Answers calls that nothing earlier resolved, in place of an interactive prompt | Expecting it to see every call. It never sees pre-approved calls, and dontAsk skips it entirely |
For an unattended bot, the reliable pattern is the one above: shrink tools, pick dontAsk, approve only the extra tools the job needs, and put every “never” rule in a hook or a deny rule. For a bot with a human reachable (a Slack approver, for example), use default mode with a can_use_tool callback that forwards the request and waits for an answer.
Hooks differ between the two SDKs. The TypeScript SDK accepts the full Claude Code event set; the Python SDK accepts a subset (PreToolUse, PostToolUse, PostToolUseFailure, UserPromptSubmit, Stop, SubagentStart, SubagentStop, PreCompact, Notification and PermissionRequest in 0.2.160), with no SessionStart or SessionEnd. If your design depends on a session hook, write it in TypeScript. For shell-command hooks shared with interactive use, see the Claude Code hooks guide.
Sessions: resume, fork and why CI loses them
Section titled “Sessions: resume, fork and why CI loses them”Every ResultMessage carries a session_id. Pass it back as resume to continue with the full context, and add fork_session=True to branch into a new session ID without changing the original. That enables a two-phase pattern: a read-only triage session, then a fork of it in acceptEdits mode that drafts a fix on a branch, reusing everything the triage already read.
The catch is where sessions live. Claude Code writes transcripts to ~/.claude/projects/<encoded-cwd>/<session-id>.jsonl on the machine that ran them, so a session created in one GitHub Actions job is gone when the next job starts on a fresh runner. You have three options:
- Run both phases in one job.
- Upload the
.jsonlfile as an artifact and restore it into~/.claude/projects/before callingresume. - Attach a
session_store(Python) orsessionStore(TypeScript) adapter that mirrors transcripts to your own storage, and resume from the samecwd.
For fire-and-forget runs where nobody will resume, TypeScript’s persistSession: false skips writing the transcript at all. In Python, set CLAUDE_CODE_SKIP_PROMPT_HISTORY in the env option instead (agent-sdk sessions doc).
How do you prove the triage bot is right?
Section titled “How do you prove the triage bot is right?”The bot is useful only if its verdicts are right more often than a tired human skimming the log, and you need to know that without reading every verdict. Four layers do the checking, and each one fails loudly.
- Schema. The CLI validates the verdict against the JSON schema. A run that cannot produce one exits non-zero and posts nothing.
- Grounding.
grounded()checks that every evidence string is at least 20 characters long and appears verbatim in the log, or is apath:linethat exists inside the checkout. An invented file or a paraphrased log line turns the verdict intounknownwith low confidence. This check is ordinary code, so you can unit-test it. - Budget and turn caps. A confused run ends with
error_max_budget_usdorerror_max_turnsinstead of burning money; the cost of each run is printed fromtotal_cost_usdinto the job log. - An eval set. Label 20 past failed runs with their known root cause, replay the bot against them, and compare. Run this locally with your own API key:
# evals/triage-cases.tsv holds lines of: <run id> <TAB> <expected category># Each case gets its own worktree; triage.py always comes from your current branch.while IFS=$'\t' read -r run expected; do sha=$(gh run view "$run" --json headSha -q .headSha) # Fork heads and force-pushed commits are often missing from a local clone. git fetch --quiet origin "$sha" || { echo -e "$run\t$expected\tskip\tFETCH-FAIL"; continue; } git worktree add --quiet --detach "/tmp/case-$run" "$sha" rm -f verdict.json if RUN_ID="$run" TARGET_DIR="/tmp/case-$run" python .github/scripts/triage.py 2>>costs.log; then got=$(jq -r .category verdict.json) else got=no-verdict fi git worktree remove --force "/tmp/case-$run" echo -e "$run\t$expected\t$got\t$([ "$got" = "$expected" ] && echo PASS || echo FAIL)"done < evals/triage-cases.tsv | tee results.tsvawk -F'\t' '$4=="PASS"{p++} END{print p+0"/"NR" passed"}' results.tsvRerun the eval whenever you change the prompt, the schema, the subagent’s model or the SDK version, and keep the results with the change. If you already use promptfoo, its anthropic:claude-agent-sdk provider runs the SDK as the system under test in the same way.
Who signs off: the tech lead who owns CI decides the pass rate at which the bot’s comments are trusted, and until the eval clears that bar the comment stays advisory and never gates a merge. When a verdict turns out wrong in production, add that run to triage-cases.tsv. This is the evidence-over-diffs habit applied to an agent: you read the eval table and the grounding result, not the agent’s transcript.
Copy-paste prompts for Agent SDK work
Section titled “Copy-paste prompts for Agent SDK work”Paste these into Claude Code in the repository that will host the bot.
What breaks when you drive Claude Code from the Agent SDK?
Section titled “What breaks when you drive Claude Code from the Agent SDK?”| Symptom | Cause | Recovery |
|---|---|---|
| The agent behaves like a generic chatbot, ignores tool conventions | system_prompt omitted, so the CLI started with an empty prompt | Pass the claude_code preset with append |
CLAUDE.md rules are ignored in CI, or a developer’s personal settings show up | setting_sources excludes "project", or is unset and loads the runner’s user settings | Set ["project"] explicitly for trusted checkouts; for untrusted ones keep [] and pass the rules in system_prompt instead |
A tool in allowed_tools is never called; permission_denials lists it | MCP tool name uses the server’s name, not its mcp_servers key | Rename to mcp__<key>__<tool> |
| The run hangs or every write is denied in CI | No human for the default mode’s prompts | Use dontAsk and approve what the job needs, or supply can_use_tool |
subtype is error_max_structured_output_retries | The schema is stricter than the evidence allows, for example culprit_file required as a string when there is none | Allow null or an unknown category, as the example does |
resume fails with an unknown session in a later job | Transcripts are local to the runner that created them | Same job, restored .jsonl artifact, or a session store |
TypeScript: “Claude Code native binary not found at …”; Python: CLINotFoundError | No usable CLI binary: in TypeScript, optional dependencies were skipped, so no platform binary | Install without --omit=optional, or set pathToClaudeCodeExecutable (TypeScript) or cli_path (Python) |
| A contributor’s branch adds an MCP server that the bot then loads | Project .mcp.json loads without a prompt in SDK sessions | strict_mcp_config=True and servers passed only in code |
The log tool fails (for example, gh missing on the runner) | Environment, not the agent | The prompt tells the agent to answer unknown; a test run on 2026-09-26 with gh missing did so instead of guessing. Fix the runner and rerun |
Where to go next with the Agent SDK
Section titled “Where to go next with the Agent SDK”For the model behind the haiku alias and current prices, see the models hub. For MCP servers beyond the in-process kind, see MCP setup for Claude Code. The Codex counterpart is building with the Codex SDK, and the Cursor one is the Cursor SDK.