Driving agents from code: Claude Agent SDK, Codex SDK and Cursor SDK
The Claude Agent SDK, the Codex SDK and the Cursor SDK let a program, not a person, run a coding agent: start it, restrict its tools, read a typed result and resume it later. Use an SDK when your code must branch on what the agent found; for a one-shot CI step, claude -p or codex exec is enough.
CI goes red at 02:10. By the time anyone looks, three more pushes have queued behind it, and the first half hour of the morning goes on reading 4,000 lines of log to learn that a test was flaky again. You already tried a bot: a workflow step that pipes the log into claude -p and posts the answer as a comment. It helps, but it cannot rerun the flaky test and wait, it cannot open a second, write-enabled session only when the verdict is “regression”, and nobody knows what it cost last month. That is the point where you stop triggering an agent and start driving one from code.
This page is for developers who build that automation and for tech leads who decide which SDK a team standardises on. It compares the three SDKs on the five things that decide the choice — sessions, tools, permissions, streaming and cost — and writes the same job in each.
What you’ll walk away with from the agent SDK comparison
Section titled “What you’ll walk away with from the agent SDK comparison”- A decision table for “trigger it or script it” that you can apply to any automation idea.
- A comparison of the three SDKs on sessions, tools, permissions, streaming, structured output and cost, checked against the published packages on 2026-09-26.
- A CI-failure triage runner written three times — Claude Agent SDK, Codex SDK, Cursor SDK — sharing one prompt and one JSON Schema, plus the GitHub Actions workflow that calls it.
- The gates that make its output trustworthy without a person reading every run, and the failure modes that bite first.
Should you script the agent or trigger it?
Section titled “Should you script the agent or trigger it?”Most “agent in CI” jobs never need an SDK. A trigger — a GitHub Action, a scheduled automation or a single headless CLI call — takes a prompt and produces a comment, a file or a pull request. An SDK earns its extra code only when the orchestration logic lives in your program.
| Your job needs… | Trigger is enough | Reach for the SDK |
|---|---|---|
| One prompt in, one result out | claude -p --output-format json, codex exec --json -o out.json, Cursor’s print mode, anthropics/claude-code-action@v1, openai/codex-action@v1, Cursor Automations | — |
| A typed result a script consumes | claude -p --json-schema, codex exec --output-schema FILE | When the next step depends on the value |
| Different next steps per result (rerun, label, open a fix session) | Only with brittle shell glue | Yes: branch in code, resume the same session |
| Tools that call your own systems (ticketing, feature flags, internal APIs) | An MCP server you deploy and maintain | In-process tools: Claude createSdkMcpServer, Cursor customTools |
A per-call permission decision (“allow git show, deny git push, ask the service for the rest”) | Static allow and deny lists | Claude’s canUseTool callback and hooks |
| Live progress in your own UI, bot or log | Tail a JSONL file | Typed event streams in all three SDKs |
| Spend accounting per run, per team or per customer | Scrape the CLI’s JSON | Cost or usage fields on the result |
| Many agents coordinated by your code | — | Yes; see multi-agent orchestration patterns |
The rule of thumb: if you can describe the job as a single prompt with a fixed output, keep it a trigger. If you catch yourself writing if statements around the agent’s answer, move to an SDK. The CLI-level options are covered per tool in Claude Code automation and Codex non-interactive mode.
How do the Claude Agent SDK, Codex SDK and Cursor SDK differ?
Section titled “How do the Claude Agent SDK, Codex SDK and Cursor SDK differ?”All three put the vendor’s own coding agent behind a library: the same tools, the same project instruction files and the same model catalogue you use interactively. They differ in what they let your code decide. Everything below was read from the published TypeScript packages on 2026-09-26.
| Claude Agent SDK | Codex SDK | Cursor SDK | |
|---|---|---|---|
| Packages checked | npm @anthropic-ai/claude-agent-sdk 0.3.283, PyPI claude-agent-sdk 0.2.160 | npm @openai/codex-sdk 0.157.1, PyPI openai-codex 0.157.1 | npm @cursor/sdk 1.0.32, PyPI cursor-sdk 1.0.32 |
| What runs underneath | The Claude Code binary, installed as a platform package | The codex CLI from @openai/codex, spoken to over JSONL on stdin and stdout | A local agent, or a cloud agent in an isolated VM when you pass cloud |
| Runtime | Node 18+ | Node 18+ | Node 22.13+ |
| Entry point | query({ prompt, options }), an async generator of messages | new Codex().startThread(), then thread.run() or thread.runStreamed() | Agent.create(), then agent.send() returning a Run |
| Sessions | session_id on every message; resume, continue, forkSession | Thread id; codex.resumeThread(id); threads persist in ~/.codex/sessions | agentId; Agent.resume(id); local agents persist in a SQLite or JSONL store, cloud agent IDs start with bc- |
| Restricting tools | tools sets the toolset; allowedTools only auto-approves; disallowedTools removes | No per-tool list; the sandbox mode bounds what commands can do | tools and disallowedTools (local agents only) |
| Your own tools | In-process MCP server via createSdkMcpServer() and tool(), plus mcpServers | MCP servers from Codex configuration | customTools callbacks (local only), plus mcpServers |
| Permissions | permissionMode (default, acceptEdits, plan, dontAsk, auto, bypassPermissions), canUseTool callback, hooks such as PreToolUse | sandboxMode (read-only, workspace-write, danger-full-access) plus approvalPolicy (never, on-request, on-failure, untrusted) | local.sandboxOptions, local.autoReview (Cursor’s classifier-backed Auto-review); a cloud agent is isolated by its VM |
| Streaming | Every message as it happens; includePartialMessages adds token deltas | runStreamed() yields thread.started, item.*, turn.completed, turn.failed, error | run.stream(), plus onStep and onDelta callbacks on send() |
| Structured output | outputFormat: { type: 'json_schema', schema } → structured_output on the result | outputSchema per turn → JSON in finalResponse | No schema option in 1.0.32; capture the report through a custom tool |
| Limits and cost | maxTurns, maxBudgetUsd; total_cost_usd on the result, documented as an estimate | Token usage on turn.completed; no dollar field and no spend cap; cancel with an AbortSignal | agent.getUsage() returns tokens and cost in cents, which can lag the run; cancel with run.cancel() |
| Model | Claude Code’s default unless you set model | Codex’s default unless you set model; modelReasoningEffort | model is required for local agents; list IDs with Cursor.models.list() |
Three consequences are easy to miss in the table:
- Only Claude lets code decide each tool call.
canUseToolandPreToolUsehooks see every call before it runs. The Codex SDK gives you a sandbox boundary instead of a callback (no per-call hook in@openai/codex-sdk0.157.1), and the Cursor SDK gives you tool lists and Auto-review. - The Agent SDK does not start in auto mode. From Claude Code v2.1.283 (the
latestchannel), interactive sessions start in auto mode, but the Agent SDK andclaude -pstill start in the Manual (default) permission mode. An unattended job that relies on prompts being answered stalls or is denied, so pickdontAskwith an explicit allowlist. - Cursor’s SDK is the only one that can hand the work to a cloud VM. Passing
cloud: { repos: [...] }runs the agent on Cursor’s infrastructure and can open the pull request itself (autoCreatePR). Tool lists, custom tools and a custom system prompt are local-only in 1.0.32, and combiningtoolswithcloudthrows aConfigurationError.
Model choice is not a reason to pick an SDK: each defaults to its tool’s default model. Start there, tune effort before you switch models, and check the current names and prices on the models hub.
The same job three times: triage a failing CI run
Section titled “The same job three times: triage a failing CI run”The job: when the CI workflow fails, download the failed steps’ log, find the first real failure, reproduce it, rerun it to rule out flakiness, and write a triage.json with a category (regression, flaky, infra, test-bug), the failing tests, a suspect commit, evidence and a next step. The agent must not edit anything. A later step decides what to do with the verdict.
The three runners share one prompt and one schema, so a switch between vendors changes only the runner file.
export const triageSchema = { type: 'object', properties: { category: { type: 'string', enum: ['regression', 'flaky', 'infra', 'test-bug'] }, failing_tests: { type: 'array', items: { type: 'string' } }, suspect_commit: { type: 'string' }, evidence: { type: 'string' }, next_step: { type: 'string' }, }, required: ['category', 'failing_tests', 'suspect_commit', 'evidence', 'next_step'], additionalProperties: false,};
export const triagePrompt = `CI failed on this commit. The job log is in ci-failure.log.Triage it; do not fix anything and do not edit any file.1. Find the first real failure in the log (skip cascading errors).2. Reproduce only the failing test(s) with: npx vitest run <file>.3. Run it twice more. Passes on a rerun => "flaky".4. Network, runner, secret or cache errors => "infra".5. Otherwise read git log -5 and the diff of the suspect commit; decide "regression" (code broke) or "test-bug" (test is wrong).Cite log lines and command output as evidence. Use "unknown" forsuspect_commit when you cannot name one.`;The Cursor SDK has no output-schema option, so the runner gives the agent a submit_triage custom tool whose input schema is the triage schema. The tool call’s arguments are the report. A local agent needs an explicit model ID; list the IDs your key can use with Cursor.models.list() and store one in CURSOR_MODEL_ID.
import { writeFile } from 'node:fs/promises';import { Agent, type SDKJsonValue } from '@cursor/sdk';import { triagePrompt, triageSchema } from './triage-spec.js';
let report: Record<string, SDKJsonValue> | undefined;
const agent = await Agent.create({ apiKey: process.env.CURSOR_API_KEY, model: { id: process.env.CURSOR_MODEL_ID! }, // required for local agents disallowedTools: ['edit', 'delete', 'applyAgentDiff'], local: { cwd: process.env.GITHUB_WORKSPACE, settingSources: ['project'], customTools: { submit_triage: { description: 'Submit the final triage report. Call it exactly once, at the end.', inputSchema: triageSchema, execute: (args) => { report = args; return 'recorded'; }, }, }, },});
try { const run = await agent.send(`${triagePrompt}\nFinish by calling submit_triage.`); const timer = setTimeout(() => void run.cancel(), 10 * 60_000); for await (const message of run.stream()) { if (message.type === 'tool_call' && message.status === 'completed') console.log('tool', message.name); } const result = await run.wait(); clearTimeout(timer); if (result.status !== 'finished' || !report) throw new Error(`triage ${result.status}, no report`); await writeFile('triage.json', JSON.stringify(report, null, 2)); const usage = await agent.getUsage(); // cost can lag the run console.log(`agent ${agent.agentId}`, usage.usage, usage.cost);} finally { agent.close();}disallowedTools removes the edit tools, but shell stays, and a shell can still write files. The git diff --exit-code step in the workflow below is what enforces “no edits”.
The Claude Agent SDK has every control this job needs in one options object: a restricted toolset, an allowlist of Bash commands, a mode that denies everything else, a turn and dollar cap, and schema-checked output.
import { writeFile } from 'node:fs/promises';import { query } from '@anthropic-ai/claude-agent-sdk';import { triagePrompt, triageSchema } from './triage-spec.js';
for await (const msg of query({ prompt: triagePrompt, options: { cwd: process.env.GITHUB_WORKSPACE, tools: ['Read', 'Grep', 'Glob', 'Bash'], // no Edit, no Write allowedTools: ['Read', 'Grep', 'Glob', 'Bash(npx vitest run *)', 'Bash(git log *)', 'Bash(git show *)'], permissionMode: 'dontAsk', // anything not pre-approved is denied settingSources: ['project'], // CLAUDE.md, not the runner's ~/.claude maxTurns: 30, maxBudgetUsd: 2, outputFormat: { type: 'json_schema', schema: triageSchema }, },})) { if (msg.type !== 'result') continue; if (msg.subtype !== 'success') throw new Error(`triage stopped: ${msg.subtype}`); await writeFile('triage.json', JSON.stringify(msg.structured_output, null, 2)); console.log(`session ${msg.session_id}, ~$${msg.total_cost_usd.toFixed(2)}, ${msg.num_turns} turns`);}allowedTools does not shrink the toolset; it only pre-approves. The tools line is what removes Edit and Write. A run that hits the caps ends with subtype set to error_max_turns or error_max_budget_usd, which the runner turns into a failed step.
The Codex SDK bounds the agent with a sandbox rather than a tool list, and it takes the schema per turn. There is no spend cap, so a timeout signal stands in for one.
import { writeFile } from 'node:fs/promises';import { Codex } from '@openai/codex-sdk';import { triagePrompt, triageSchema } from './triage-spec.js';
const codex = new Codex({ apiKey: process.env.CODEX_API_KEY });const thread = codex.startThread({ workingDirectory: process.env.GITHUB_WORKSPACE, sandboxMode: 'workspace-write', // tests may write caches; the git-diff gate catches edits approvalPolicy: 'never', // nobody is there to approve networkAccessEnabled: false,});
const timeout = AbortSignal.timeout(10 * 60_000);const { events } = await thread.runStreamed(triagePrompt, { outputSchema: triageSchema, signal: timeout });
let report = '';for await (const event of events) { if (event.type === 'item.completed' && event.item.type === 'command_execution') { console.log(`$ ${event.item.command} -> exit ${event.item.exit_code}`); } if (event.type === 'item.completed' && event.item.type === 'agent_message') report = event.item.text; if (event.type === 'turn.failed') throw new Error(event.error.message); if (event.type === 'error') throw new Error(event.message); if (event.type === 'turn.completed') console.log('tokens', event.usage);}if (!report) throw new Error('triage produced no report');await writeFile('triage.json', report); // JSON, because outputSchema was setconsole.log(`thread ${thread.id}`);read-only is the stricter sandbox, but a test runner that writes a cache or a coverage file fails under it and the agent then reports a false “infra”. workspace-write with the network off plus the diff gate is the practical middle.
All three files type-check with TypeScript 5 (strict, NodeNext) against the package versions in the comparison table. Run them with npx tsx ci/triage-<tool>.ts.
Wire the runner into GitHub Actions
Section titled “Wire the runner into GitHub Actions”The workflow is the same for all three SDKs apart from the last run line and its secret.
name: ci-triageon: workflow_run: workflows: [CI] types: [completed]permissions: contents: read actions: readjobs: triage: # Only failed runs, and never code from a fork: this job holds API keys. if: >- github.event.workflow_run.conclusion == 'failure' && github.event.workflow_run.head_repository.full_name == github.repository runs-on: ubuntu-latest timeout-minutes: 15 steps: - uses: actions/checkout@v4 with: ref: ${{ github.event.workflow_run.head_sha }} fetch-depth: 20 - uses: actions/setup-node@v4 with: node-version: 22 # the Cursor SDK needs 22.13+ - run: npm ci - run: gh run view ${{ github.event.workflow_run.id }} --log-failed > ci-failure.log env: GH_TOKEN: ${{ github.token }} - run: npx tsx ci/triage-claude.ts # or triage-codex.ts / triage-cursor.ts env: ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }} - run: git diff --exit-code # the triage must leave tracked files untouched - uses: actions/upload-artifact@v4 with: name: triage path: triage.jsonFor the Codex runner, pass CODEX_API_KEY; for the Cursor runner, pass CURSOR_API_KEY and CURSOR_MODEL_ID. Check your vendor’s terms for running unattended on an API key versus a subscription login before you pick the credential.
Branch on the verdict: where the SDK pays for itself
Section titled “Branch on the verdict: where the SDK pays for itself”Up to here, claude -p --json-schema or codex exec --output-schema would have done the same job. The SDK becomes worth its code when the next step depends on the answer and reuses the agent’s context:
-
Read
triage.jsonand validate it againsttriageSchemawith a JSON Schema validator. A report that fails validation fails the job; nobody acts on it. -
On
flaky, rerun the failed jobs (gh run rerun <run-id> --failed; this step needsactions: writeon its job) and add the test to your quarantine list. No second agent session. -
On
infra, post the evidence to the on-call channel and stop. -
On
regressionortest-bug, resume the same session with write access in a new branch, so the agent keeps what it already read:- Claude Agent SDK:
query({ prompt, options: { resume: sessionId, tools: { type: 'preset', preset: 'claude_code' }, permissionMode: 'acceptEdits' } }). - Codex SDK:
codex.resumeThread(threadId, { sandboxMode: 'workspace-write' }), thenthread.run(...). - Cursor SDK:
agent.send(...)on the same agent, orAgent.resume(agentId)from a later job; the tool restrictions are not persisted, so pass them again.
- Claude Agent SDK:
-
The fix session opens a pull request that carries the triage report as evidence, and it goes through the same gates as any other agent change: see the evidence bundle an agent’s pull request must carry.
The GitHub Actions runner is ephemeral, so a session saved on it is gone when the job ends. To resume in a later job, persist the session store (~/.claude/projects, ~/.codex/sessions or the Cursor local store) as an artifact, or keep the whole flow in one job.
Copy-paste prompts for building an SDK runner
Section titled “Copy-paste prompts for building an SDK runner”Run the audit prompt with a different model or vendor than the runner, so the author never grades itself; model-graded checks covers how to calibrate such a judge.
How do you prove the triage is right without reading every run?
Section titled “How do you prove the triage is right without reading every run?”Treat the runner like any other production code path whose output other automation acts on. The gates, in the order they fire:
- Schema. The SDK validates the output shape (Claude, Codex) or the custom tool’s input schema constrains it (Cursor). Your step validates it again before acting, because a run that hits a cap can end with no report at all.
- No side effects.
git diff --exit-codefails the job if the agent changed a tracked file. Tool restrictions reduce the chance; the diff proves it. - Caps.
maxTurnsandmaxBudgetUsdon Claude, a timeout on Codex and Cursor, andtimeout-minuteson the job. A capped run is a failed run, never a partial verdict. - A golden set. Collect past failed runs with a label a person already agreed on, including flaky, infra and regression cases. Replay the runner on them after every SDK, model or prompt change, and track agreement. Decide in advance what agreement you require before the runner may rerun jobs or label issues on its own; until then it only comments.
- An independent audit. The audit prompt above, run by another model, rejects verdicts whose evidence does not appear in the log.
- Telemetry. Log the session or thread ID, turns, tokens and cost for every run, and join them to the CI run ID. The agent observability page shows the dashboard: run success, cost per accepted verdict and override rate.
Who signs off: the on-call developer owns every action the runner proposes until the golden-set agreement meets the bar the tech lead set. After that, the tech lead signs off on widening its authority one action at a time, starting with reruns of flaky tests.
What breaks when you drive agents from an SDK?
Section titled “What breaks when you drive agents from an SDK?”| Symptom | Cause | Recovery |
|---|---|---|
The Claude run ends with error_max_budget_usd or error_max_turns and no report | The cap is too low for the repository’s log size, or the agent is looping on a noisy log | Trim the log to the failed step before the run (--log-failed already helps), then raise one cap at a time and record the new cost in the golden-set run |
| The Codex run hangs until the timeout | A command waited for input or the network with networkAccessEnabled: false | Read the last command_execution item in the streamed log; add the non-interactive flag to the reproduce command in the prompt |
| Every Codex verdict is “infra” | read-only sandbox, and the test runner needs to write a cache | Switch to workspace-write and keep the diff gate |
Cursor’s Agent.create throws ConfigurationError | tools, disallowedTools or systemPrompt combined with cloud (and customTools is local-only too) | Run the triage as a local agent, or drop those options for the cloud agent and restrict it through the prompt and the repository access you grant |
Cursor reports finished but report is undefined | The agent answered in text and never called submit_triage | Treat it as a failure (the runner already throws); repeat the instruction to call the tool at the end of the prompt, and send one follow-up on the same agent asking for the call |
| Cost on the Cursor usage call is missing | Billing data lags the run | Read agent.getUsage() again in a later step, or join the usage export by agent ID |
| A resumed session in a later job is “not found” | The session lived on the previous ephemeral runner | Upload the session store as an artifact, or keep triage and fix in one job |
| The SDK and a globally installed CLI behave differently | The SDK drives its own pinned binary, and your global CLI is another version | Pin the SDK version in package.json and upgrade on purpose; the Claude Agent SDK accepts pathToClaudeCodeExecutable when you must use a specific binary |
Which SDK should a team standardise on?
Section titled “Which SDK should a team standardise on?”Pick by where the team already works and by the control the job needs, not by the model: each SDK drives its own vendor’s agent and default model.
- Choose the Claude Agent SDK when the job needs per-call permission decisions, hooks, in-process tools and a hard dollar cap in the same place. It has the widest control surface of the three.
- Choose the Codex SDK when the team already runs Codex and a sandbox boundary is the control you trust. It is the thinnest wrapper, and its per-turn schema makes multi-step pipelines on one thread straightforward.
- Choose the Cursor SDK when you want the agent to run on Cursor’s cloud VMs and open the pull request itself, or when the team’s rules and skills already live in Cursor.
For a lead, the standardisation decision is mostly about the shared parts: one prompt and schema file per job, one set of gates, one golden set and one cost dashboard. With those in place, switching the runner file is a small change, as the three tabs above show.
Where to go next with agent SDKs
Section titled “Where to go next with agent SDKs”Frequently asked questions
When should I use an agent SDK instead of claude -p or codex exec?
Use the CLI when the job is one prompt in and one result out. Use an SDK when your code must branch on the agent's result, resume the same session, register in-process tools, decide permissions per tool call, or stream progress into your own UI.
Which agent SDKs return structured output against a JSON Schema?
The Claude Agent SDK (outputFormat with type json_schema) and the Codex SDK (outputSchema per turn) do. Cursor SDK 1.0.32 has no schema option; a local custom tool that receives the report as its arguments gives you the same result.
How do I cap what an SDK-driven agent spends?
The Claude Agent SDK has maxBudgetUsd and maxTurns. The Codex SDK has no spend cap, so pass an AbortSignal with a timeout. The Cursor SDK has no cap either; cancel the run on a timer and read agent.getUsage() afterwards.