Skip to content

Driving agents from code: Claude Agent SDK, Codex SDK and Cursor SDK

The Claude Agent SDK, the Codex SDK and the Cursor SDK let a program, not a person, run a coding agent: start it, restrict its tools, read a typed result and resume it later. Use an SDK when your code must branch on what the agent found; for a one-shot CI step, claude -p or codex exec is enough.

CI goes red at 02:10. By the time anyone looks, three more pushes have queued behind it, and the first half hour of the morning goes on reading 4,000 lines of log to learn that a test was flaky again. You already tried a bot: a workflow step that pipes the log into claude -p and posts the answer as a comment. It helps, but it cannot rerun the flaky test and wait, it cannot open a second, write-enabled session only when the verdict is “regression”, and nobody knows what it cost last month. That is the point where you stop triggering an agent and start driving one from code.

This page is for developers who build that automation and for tech leads who decide which SDK a team standardises on. It compares the three SDKs on the five things that decide the choice — sessions, tools, permissions, streaming and cost — and writes the same job in each.

What you’ll walk away with from the agent SDK comparison

Section titled “What you’ll walk away with from the agent SDK comparison”
  • A decision table for “trigger it or script it” that you can apply to any automation idea.
  • A comparison of the three SDKs on sessions, tools, permissions, streaming, structured output and cost, checked against the published packages on 2026-09-26.
  • A CI-failure triage runner written three times — Claude Agent SDK, Codex SDK, Cursor SDK — sharing one prompt and one JSON Schema, plus the GitHub Actions workflow that calls it.
  • The gates that make its output trustworthy without a person reading every run, and the failure modes that bite first.

Should you script the agent or trigger it?

Section titled “Should you script the agent or trigger it?”

Most “agent in CI” jobs never need an SDK. A trigger — a GitHub Action, a scheduled automation or a single headless CLI call — takes a prompt and produces a comment, a file or a pull request. An SDK earns its extra code only when the orchestration logic lives in your program.

Your job needs…Trigger is enoughReach for the SDK
One prompt in, one result outclaude -p --output-format json, codex exec --json -o out.json, Cursor’s print mode, anthropics/claude-code-action@v1, openai/codex-action@v1, Cursor Automations—
A typed result a script consumesclaude -p --json-schema, codex exec --output-schema FILEWhen the next step depends on the value
Different next steps per result (rerun, label, open a fix session)Only with brittle shell glueYes: branch in code, resume the same session
Tools that call your own systems (ticketing, feature flags, internal APIs)An MCP server you deploy and maintainIn-process tools: Claude createSdkMcpServer, Cursor customTools
A per-call permission decision (“allow git show, deny git push, ask the service for the rest”)Static allow and deny listsClaude’s canUseTool callback and hooks
Live progress in your own UI, bot or logTail a JSONL fileTyped event streams in all three SDKs
Spend accounting per run, per team or per customerScrape the CLI’s JSONCost or usage fields on the result
Many agents coordinated by your code—Yes; see multi-agent orchestration patterns

The rule of thumb: if you can describe the job as a single prompt with a fixed output, keep it a trigger. If you catch yourself writing if statements around the agent’s answer, move to an SDK. The CLI-level options are covered per tool in Claude Code automation and Codex non-interactive mode.

How do the Claude Agent SDK, Codex SDK and Cursor SDK differ?

Section titled “How do the Claude Agent SDK, Codex SDK and Cursor SDK differ?”

All three put the vendor’s own coding agent behind a library: the same tools, the same project instruction files and the same model catalogue you use interactively. They differ in what they let your code decide. Everything below was read from the published TypeScript packages on 2026-09-26.

Claude Agent SDKCodex SDKCursor SDK
Packages checkednpm @anthropic-ai/claude-agent-sdk 0.3.283, PyPI claude-agent-sdk 0.2.160npm @openai/codex-sdk 0.157.1, PyPI openai-codex 0.157.1npm @cursor/sdk 1.0.32, PyPI cursor-sdk 1.0.32
What runs underneathThe Claude Code binary, installed as a platform packageThe codex CLI from @openai/codex, spoken to over JSONL on stdin and stdoutA local agent, or a cloud agent in an isolated VM when you pass cloud
RuntimeNode 18+Node 18+Node 22.13+
Entry pointquery({ prompt, options }), an async generator of messagesnew Codex().startThread(), then thread.run() or thread.runStreamed()Agent.create(), then agent.send() returning a Run
Sessionssession_id on every message; resume, continue, forkSessionThread id; codex.resumeThread(id); threads persist in ~/.codex/sessionsagentId; Agent.resume(id); local agents persist in a SQLite or JSONL store, cloud agent IDs start with bc-
Restricting toolstools sets the toolset; allowedTools only auto-approves; disallowedTools removesNo per-tool list; the sandbox mode bounds what commands can dotools and disallowedTools (local agents only)
Your own toolsIn-process MCP server via createSdkMcpServer() and tool(), plus mcpServersMCP servers from Codex configurationcustomTools callbacks (local only), plus mcpServers
PermissionspermissionMode (default, acceptEdits, plan, dontAsk, auto, bypassPermissions), canUseTool callback, hooks such as PreToolUsesandboxMode (read-only, workspace-write, danger-full-access) plus approvalPolicy (never, on-request, on-failure, untrusted)local.sandboxOptions, local.autoReview (Cursor’s classifier-backed Auto-review); a cloud agent is isolated by its VM
StreamingEvery message as it happens; includePartialMessages adds token deltasrunStreamed() yields thread.started, item.*, turn.completed, turn.failed, errorrun.stream(), plus onStep and onDelta callbacks on send()
Structured outputoutputFormat: { type: 'json_schema', schema } → structured_output on the resultoutputSchema per turn → JSON in finalResponseNo schema option in 1.0.32; capture the report through a custom tool
Limits and costmaxTurns, maxBudgetUsd; total_cost_usd on the result, documented as an estimateToken usage on turn.completed; no dollar field and no spend cap; cancel with an AbortSignalagent.getUsage() returns tokens and cost in cents, which can lag the run; cancel with run.cancel()
ModelClaude Code’s default unless you set modelCodex’s default unless you set model; modelReasoningEffortmodel is required for local agents; list IDs with Cursor.models.list()

Three consequences are easy to miss in the table:

  • Only Claude lets code decide each tool call. canUseTool and PreToolUse hooks see every call before it runs. The Codex SDK gives you a sandbox boundary instead of a callback (no per-call hook in @openai/codex-sdk 0.157.1), and the Cursor SDK gives you tool lists and Auto-review.
  • The Agent SDK does not start in auto mode. From Claude Code v2.1.283 (the latest channel), interactive sessions start in auto mode, but the Agent SDK and claude -p still start in the Manual (default) permission mode. An unattended job that relies on prompts being answered stalls or is denied, so pick dontAsk with an explicit allowlist.
  • Cursor’s SDK is the only one that can hand the work to a cloud VM. Passing cloud: { repos: [...] } runs the agent on Cursor’s infrastructure and can open the pull request itself (autoCreatePR). Tool lists, custom tools and a custom system prompt are local-only in 1.0.32, and combining tools with cloud throws a ConfigurationError.

Model choice is not a reason to pick an SDK: each defaults to its tool’s default model. Start there, tune effort before you switch models, and check the current names and prices on the models hub.

The same job three times: triage a failing CI run

Section titled “The same job three times: triage a failing CI run”

The job: when the CI workflow fails, download the failed steps’ log, find the first real failure, reproduce it, rerun it to rule out flakiness, and write a triage.json with a category (regression, flaky, infra, test-bug), the failing tests, a suspect commit, evidence and a next step. The agent must not edit anything. A later step decides what to do with the verdict.

The three runners share one prompt and one schema, so a switch between vendors changes only the runner file.

ci/triage-spec.ts
export const triageSchema = {
type: 'object',
properties: {
category: { type: 'string', enum: ['regression', 'flaky', 'infra', 'test-bug'] },
failing_tests: { type: 'array', items: { type: 'string' } },
suspect_commit: { type: 'string' },
evidence: { type: 'string' },
next_step: { type: 'string' },
},
required: ['category', 'failing_tests', 'suspect_commit', 'evidence', 'next_step'],
additionalProperties: false,
};
export const triagePrompt = `CI failed on this commit. The job log is in ci-failure.log.
Triage it; do not fix anything and do not edit any file.
1. Find the first real failure in the log (skip cascading errors).
2. Reproduce only the failing test(s) with: npx vitest run <file>.
3. Run it twice more. Passes on a rerun => "flaky".
4. Network, runner, secret or cache errors => "infra".
5. Otherwise read git log -5 and the diff of the suspect commit;
decide "regression" (code broke) or "test-bug" (test is wrong).
Cite log lines and command output as evidence. Use "unknown" for
suspect_commit when you cannot name one.`;

The Cursor SDK has no output-schema option, so the runner gives the agent a submit_triage custom tool whose input schema is the triage schema. The tool call’s arguments are the report. A local agent needs an explicit model ID; list the IDs your key can use with Cursor.models.list() and store one in CURSOR_MODEL_ID.

ci/triage-cursor.ts
import { writeFile } from 'node:fs/promises';
import { Agent, type SDKJsonValue } from '@cursor/sdk';
import { triagePrompt, triageSchema } from './triage-spec.js';
let report: Record<string, SDKJsonValue> | undefined;
const agent = await Agent.create({
apiKey: process.env.CURSOR_API_KEY,
model: { id: process.env.CURSOR_MODEL_ID! }, // required for local agents
disallowedTools: ['edit', 'delete', 'applyAgentDiff'],
local: {
cwd: process.env.GITHUB_WORKSPACE,
settingSources: ['project'],
customTools: {
submit_triage: {
description: 'Submit the final triage report. Call it exactly once, at the end.',
inputSchema: triageSchema,
execute: (args) => {
report = args;
return 'recorded';
},
},
},
},
});
try {
const run = await agent.send(`${triagePrompt}\nFinish by calling submit_triage.`);
const timer = setTimeout(() => void run.cancel(), 10 * 60_000);
for await (const message of run.stream()) {
if (message.type === 'tool_call' && message.status === 'completed') console.log('tool', message.name);
}
const result = await run.wait();
clearTimeout(timer);
if (result.status !== 'finished' || !report) throw new Error(`triage ${result.status}, no report`);
await writeFile('triage.json', JSON.stringify(report, null, 2));
const usage = await agent.getUsage(); // cost can lag the run
console.log(`agent ${agent.agentId}`, usage.usage, usage.cost);
} finally {
agent.close();
}

disallowedTools removes the edit tools, but shell stays, and a shell can still write files. The git diff --exit-code step in the workflow below is what enforces “no edits”.

All three files type-check with TypeScript 5 (strict, NodeNext) against the package versions in the comparison table. Run them with npx tsx ci/triage-<tool>.ts.

The workflow is the same for all three SDKs apart from the last run line and its secret.

.github/workflows/ci-triage.yml
name: ci-triage
on:
workflow_run:
workflows: [CI]
types: [completed]
permissions:
contents: read
actions: read
jobs:
triage:
# Only failed runs, and never code from a fork: this job holds API keys.
if: >-
github.event.workflow_run.conclusion == 'failure' &&
github.event.workflow_run.head_repository.full_name == github.repository
runs-on: ubuntu-latest
timeout-minutes: 15
steps:
- uses: actions/checkout@v4
with:
ref: ${{ github.event.workflow_run.head_sha }}
fetch-depth: 20
- uses: actions/setup-node@v4
with:
node-version: 22 # the Cursor SDK needs 22.13+
- run: npm ci
- run: gh run view ${{ github.event.workflow_run.id }} --log-failed > ci-failure.log
env:
GH_TOKEN: ${{ github.token }}
- run: npx tsx ci/triage-claude.ts # or triage-codex.ts / triage-cursor.ts
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
- run: git diff --exit-code # the triage must leave tracked files untouched
- uses: actions/upload-artifact@v4
with:
name: triage
path: triage.json

For the Codex runner, pass CODEX_API_KEY; for the Cursor runner, pass CURSOR_API_KEY and CURSOR_MODEL_ID. Check your vendor’s terms for running unattended on an API key versus a subscription login before you pick the credential.

Branch on the verdict: where the SDK pays for itself

Section titled “Branch on the verdict: where the SDK pays for itself”

Up to here, claude -p --json-schema or codex exec --output-schema would have done the same job. The SDK becomes worth its code when the next step depends on the answer and reuses the agent’s context:

  1. Read triage.json and validate it against triageSchema with a JSON Schema validator. A report that fails validation fails the job; nobody acts on it.

  2. On flaky, rerun the failed jobs (gh run rerun <run-id> --failed; this step needs actions: write on its job) and add the test to your quarantine list. No second agent session.

  3. On infra, post the evidence to the on-call channel and stop.

  4. On regression or test-bug, resume the same session with write access in a new branch, so the agent keeps what it already read:

    • Claude Agent SDK: query({ prompt, options: { resume: sessionId, tools: { type: 'preset', preset: 'claude_code' }, permissionMode: 'acceptEdits' } }).
    • Codex SDK: codex.resumeThread(threadId, { sandboxMode: 'workspace-write' }), then thread.run(...).
    • Cursor SDK: agent.send(...) on the same agent, or Agent.resume(agentId) from a later job; the tool restrictions are not persisted, so pass them again.
  5. The fix session opens a pull request that carries the triage report as evidence, and it goes through the same gates as any other agent change: see the evidence bundle an agent’s pull request must carry.

The GitHub Actions runner is ephemeral, so a session saved on it is gone when the job ends. To resume in a later job, persist the session store (~/.claude/projects, ~/.codex/sessions or the Cursor local store) as an artifact, or keep the whole flow in one job.

Copy-paste prompts for building an SDK runner

Section titled “Copy-paste prompts for building an SDK runner”

Run the audit prompt with a different model or vendor than the runner, so the author never grades itself; model-graded checks covers how to calibrate such a judge.

How do you prove the triage is right without reading every run?

Section titled “How do you prove the triage is right without reading every run?”

Treat the runner like any other production code path whose output other automation acts on. The gates, in the order they fire:

  1. Schema. The SDK validates the output shape (Claude, Codex) or the custom tool’s input schema constrains it (Cursor). Your step validates it again before acting, because a run that hits a cap can end with no report at all.
  2. No side effects. git diff --exit-code fails the job if the agent changed a tracked file. Tool restrictions reduce the chance; the diff proves it.
  3. Caps. maxTurns and maxBudgetUsd on Claude, a timeout on Codex and Cursor, and timeout-minutes on the job. A capped run is a failed run, never a partial verdict.
  4. A golden set. Collect past failed runs with a label a person already agreed on, including flaky, infra and regression cases. Replay the runner on them after every SDK, model or prompt change, and track agreement. Decide in advance what agreement you require before the runner may rerun jobs or label issues on its own; until then it only comments.
  5. An independent audit. The audit prompt above, run by another model, rejects verdicts whose evidence does not appear in the log.
  6. Telemetry. Log the session or thread ID, turns, tokens and cost for every run, and join them to the CI run ID. The agent observability page shows the dashboard: run success, cost per accepted verdict and override rate.

Who signs off: the on-call developer owns every action the runner proposes until the golden-set agreement meets the bar the tech lead set. After that, the tech lead signs off on widening its authority one action at a time, starting with reruns of flaky tests.

What breaks when you drive agents from an SDK?

Section titled “What breaks when you drive agents from an SDK?”
SymptomCauseRecovery
The Claude run ends with error_max_budget_usd or error_max_turns and no reportThe cap is too low for the repository’s log size, or the agent is looping on a noisy logTrim the log to the failed step before the run (--log-failed already helps), then raise one cap at a time and record the new cost in the golden-set run
The Codex run hangs until the timeoutA command waited for input or the network with networkAccessEnabled: falseRead the last command_execution item in the streamed log; add the non-interactive flag to the reproduce command in the prompt
Every Codex verdict is “infra”read-only sandbox, and the test runner needs to write a cacheSwitch to workspace-write and keep the diff gate
Cursor’s Agent.create throws ConfigurationErrortools, disallowedTools or systemPrompt combined with cloud (and customTools is local-only too)Run the triage as a local agent, or drop those options for the cloud agent and restrict it through the prompt and the repository access you grant
Cursor reports finished but report is undefinedThe agent answered in text and never called submit_triageTreat it as a failure (the runner already throws); repeat the instruction to call the tool at the end of the prompt, and send one follow-up on the same agent asking for the call
Cost on the Cursor usage call is missingBilling data lags the runRead agent.getUsage() again in a later step, or join the usage export by agent ID
A resumed session in a later job is “not found”The session lived on the previous ephemeral runnerUpload the session store as an artifact, or keep triage and fix in one job
The SDK and a globally installed CLI behave differentlyThe SDK drives its own pinned binary, and your global CLI is another versionPin the SDK version in package.json and upgrade on purpose; the Claude Agent SDK accepts pathToClaudeCodeExecutable when you must use a specific binary

Pick by where the team already works and by the control the job needs, not by the model: each SDK drives its own vendor’s agent and default model.

  • Choose the Claude Agent SDK when the job needs per-call permission decisions, hooks, in-process tools and a hard dollar cap in the same place. It has the widest control surface of the three.
  • Choose the Codex SDK when the team already runs Codex and a sandbox boundary is the control you trust. It is the thinnest wrapper, and its per-turn schema makes multi-step pipelines on one thread straightforward.
  • Choose the Cursor SDK when you want the agent to run on Cursor’s cloud VMs and open the pull request itself, or when the team’s rules and skills already live in Cursor.

For a lead, the standardisation decision is mostly about the shared parts: one prompt and schema file per job, one set of gates, one golden set and one cost dashboard. With those in place, switching the runner file is a small change, as the three tabs above show.

Frequently asked questions

When should I use an agent SDK instead of claude -p or codex exec?

Use the CLI when the job is one prompt in and one result out. Use an SDK when your code must branch on the agent's result, resume the same session, register in-process tools, decide permissions per tool call, or stream progress into your own UI.

Which agent SDKs return structured output against a JSON Schema?

The Claude Agent SDK (outputFormat with type json_schema) and the Codex SDK (outputSchema per turn) do. Cursor SDK 1.0.32 has no schema option; a local custom tool that receives the report as its arguments gives you the same result.

How do I cap what an SDK-driven agent spends?

The Claude Agent SDK has maxBudgetUsd and maxTurns. The Codex SDK has no spend cap, so pass an AbortSignal with a timeout. The Cursor SDK has no cap either; cancel the run on a timer and read agent.getUsage() afterwards.