Skip to content

Kiro: spec files, property-based tests and Kiro Crew

Kiro is AWS’s agentic development platform: an IDE, CLI, Web, Mobile (preview) and the open-source Kiro Crew, all on one agent harness. Its distinctive pieces are spec files checked by a requirements-analysis pass, property-based tests that check code against the spec, and Crew’s persistent, scheduled workspace. Teams on Claude Code, Codex or Cursor can borrow all three without switching.

Someone on your team demos Kiro: a one-line idea becomes requirements.md, design.md and tasks.md, the agent flags two contradictory requirements before writing any code, and the finished feature arrives with property tests traced to requirement IDs. Your team runs Claude Code and Cursor, and you have to decide whether that justifies a second tool, or whether the method can move without the product.

  • A dated map of Kiro’s five surfaces and the features that set it apart.
  • How Kiro’s spec flow and its property-based “correctness” check fit together, and which parts are verified.
  • What Kiro Crew is, how to install it, and what it adds beyond a chat session.
  • A borrow table: each Kiro idea mapped to its equivalent in Claude Code, Codex and Cursor, with verified install commands.
  • Three copy-paste prompts (requirements analysis, properties from requirements, bugfix spec) and a verification checklist for the borrowed loop.

Kiro started as an agentic IDE in 2025 and is now a family of surfaces. Its README states that “one unified agent harness powers every Kiro surface, so your specs, steering, permissions, hooks, MCP servers, and custom agents can follow your project across workflows” (kirodotdev/Kiro, read 2026-09-26).

SurfaceWhat the README says it is forStatus
IDELocal development with editor integration, chat, specs and hooksAvailable
CLI (kiro-cli)Terminal-native work, headless automation and CI/CDAvailable
WebDelegating multi-repository tasks in isolated cloud sandboxesAvailable
MobileMonitoring tasks, reviewing pull requests, working with agents on the goPreview
CrewA persistent open-source workspace with memory, scheduling and multi-channel accessOpen source, Apache-2.0

Around the spec workflow, the README lists custom agents, Steering (project conventions), Agent Skills, Hooks, Powers (tools and knowledge loaded on demand), MCP, permissions, a .kiroignore file for sensitive paths, checkpoints with rewind, cloud sessions (preview), and enterprise governance. The product source is closed: kirodotdev/Kiro is the public issue tracker (4,330 stars on 2026-09-26), not the code.

Kiro sells credit-based plans with a free tier (secondary: aitoolpick, September 2026). Check the vendor’s pricing page before you budget; this site keeps tool plan prices on the pricing comparison.

How does Kiro’s spec workflow turn intent into tasks?

Section titled “How does Kiro’s spec workflow turn intent into tasks?”

A Kiro spec is three Markdown files per feature, generated in order, with a human approval gate between each. According to secondary sources (kiro.dev search extracts, September 2026), they live under .kiro/specs/<feature>/:

FileHoldsWho approves it
requirements.mdUser stories with acceptance criteria in EARS form: “WHEN [condition] THE SYSTEM SHALL [behaviour]”Product owner or tech lead
design.mdArchitecture, data model, interfaces and error handling for those requirementsTech lead
tasks.mdNumbered implementation tasks, each traced to requirement IDsThe developer who runs them

The same sources describe three variants: Feature Specs (requirements-first or design-first), Bugfix Specs (root cause, fix design and the behaviour that must be preserved), and a Quick Plan that writes all three files without gates. Kiro CLI shares .kiro/specs/ with the IDE.

Two parts of this flow are verified in Kiro’s own README. Requirements Analysis runs over the spec “to find contradictions, ambiguities, and gaps before coding begins”. Correctness with property-based testing comes next.

An EARS requirement is useful because a machine can check it. Compare the two forms:

.kiro/specs/cart-discounts/requirements.md (excerpt)
### CART-3: Discounted total stays in range
WHEN a discount code is applied to a cart
THE SYSTEM SHALL return a total that is a whole number of cents,
not less than 0 and not greater than the undiscounted subtotal.
Example: subtotal 1999, code SAVE100PCT (100%) → total 0.

“Discounts should work correctly” gives an agent nothing to test. CART-3 names an input class, an observable output and two bounds, which is exactly what a property test needs.

How does Kiro check the code against the spec?

Section titled “How does Kiro check the code against the spec?”

Kiro’s README describes it as “correctness with property-based testing”: the agent turns requirements into executable properties and exercises them “across generated inputs that example-based tests may miss” (kirodotdev/Kiro, read 2026-09-26). Which property library Kiro uses for each language is not stated in the README, and kiro.dev was unreachable, so this page does not name one.

The idea is portable, and it matters more with agents than without them. An agent that writes both the code and the example tests can pick examples its code passes. A property test lets the library pick the inputs, usually 100 or more per run, and shrink any failure to a minimal counterexample. For CART-3 in TypeScript with fast-check 4.10.2 and @fast-check/vitest. Install both first; @fast-check/vitest needs Vitest 4.1 or later (npm peer range, checked 2026-09-26):

Terminal window
npm install -D fast-check @fast-check/vitest

Then write the property:

tests/pricing/apply-discount.property.test.ts
import { expect } from 'vitest';
import { fc, test } from '@fast-check/vitest';
import { applyDiscount } from '../../src/pricing/apply-discount';
const lineCents = fc.array(fc.integer({ min: 1, max: 1_000_000 }), { minLength: 1, maxLength: 50 });
const discount = fc.oneof(
fc.record({ kind: fc.constant('percent' as const), value: fc.integer({ min: 0, max: 100 }) }),
fc.record({ kind: fc.constant('fixed' as const), value: fc.integer({ min: 0, max: 5_000_000 }) }),
);
// CART-3: WHEN a discount code is applied THE SYSTEM SHALL return a whole-cent total in [0, subtotal].
test.prop([lineCents, discount])('CART-3 discounted total stays in range', (lines, code) => {
const subtotal = lines.reduce((sum, cents) => sum + cents, 0);
const total = applyDiscount(lines, code);
expect(Number.isInteger(total)).toBe(true);
expect(total).toBeGreaterThanOrEqual(0);
expect(total).toBeLessThanOrEqual(subtotal);
});

The requirement ID in the test name is the traceability link: a reviewer reads requirements.md and the list of passing property names, not the diff. The full method, including Hypothesis and proptest versions and how to pin shrunk failures as regression tests, is on property-based testing for agent-written code.

Kiro Crew is “a persistent workspace for development work that self-improves and continues beyond one session” (kirodotdev/KiroCrew README, read 2026-09-26). It is open source under Apache-2.0, created on 2026-07-16, and had about 4,200 stars (GitHub, 2026-09-26). It runs on macOS, Linux or Windows, in a local container, or on a remote machine you control (README, 2026-09-26).

A Gateway process keeps sessions, memory, schedules, approvals and task checkpoints on the host, and you reach the same work from the desktop app, a web dashboard, the CLI, Slack or Discord. The README lists what it adds over a chat session:

  • Long-running tasks. You give Crew a task spec; it plans, executes, validates each step, retries failures and resumes from checkpoints. The README’s example: “Implement this migration plan and stop if the tests fail.”
  • Unattended autonomy. Scheduled agent jobs, deterministic scripts that need no model call, heartbeats that watch a system until something needs attention, and reactions to messaging events and authenticated webhooks.
  • Self-learning. Corrections and task failures become durable lessons, optionally scoped to one repository with repo_scope.
  • Self-evolving skills. Repeated patterns become skills you can inspect, refine or delete.
  • Defense in depth. Tool approvals, OS sandboxing where available, sensitive-path and credential guards, deny rules, audit events and governance profiles. The dashboard binds to loopback by default.

Crew drives its agent over the Agent Client Protocol (ACP). A fresh configuration uses kiro-cli, so the default path needs a Kiro login. The agent.acp_backend setting selects a different harness, and the user configuration reference lists kas (Kiro Agent, through a kiro-cli relay), claude (Claude Code, through the public claude-agent-acp adapter), codex (Codex), opencode, pi and goose besides the default (kirodotdev/KiroCrew src/kiro_crew/docs/configuration.md, read 2026-09-26). The changelog labels these backends Preview: Claude Code and Codex arrived in 0.6.0 and OpenCode in 0.7.0, and you enable them under Settings → Developer (Developer Mode, off by default). The latest release is 0.7.1 (2026-09-24).

That changes the borrow question. A team on Claude Code or Codex can try Crew’s persistence and scheduling without switching agents, with two caveats from the same changelog: a tool pre-approved in Claude Code’s own settings “never reaches Crew’s approval path”, so Crew’s deny rules and audit log do not see that call; and Codex refuses to start while Crew’s sandbox is off.

To try it on a workstation (terminal):

Terminal window
# Install kiro-cli from Kiro's CLI setup docs first, then sign in:
kiro-cli login
# Install the signed Stable Crew build, then check the setup
curl -fsSL https://download.crew.kiro.dev/cli.sh | sh
kirocrew doctor
kirocrew gateway # dashboard on http://localhost:5476
# Crew sends one anonymous heartbeat per day by default; to turn it off:
kirocrew telemetry disable

For an always-on server the README ships a Docker image, ghcr.io/kirodotdev/kirocrew:stable, published on 127.0.0.1:5476. Keep it on loopback and reach it through an SSH tunnel or your VPN rather than a public port.

What can a team on Claude Code, Codex or Cursor borrow from Kiro?

Section titled “What can a team on Claude Code, Codex or Cursor borrow from Kiro?”

Each Kiro idea has a portable equivalent. The method moves; the product-enforced gates do not, so you enforce them in CI and review instead.

Kiro ideaBorrow it asClaude CodeCodexCursor
Three-file spec with gatescc-sdd 3.1.0 skills, or one spec.md per capabilitynpx cc-sdd@latest (default target)npx cc-sdd@latest --codex-skillsnpx cc-sdd@latest --cursor-skills (beta)
Requirements AnalysisThe first prompt below, run before designPlan mode (/plan)Plan mode (/plan)Plan Mode
Property-based correctnessfast-check or Hypothesis in the test suite, plus Trail of Bits’ property-based-testing skillPluginPluginnpx skills add
Bugfix Spec, preserved behaviourCharacterization tests before the fixCharacterization testssamesame
SteeringProject instruction filesCLAUDE.mdAGENTS.mdRules
Crew schedules and heartbeatsThe tool’s own scheduler, or Crew itself with agent.acp_backend set to claude or codex (preview)RoutinesAutomationsCloud agents and automations
Crew lessonsMemory and instruction files you reviewMemory patternssamesame

Set up the borrowed spec-and-properties loop

Section titled “Set up the borrowed spec-and-properties loop”
  1. Install the spec skills. cc-sdd (“Kiro-style” specs for other agents, npm 3.1.0 on 2026-09-26) writes the same requirements.md (EARS), design.md and tasks.md shape. Preview what it will write with --dry-run first. cc-sdd also writes CLAUDE.md (Claude Code) or AGENTS.md (Codex, Cursor) and .kiro/settings/; if you already have an instruction file, run with --overwrite skip or --backup and merge by hand.

    Terminal window
    npx cc-sdd@latest --dry-run
    npx cc-sdd@latest # Claude Code skills are the default target

    Then, in Claude Code: /kiro-discovery CSV export of the audit log for team admins.

  2. Run discovery, then one phase at a time. The cc-sdd flow is kiro-discovery → kiro-spec-init → kiro-spec-requirements → kiro-spec-design → kiro-spec-tasks → kiro-impl. Stop after kiro-spec-requirements and run the requirements-analysis prompt below before you approve anything.

  3. Install a property-based-testing skill. Trail of Bits’ property-based-testing plugin (1.2.2) writes, reviews and triages property tests for Hypothesis, fast-check, proptest and others. Only the skill’s description stays in context until it fires.

    Terminal window
    claude plugin marketplace add trailofbits/skills
    claude plugin install property-based-testing@trailofbits
  4. Write the properties before the implementation. Use the second prompt below. Commit the property tests with the approved requirements, so the implementation task starts red.

  5. Implement task by task. kiro-impl gives each task a fresh implementer running red-green TDD and an independent reviewer when the host has subagents; otherwise it runs inline. The property tests are the stop condition.

  6. Gate the pull request. CI runs the property tests at their default run count on every pull request and a deeper sweep nightly. To make the count settable, register a setup file in the Vitest config and call fc.configureGlobal there:

    vitest.config.ts
    import { defineConfig } from 'vitest/config';
    export default defineConfig({
    test: { setupFiles: ['tests/setup/fast-check.ts'] },
    });
    tests/setup/fast-check.ts
    import fc from 'fast-check';
    fc.configureGlobal({ numRuns: Number(process.env.FC_NUM_RUNS ?? 100) });

    Set FC_NUM_RUNS=10000 in the nightly job. The pull request lists each requirement ID with the test that proves it.

How do you prove the borrowed loop works without reading every line?

Section titled “How do you prove the borrowed loop works without reading every line?”

Kiro enforces its gates in the product. When you borrow the method, the evidence has to be explicit, and each gate needs an owner:

GateEvidenceWho signs off
RequirementsThe requirements-analysis table has no open rows; every requirement is in EARS form with an exampleProduct owner or tech lead approves requirements.md
Oracle strengthEvery requirement maps to a named property or example test; the test file changed in the same pull request as requirements.md, not after the codeTech lead, from the pull request’s traceability list
CorrectnessProperty tests pass in CI at the default run count; the nightly deep run is green; every shrunk failure is pinned as a regression caseCI
DriftA change to behaviour without a spec delta fails review, as in spec-driven developmentReviewer

Two further checks keep the oracle honest. Mutation testing on the module under test shows whether the properties can fail at all; Trail of Bits ships a mutation-testing skill in the same marketplace. And the agent that implements a task must not edit the property tests in the same session: see protecting the oracle for the hook and review rules that enforce this, and oracle strength for how to measure it.

Should your team switch to Kiro or borrow from it?

Section titled “Should your team switch to Kiro or borrow from it?”
SituationRecommendation
Your team is on Claude Code, Codex or Cursor and likes the spec disciplineBorrow: cc-sdd plus property tests in CI. No second tool, no second instruction format
You want the gates enforced by the product, not by convention, and your stack is AWS-centricPilot Kiro on one team, with the bake-off protocol on real closed issues
You want an always-on agent on your own hardware, with memory and schedulesTry Kiro Crew on a spare machine. The default harness needs a Kiro login; the Claude Code and Codex harnesses are preview. Compare it with your tool’s own scheduler first
You need one harness across many agentsStay portable: Agent Skills and one instruction file work across tools; see lock-in and portability

What breaks when you borrow Kiro’s spec and property workflow?

Section titled “What breaks when you borrow Kiro’s spec and property workflow?”

Properties that restate the implementation. The agent writes expect(applyDiscount(l, d)).toBe(subtotal - discountFor(l, d)), which passes for any bug the helper shares. Recovery: ban calls into the module under test from the expected side; accept only bounds, round-trips, idempotence or a separate simple reference model. Rerun the second prompt with that rule quoted.

Generators narrowed until green. A failing property is “fixed” by shrinking the input range to values the code handles. Recovery: review generator diffs as carefully as code diffs, and fail the pull request when a generator’s domain shrinks without a matching change to requirements.md.

EARS in form, not in substance. “WHEN the user exports THE SYSTEM SHALL export correctly” has the shape and none of the content. Recovery: the requirements-analysis prompt flags missing outcomes, units and bounds; do not approve design until its table is empty.

Spec and code drift apart after the first release. The three files describe the feature as planned, and later pull requests change behaviour without touching them. Recovery: require a spec delta in every behaviour-changing pull request and audit drift on a schedule, as described in spec-driven development.

cc-sdd installed in legacy mode. --claude, --cursor and the other command modes are deprecated and --codex is blocked (cc-sdd 3.1.0). Recovery: reinstall with the --*-skills flag, after a --dry-run into an empty directory if you customised the old files.

Crew lessons and skills grow without review. Self-evolving skills and lessons change later behaviour. Recovery: review them in the dashboard on a schedule, delete anything you would not accept in a pull request, and scope lessons to one repository with repo_scope.

Crew exposed beyond loopback. A port published on 0.0.0.0 gives anyone on the network an agent with your credentials. Recovery: bind to 127.0.0.1, use an SSH tunnel or VPN, and read Crew’s security model before widening access. See permissions and sandboxing.

Claude Code under Crew skips Crew’s approval path. With agent.acp_backend set to claude, any tool you pre-approved in Claude Code’s own settings (user, project, local or managed) runs without reaching Crew’s deny rules or audit log (Crew changelog 0.6.0). Recovery: keep Claude Code’s allow list minimal on the Crew host, put deny rules in Claude Code’s own settings as well as Crew’s, and treat Crew’s audit log as incomplete for that backend.

Vendor details change under you. Kiro’s spec variants and CLI commands above come from secondary sources, and Kiro ships often. Recovery: re-check kiro.dev before you standardise on a specific command, and pin the version you tested.

Frequently asked questions

What is Kiro?

Kiro is AWS's agentic development platform. As of 26 September 2026 one agent harness powers its IDE, CLI, Web and Mobile (preview) surfaces and the open-source Kiro Crew workspace, so specs, steering, permissions, hooks, MCP servers and custom agents follow a project across all of them.

How does Kiro use property-based testing?

Kiro turns spec requirements into executable properties and exercises them across generated inputs, to check that the code matches the spec. You can reproduce the same loop on Claude Code, Codex or Cursor with fast-check or Hypothesis and a property-based-testing skill.

What is Kiro Crew?

Kiro Crew is an open-source (Apache-2.0) workspace that runs on your own hardware and keeps working between sessions: persistent memory, checkpointed long-running tasks, schedules, heartbeats, and access from a desktop app, web dashboard, CLI, Slack and Discord. Its default agent runs through kiro-cli and a Kiro login; Claude Code, Codex and OpenCode are selectable harnesses in preview.

Should a team on Claude Code, Codex or Cursor switch to Kiro for specs?

Usually not. The spec shape (requirements, design, tasks), the requirements-analysis pass and the property tests are all portable: cc-sdd installs Kiro-style spec skills for Claude Code and Codex (Cursor in beta), and the property tests are plain test code. Switch only if you want those gates enforced by the product rather than by your own conventions.