Kiro: spec files, property-based tests and Kiro Crew
Kiro is AWS’s agentic development platform: an IDE, CLI, Web, Mobile (preview) and the open-source Kiro Crew, all on one agent harness. Its distinctive pieces are spec files checked by a requirements-analysis pass, property-based tests that check code against the spec, and Crew’s persistent, scheduled workspace. Teams on Claude Code, Codex or Cursor can borrow all three without switching.
Someone on your team demos Kiro: a one-line idea becomes requirements.md, design.md and tasks.md, the agent flags two contradictory requirements before writing any code, and the finished feature arrives with property tests traced to requirement IDs. Your team runs Claude Code and Cursor, and you have to decide whether that justifies a second tool, or whether the method can move without the product.
What this Kiro reference gives you
Section titled “What this Kiro reference gives you”- A dated map of Kiro’s five surfaces and the features that set it apart.
- How Kiro’s spec flow and its property-based “correctness” check fit together, and which parts are verified.
- What Kiro Crew is, how to install it, and what it adds beyond a chat session.
- A borrow table: each Kiro idea mapped to its equivalent in Claude Code, Codex and Cursor, with verified install commands.
- Three copy-paste prompts (requirements analysis, properties from requirements, bugfix spec) and a verification checklist for the borrowed loop.
What is Kiro in September 2026?
Section titled “What is Kiro in September 2026?”Kiro started as an agentic IDE in 2025 and is now a family of surfaces. Its README states that “one unified agent harness powers every Kiro surface, so your specs, steering, permissions, hooks, MCP servers, and custom agents can follow your project across workflows” (kirodotdev/Kiro, read 2026-09-26).
| Surface | What the README says it is for | Status |
|---|---|---|
| IDE | Local development with editor integration, chat, specs and hooks | Available |
CLI (kiro-cli) | Terminal-native work, headless automation and CI/CD | Available |
| Web | Delegating multi-repository tasks in isolated cloud sandboxes | Available |
| Mobile | Monitoring tasks, reviewing pull requests, working with agents on the go | Preview |
| Crew | A persistent open-source workspace with memory, scheduling and multi-channel access | Open source, Apache-2.0 |
Around the spec workflow, the README lists custom agents, Steering (project conventions), Agent Skills, Hooks, Powers (tools and knowledge loaded on demand), MCP, permissions, a .kiroignore file for sensitive paths, checkpoints with rewind, cloud sessions (preview), and enterprise governance. The product source is closed: kirodotdev/Kiro is the public issue tracker (4,330 stars on 2026-09-26), not the code.
Kiro sells credit-based plans with a free tier (secondary: aitoolpick, September 2026). Check the vendor’s pricing page before you budget; this site keeps tool plan prices on the pricing comparison.
How does Kiro’s spec workflow turn intent into tasks?
Section titled “How does Kiro’s spec workflow turn intent into tasks?”A Kiro spec is three Markdown files per feature, generated in order, with a human approval gate between each. According to secondary sources (kiro.dev search extracts, September 2026), they live under .kiro/specs/<feature>/:
| File | Holds | Who approves it |
|---|---|---|
requirements.md | User stories with acceptance criteria in EARS form: “WHEN [condition] THE SYSTEM SHALL [behaviour]” | Product owner or tech lead |
design.md | Architecture, data model, interfaces and error handling for those requirements | Tech lead |
tasks.md | Numbered implementation tasks, each traced to requirement IDs | The developer who runs them |
The same sources describe three variants: Feature Specs (requirements-first or design-first), Bugfix Specs (root cause, fix design and the behaviour that must be preserved), and a Quick Plan that writes all three files without gates. Kiro CLI shares .kiro/specs/ with the IDE.
Two parts of this flow are verified in Kiro’s own README. Requirements Analysis runs over the spec “to find contradictions, ambiguities, and gaps before coding begins”. Correctness with property-based testing comes next.
An EARS requirement is useful because a machine can check it. Compare the two forms:
### CART-3: Discounted total stays in rangeWHEN a discount code is applied to a cartTHE SYSTEM SHALL return a total that is a whole number of cents,not less than 0 and not greater than the undiscounted subtotal.
Example: subtotal 1999, code SAVE100PCT (100%) → total 0.“Discounts should work correctly” gives an agent nothing to test. CART-3 names an input class, an observable output and two bounds, which is exactly what a property test needs.
How does Kiro check the code against the spec?
Section titled “How does Kiro check the code against the spec?”Kiro’s README describes it as “correctness with property-based testing”: the agent turns requirements into executable properties and exercises them “across generated inputs that example-based tests may miss” (kirodotdev/Kiro, read 2026-09-26). Which property library Kiro uses for each language is not stated in the README, and kiro.dev was unreachable, so this page does not name one.
The idea is portable, and it matters more with agents than without them. An agent that writes both the code and the example tests can pick examples its code passes. A property test lets the library pick the inputs, usually 100 or more per run, and shrink any failure to a minimal counterexample. For CART-3 in TypeScript with fast-check 4.10.2 and @fast-check/vitest. Install both first; @fast-check/vitest needs Vitest 4.1 or later (npm peer range, checked 2026-09-26):
npm install -D fast-check @fast-check/vitestThen write the property:
import { expect } from 'vitest';import { fc, test } from '@fast-check/vitest';import { applyDiscount } from '../../src/pricing/apply-discount';
const lineCents = fc.array(fc.integer({ min: 1, max: 1_000_000 }), { minLength: 1, maxLength: 50 });const discount = fc.oneof( fc.record({ kind: fc.constant('percent' as const), value: fc.integer({ min: 0, max: 100 }) }), fc.record({ kind: fc.constant('fixed' as const), value: fc.integer({ min: 0, max: 5_000_000 }) }),);
// CART-3: WHEN a discount code is applied THE SYSTEM SHALL return a whole-cent total in [0, subtotal].test.prop([lineCents, discount])('CART-3 discounted total stays in range', (lines, code) => { const subtotal = lines.reduce((sum, cents) => sum + cents, 0); const total = applyDiscount(lines, code); expect(Number.isInteger(total)).toBe(true); expect(total).toBeGreaterThanOrEqual(0); expect(total).toBeLessThanOrEqual(subtotal);});The requirement ID in the test name is the traceability link: a reviewer reads requirements.md and the list of passing property names, not the diff. The full method, including Hypothesis and proptest versions and how to pin shrunk failures as regression tests, is on property-based testing for agent-written code.
What is Kiro Crew and what does it add?
Section titled “What is Kiro Crew and what does it add?”Kiro Crew is “a persistent workspace for development work that self-improves and continues beyond one session” (kirodotdev/KiroCrew README, read 2026-09-26). It is open source under Apache-2.0, created on 2026-07-16, and had about 4,200 stars (GitHub, 2026-09-26). It runs on macOS, Linux or Windows, in a local container, or on a remote machine you control (README, 2026-09-26).
A Gateway process keeps sessions, memory, schedules, approvals and task checkpoints on the host, and you reach the same work from the desktop app, a web dashboard, the CLI, Slack or Discord. The README lists what it adds over a chat session:
- Long-running tasks. You give Crew a task spec; it plans, executes, validates each step, retries failures and resumes from checkpoints. The README’s example: “Implement this migration plan and stop if the tests fail.”
- Unattended autonomy. Scheduled agent jobs, deterministic scripts that need no model call, heartbeats that watch a system until something needs attention, and reactions to messaging events and authenticated webhooks.
- Self-learning. Corrections and task failures become durable lessons, optionally scoped to one repository with
repo_scope. - Self-evolving skills. Repeated patterns become skills you can inspect, refine or delete.
- Defense in depth. Tool approvals, OS sandboxing where available, sensitive-path and credential guards, deny rules, audit events and governance profiles. The dashboard binds to loopback by default.
Crew drives its agent over the Agent Client Protocol (ACP). A fresh configuration uses kiro-cli, so the default path needs a Kiro login. The agent.acp_backend setting selects a different harness, and the user configuration reference lists kas (Kiro Agent, through a kiro-cli relay), claude (Claude Code, through the public claude-agent-acp adapter), codex (Codex), opencode, pi and goose besides the default (kirodotdev/KiroCrew src/kiro_crew/docs/configuration.md, read 2026-09-26). The changelog labels these backends Preview: Claude Code and Codex arrived in 0.6.0 and OpenCode in 0.7.0, and you enable them under Settings → Developer (Developer Mode, off by default). The latest release is 0.7.1 (2026-09-24).
That changes the borrow question. A team on Claude Code or Codex can try Crew’s persistence and scheduling without switching agents, with two caveats from the same changelog: a tool pre-approved in Claude Code’s own settings “never reaches Crew’s approval path”, so Crew’s deny rules and audit log do not see that call; and Codex refuses to start while Crew’s sandbox is off.
To try it on a workstation (terminal):
# Install kiro-cli from Kiro's CLI setup docs first, then sign in:kiro-cli login
# Install the signed Stable Crew build, then check the setupcurl -fsSL https://download.crew.kiro.dev/cli.sh | shkirocrew doctorkirocrew gateway # dashboard on http://localhost:5476
# Crew sends one anonymous heartbeat per day by default; to turn it off:kirocrew telemetry disableFor an always-on server the README ships a Docker image, ghcr.io/kirodotdev/kirocrew:stable, published on 127.0.0.1:5476. Keep it on loopback and reach it through an SSH tunnel or your VPN rather than a public port.
What can a team on Claude Code, Codex or Cursor borrow from Kiro?
Section titled “What can a team on Claude Code, Codex or Cursor borrow from Kiro?”Each Kiro idea has a portable equivalent. The method moves; the product-enforced gates do not, so you enforce them in CI and review instead.
| Kiro idea | Borrow it as | Claude Code | Codex | Cursor |
|---|---|---|---|---|
| Three-file spec with gates | cc-sdd 3.1.0 skills, or one spec.md per capability | npx cc-sdd@latest (default target) | npx cc-sdd@latest --codex-skills | npx cc-sdd@latest --cursor-skills (beta) |
| Requirements Analysis | The first prompt below, run before design | Plan mode (/plan) | Plan mode (/plan) | Plan Mode |
| Property-based correctness | fast-check or Hypothesis in the test suite, plus Trail of Bits’ property-based-testing skill | Plugin | Plugin | npx skills add |
| Bugfix Spec, preserved behaviour | Characterization tests before the fix | Characterization tests | same | same |
| Steering | Project instruction files | CLAUDE.md | AGENTS.md | Rules |
| Crew schedules and heartbeats | The tool’s own scheduler, or Crew itself with agent.acp_backend set to claude or codex (preview) | Routines | Automations | Cloud agents and automations |
| Crew lessons | Memory and instruction files you review | Memory patterns | same | same |
Set up the borrowed spec-and-properties loop
Section titled “Set up the borrowed spec-and-properties loop”-
Install the spec skills. cc-sdd (“Kiro-style” specs for other agents, npm 3.1.0 on 2026-09-26) writes the same
requirements.md(EARS),design.mdandtasks.mdshape. Preview what it will write with--dry-runfirst. cc-sdd also writesCLAUDE.md(Claude Code) orAGENTS.md(Codex, Cursor) and.kiro/settings/; if you already have an instruction file, run with--overwrite skipor--backupand merge by hand.Terminal window npx cc-sdd@latest --dry-runnpx cc-sdd@latest # Claude Code skills are the default targetThen, in Claude Code:
/kiro-discovery CSV export of the audit log for team admins.Terminal window npx cc-sdd@latest --codex-skillsThe skills land in
.agents/skills/and are invoked as$kiro-discovery,$kiro-spec-initand so on (cc-sdd agent-compatibility guide, 2026-09-26). Use--codex-skills: the legacy--codexmode is blocked.Terminal window npx cc-sdd@latest --cursor-skillscc-sdd labels its Cursor adapter beta. Confirm the
kiro-*skills appear in Cursor before you depend on them; if they do not, keep a plainspec.mdas described in spec-driven development. -
Run discovery, then one phase at a time. The cc-sdd flow is
kiro-discovery→kiro-spec-init→kiro-spec-requirements→kiro-spec-design→kiro-spec-tasks→kiro-impl. Stop afterkiro-spec-requirementsand run the requirements-analysis prompt below before you approve anything. -
Install a property-based-testing skill. Trail of Bits’
property-based-testingplugin (1.2.2) writes, reviews and triages property tests for Hypothesis, fast-check, proptest and others. Only the skill’s description stays in context until it fires.Terminal window claude plugin marketplace add trailofbits/skillsclaude plugin install property-based-testing@trailofbitsTerminal window codex plugin marketplace add trailofbits/skillscodex plugin add property-based-testing@trailofbitsTerminal window npx skills add trailofbits/skills --skill property-based-testing -a cursor -
Write the properties before the implementation. Use the second prompt below. Commit the property tests with the approved requirements, so the implementation task starts red.
-
Implement task by task.
kiro-implgives each task a fresh implementer running red-green TDD and an independent reviewer when the host has subagents; otherwise it runs inline. The property tests are the stop condition. -
Gate the pull request. CI runs the property tests at their default run count on every pull request and a deeper sweep nightly. To make the count settable, register a setup file in the Vitest config and call
fc.configureGlobalthere:vitest.config.ts import { defineConfig } from 'vitest/config';export default defineConfig({test: { setupFiles: ['tests/setup/fast-check.ts'] },});tests/setup/fast-check.ts import fc from 'fast-check';fc.configureGlobal({ numRuns: Number(process.env.FC_NUM_RUNS ?? 100) });Set
FC_NUM_RUNS=10000in the nightly job. The pull request lists each requirement ID with the test that proves it.
How do you prove the borrowed loop works without reading every line?
Section titled “How do you prove the borrowed loop works without reading every line?”Kiro enforces its gates in the product. When you borrow the method, the evidence has to be explicit, and each gate needs an owner:
| Gate | Evidence | Who signs off |
|---|---|---|
| Requirements | The requirements-analysis table has no open rows; every requirement is in EARS form with an example | Product owner or tech lead approves requirements.md |
| Oracle strength | Every requirement maps to a named property or example test; the test file changed in the same pull request as requirements.md, not after the code | Tech lead, from the pull request’s traceability list |
| Correctness | Property tests pass in CI at the default run count; the nightly deep run is green; every shrunk failure is pinned as a regression case | CI |
| Drift | A change to behaviour without a spec delta fails review, as in spec-driven development | Reviewer |
Two further checks keep the oracle honest. Mutation testing on the module under test shows whether the properties can fail at all; Trail of Bits ships a mutation-testing skill in the same marketplace. And the agent that implements a task must not edit the property tests in the same session: see protecting the oracle for the hook and review rules that enforce this, and oracle strength for how to measure it.
Should your team switch to Kiro or borrow from it?
Section titled “Should your team switch to Kiro or borrow from it?”| Situation | Recommendation |
|---|---|
| Your team is on Claude Code, Codex or Cursor and likes the spec discipline | Borrow: cc-sdd plus property tests in CI. No second tool, no second instruction format |
| You want the gates enforced by the product, not by convention, and your stack is AWS-centric | Pilot Kiro on one team, with the bake-off protocol on real closed issues |
| You want an always-on agent on your own hardware, with memory and schedules | Try Kiro Crew on a spare machine. The default harness needs a Kiro login; the Claude Code and Codex harnesses are preview. Compare it with your tool’s own scheduler first |
| You need one harness across many agents | Stay portable: Agent Skills and one instruction file work across tools; see lock-in and portability |
What breaks when you borrow Kiro’s spec and property workflow?
Section titled “What breaks when you borrow Kiro’s spec and property workflow?”Properties that restate the implementation. The agent writes expect(applyDiscount(l, d)).toBe(subtotal - discountFor(l, d)), which passes for any bug the helper shares. Recovery: ban calls into the module under test from the expected side; accept only bounds, round-trips, idempotence or a separate simple reference model. Rerun the second prompt with that rule quoted.
Generators narrowed until green. A failing property is “fixed” by shrinking the input range to values the code handles. Recovery: review generator diffs as carefully as code diffs, and fail the pull request when a generator’s domain shrinks without a matching change to requirements.md.
EARS in form, not in substance. “WHEN the user exports THE SYSTEM SHALL export correctly” has the shape and none of the content. Recovery: the requirements-analysis prompt flags missing outcomes, units and bounds; do not approve design until its table is empty.
Spec and code drift apart after the first release. The three files describe the feature as planned, and later pull requests change behaviour without touching them. Recovery: require a spec delta in every behaviour-changing pull request and audit drift on a schedule, as described in spec-driven development.
cc-sdd installed in legacy mode. --claude, --cursor and the other command modes are deprecated and --codex is blocked (cc-sdd 3.1.0). Recovery: reinstall with the --*-skills flag, after a --dry-run into an empty directory if you customised the old files.
Crew lessons and skills grow without review. Self-evolving skills and lessons change later behaviour. Recovery: review them in the dashboard on a schedule, delete anything you would not accept in a pull request, and scope lessons to one repository with repo_scope.
Crew exposed beyond loopback. A port published on 0.0.0.0 gives anyone on the network an agent with your credentials. Recovery: bind to 127.0.0.1, use an SSH tunnel or VPN, and read Crew’s security model before widening access. See permissions and sandboxing.
Claude Code under Crew skips Crew’s approval path. With agent.acp_backend set to claude, any tool you pre-approved in Claude Code’s own settings (user, project, local or managed) runs without reaching Crew’s deny rules or audit log (Crew changelog 0.6.0). Recovery: keep Claude Code’s allow list minimal on the Crew host, put deny rules in Claude Code’s own settings as well as Crew’s, and treat Crew’s audit log as incomplete for that backend.
Vendor details change under you. Kiro’s spec variants and CLI commands above come from secondary sources, and Kiro ships often. Recovery: re-check kiro.dev before you standardise on a specific command, and pin the version you tested.
Where to go next with Kiro’s ideas
Section titled “Where to go next with Kiro’s ideas”Frequently asked questions
What is Kiro?
Kiro is AWS's agentic development platform. As of 26 September 2026 one agent harness powers its IDE, CLI, Web and Mobile (preview) surfaces and the open-source Kiro Crew workspace, so specs, steering, permissions, hooks, MCP servers and custom agents follow a project across all of them.
How does Kiro use property-based testing?
Kiro turns spec requirements into executable properties and exercises them across generated inputs, to check that the code matches the spec. You can reproduce the same loop on Claude Code, Codex or Cursor with fast-check or Hypothesis and a property-based-testing skill.
What is Kiro Crew?
Kiro Crew is an open-source (Apache-2.0) workspace that runs on your own hardware and keeps working between sessions: persistent memory, checkpointed long-running tasks, schedules, heartbeats, and access from a desktop app, web dashboard, CLI, Slack and Discord. Its default agent runs through kiro-cli and a Kiro login; Claude Code, Codex and OpenCode are selectable harnesses in preview.
Should a team on Claude Code, Codex or Cursor switch to Kiro for specs?
Usually not. The spec shape (requirements, design, tasks), the requirements-analysis pass and the property tests are all portable: cc-sdd installs Kiro-style spec skills for Claude Code and Codex (Cursor in beta), and the property tests are plain test code. Switch only if you want those gates enforced by the product rather than by your own conventions.