Skip to content

AI-native delivery for software houses and agencies

An AI-native software house sells verified outcomes, not hours. When coding agents do most of the build, time-and-materials billing shrinks revenue, so the firm moves to fixed or unit pricing backed by executable acceptance criteria, discloses agent use in the contract, and rewrites its IP, data, and warranty clauses around the evidence it hands over.

Your sales team is quoting a 1,200-hour time-and-materials engagement. Your delivery lead expects agents to do most of the build in a fraction of that, and the client’s procurement questionnaire asks which AI tools touch its code. Bill the hours you no longer spend and you have a disclosure problem; bill the hours you do spend and revenue falls while review, verification, and model costs rise.

This page is for the owner, managing director, and CTO of a software house, agency, or small studio, including the solo freelancer. It covers the commercial side; the engineering side lives on the linked pages.

What this page gives a software house or agency

Section titled “What this page gives a software house or agency”
  • A decision table that matches five pricing models to the engagements they fit.
  • A fixed-price bid worksheet that prices specification, verification, and warranty explicitly.
  • Three margin metrics for a delivery business where hours no longer track output.
  • A client disclosure appendix and a contract clause checklist on IP, data, and warranties, with vendor terms checked on 2026-09-26.
  • A per-client setup for Claude Code, Codex, and Cursor that keeps client data, credentials, and spend separate.

Why does time and materials break when agents write the code?

Section titled “Why does time and materials break when agents write the code?”

Time and materials pays for effort. Agents cut the effort in the build, so the same scope earns less, and the remaining hours move from writing code to specifying, verifying, and reviewing it.

Faros AI’s AI Engineering Report 2026 (April 2026, 22,000 developers, vendor data) found task throughput per developer up 33.7%, bugs per pull request up 28.7% and median time in review up 441.5%, so the gain and the verification bill arrive together. No speed-up figure is safe to promise a client either: METR’s randomized trial found experienced open-source developers took 19% longer to complete issues with AI tools (METR, 2025-07-10), and METR calls its 2026 follow-up “an unreliable signal” (METR, 2026-02-24).

So the firm that keeps billing hours hands the agent gain to the client and carries the verification cost itself. The firm that prices the outcome keeps the gain, provided it can prove the outcome was delivered.

Which pricing model fits which engagement?

Section titled “Which pricing model fits which engagement?”

The deciding question: can you write down, before the work starts, a test that says the work is done?

ModelFits whenWho carries estimation riskWhere the agent gain goesWhat must exist firstThe trap
Time and materialsDiscovery, research spikes, legacy code with no tests, unclear scopeClientClient (fewer billed hours)Nothing newRevenue falls as delivery speeds up
Capped T&M or a team retainerOngoing product work for one client, backlog changes weeklyShared: the cap moves overrun risk to youSplit, depending on how much the team ships inside the capA backlog shaped into verifiable ticketsJudged on visible throughput, so it invites ticket splitting
Fixed price per milestoneScope fits executable acceptance criteria agreed before the buildYouYou, after verification and warranty costsExecutable acceptance criteria the client signs, a baseline of your own cost per accepted changeUnderbidding because specification, verification, and warranty were not priced
Unit price per accepted changeHigh-volume, similar work: migrations, integrations, test backfill, UI screens from a design systemYou, per unitYouA unit definition, a rate card, a shared acceptance gateDisputes over what counts as one unit; define it in the contract
Outcome-basedYou control enough of the system to move a business metric, such as conversion or latencyYou, plus attribution riskYou, if the outcome arrivesAn agreed baseline, a measurement window, a way to separate your effect from theirsPaying for outcomes you do not control; keep a fixed floor fee

Most firms mix them: time and materials for discovery, where the acceptance criteria get written, fixed or unit price for the build, and a retainer for maintenance.

How do you price a fixed-price bid when agents do the build?

Section titled “How do you price a fixed-price bid when agents do the build?”

The build becomes the cheapest line in the bid. Specification, verification, acceptance, and warranty used to hide inside “development hours”; now they decide your margin, so price each on its own.

  1. Baseline your own unit cost. Compute your cost per accepted change on two or three recent agent-delivered projects with the economics method. Without this number a fixed price is a guess.
  2. Slice the scope into verifiable units during paid discovery. Each unit gets an acceptance criterion the client signs. Units you cannot write a test for stay on time and materials.
  3. Price discovery, specification, and acceptance as hours. Agents shorten this work with the client least.
  4. Price the build as units × your unit cost, plus a contingency that rises as the test suite weakens. Legacy code with no tests needs billable characterization tests first.
  5. Add a warranty reserve from your own history of defects found during warranty, not from a hopeful percentage.
  6. Publish a price per unit for change requests, which ends the most common fixed-price argument.

Paste the block into cell A1 of an empty Google Sheets or Excel sheet; the columns are tab-separated and the formulas calculate. The values are illustrative assumptions, not benchmarks: replace every number in column B with your own.

Line Value Source or owner
Accepted changes in scope 120 Scope sliced into verifiable units during discovery — tech lead
Cost per accepted change on comparable past work (USD) 180 Your last two or three agent-delivered projects — finance and delivery
Build, verification, and review cost (USD) =B2*B3
Discovery and specification hours 160 Estimate before signature, actuals after
Acceptance, demo, and handover hours 60 Include client workshops and the handover of the harness
Loaded cost per hour (USD) 70 Finance
Discovery, acceptance, and handover cost (USD) =(B5+B6)*B7
Warranty reserve rate 0.08 Cost of fixing warranty defects ÷ delivery cost on past projects
Warranty reserve (USD) =(B4+B8)*B9
Contingency rate 0.15 Raise it for legacy code with weak tests
Contingency (USD) =(B4+B8)*B11
Total delivery cost (USD) =B4+B8+B10+B12
Target gross margin 0.35 Management
Fixed price (USD) =B13/(1-B14)
Price per unit for change requests (USD) =B15/B2 Goes into the contract's change-request clause

The agent finds ambiguity well but does not know your unit cost, so do not take hour estimates from it.

How do margin and staffing change in an AI-native delivery firm?

Section titled “How do margin and staffing change in an AI-native delivery firm?”

Agents take over much of the junior work a services pyramid used to bill, so leverage moves from people to the harness: the rules, skills, test oracles, and review agents reused on every engagement. The unit of staffing becomes a small pod.

Pyramid firmAI-native pod
ShapeOne senior, several mid-level, many juniorsA tech lead, one or two engineers who specify and verify, agents doing the build
What seniors doReview, unblock, estimateWrite acceptance criteria with the client, design verification, own the evidence
What the firm reusesPeople on the benchHarness assets: AGENTS.md or CLAUDE.md templates, skills, eval suites, CI gates
Margin driverUtilisation × rate × leveragePrice per outcome − (people + usage + verification + warranty)
Main riskBench timeUnderpriced verification, and no juniors growing into leads

Utilisation stops running the business. Track these three per engagement and per quarter:

  • Delivery margin = (fee − people cost − agent usage − CI and verification − warranty cost) ÷ fee. Bill agent usage through or price it in.
  • Cost per accepted change per client, by the economics method: the input to your next bid.
  • Warranty defect rate = defects the client reports during warranty ÷ accepted changes delivered. If it rises with margin, you are borrowing against future disputes.

What changes for freelancers and two-person studios?

Section titled “What changes for freelancers and two-person studios?”

Sell packages, not hours: a fixed price for a defined deliverable (“a Stripe checkout with passing end-to-end tests and a staging deploy”) keeps the speed gain that an hourly rate gives away. Keep one commercial account or API key per client, and hand over the tests, acceptance criteria, and rules file with the code.

How do you keep client data, credentials, and spend separate per tool?

Section titled “How do you keep client data, credentials, and spend separate per tool?”

Every client needs its own credentials, data route, and cost attribution.

  • Credentials per client: keep one settings file per client outside the repository and start sessions with it, for example claude --settings ~/clients/acme/settings.json. Its apiKeyHelper fetches the client’s API key from your secret manager, so the key never sits in the repository or the model’s context.
  • Tools per client: load only that client’s MCP servers with --mcp-config ~/clients/acme/mcp.json --strict-mcp-config, which ignores every other configured server.
  • Spend per client: set CLAUDE_CODE_ENABLE_TELEMETRY=1 and a client tag such as OTEL_RESOURCE_ATTRIBUTES="client=acme,engagement=acme-2026-q4". Claude Code ignores the OpenTelemetry exporter variables, including CLAUDE_CODE_ENABLE_TELEMETRY, in a repository’s .claude/settings.json and .claude/settings.local.json (from v2.1.282, the latest channel; on stable 2.1.274 a project file can still set them), so set telemetry in the shell, user settings, managed settings, or the per-client --settings file. The metric claude_code.cost.usage then carries the client tag. Cost governance has the full telemetry setup.
  • Unattended runs: cap each CI run with claude -p --max-budget-usd 5 "…"; the flag works only with --print (Claude Code 2.1.283).
  • For the admin: on machines that sign in to your own Claude for Enterprise or Team organization, set forceLoginMethod and forceLoginOrgUUID in managed settings. Those keys block ANTHROPIC_API_KEY, ANTHROPIC_AUTH_TOKEN, and apiKeyHelper at startup (Claude Code 2.1.283), so leave them off machines that run on client-owned keys. There, the per-client --settings file with its apiKeyHelper is the setup, and ANTHROPIC_API_KEY stays unset because it outranks apiKeyHelper in authentication precedence.

What should you disclose to clients about agent use?

Section titled “What should you disclose to clients about agent use?”

Disclose enough for the client to make its own data and risk decisions, and no less than its procurement questionnaire asks. A client who finds agent use in a commit trailer or an invoice line talks about trust, not price.

Data terms differ by vendor, plan, and model, so the appendix names models and routes, not only tools. On 2026-09-26, Anthropic’s Commercial Terms of Service (effective 2025-06-17) state that “Anthropic may not train models on Customer Content from Services”. Zero data retention (ZDR) for Claude Code is available to qualified Claude for Enterprise accounts, enabled per organization by Anthropic’s account team; API keys from a commercial organization already under ZDR are covered too, and ZDR never covers data processed by MCP servers or other third-party integrations; where the model runs has the per-model and regional detail. OpenAI’s and Cursor’s current terms were not checked for this page on 2026-09-26; read them on the vendor’s site before you name them.

Adapt this text with counsel, fill the bracketed fields per engagement, and attach it to the statement of work. Confirm the training clause in item 2 against each named vendor’s current commercial terms before you sign, and delete it for any vendor whose terms you could not check.

Appendix C — Use of AI coding agents
1. Tools and models. The Supplier uses the following AI coding agents and models to
deliver the Services: [Claude Code with Claude Opus 5.5 via the Anthropic API;
Codex with GPT-6 Astra]. The Supplier will notify the Client in writing at least
[14] days before adding a tool or model whose data terms are less protective.
2. Data route. Client code and data are processed by [vendor] under the Supplier's
commercial agreement, in [region], with [zero data retention | retention of N days].
[The vendor's commercial terms, checked on (date), do not permit training on this
data.] Personal accounts and consumer plans are never used for Client work.
3. Data classes. Agents may read: source code, test data, [synthetic data only].
Agents may not read: production personal data, secrets, [list]. Secrets are
injected at runtime and never placed in prompts or context files.
4. Human accountability. A named engineer of the Supplier is accountable for every
change delivered, whoever or whatever wrote it. Agents do not approve their own work.
5. Verification. Each milestone is delivered with an evidence bundle: acceptance tests
agreed with the Client, CI results, security and license scans, and a summary of
what was not verified. Acceptance is decided against that evidence.
6. Client choices. The Client may (a) exclude named repositories or data from agent
processing, (b) require use of Client-owned accounts or API keys, and (c) request
the agent run logs for any change.
7. Handover. At the end of the engagement the Supplier delivers the rules files, skills,
prompts, acceptance tests, and CI configuration used to build the deliverables.

Name the model your sessions actually run (check /status). Opus 5.5 needs Claude Code v2.1.280+, which on 2026-09-26 is the latest channel only.

The privacy and data-handling policy has the data classification behind clause 3; for EU clients, see the EU AI Act page.

Which contract clauses on IP, data, and warranties need rewriting?

Section titled “Which contract clauses on IP, data, and warranties need rewriting?”

Six clauses carry most of the risk once agents write the code. This is a checklist to take to counsel, not legal advice, and the legal and IP page goes deeper on each question.

ClauseThe old wording assumesWhat to write insteadWhy
Ownership of deliverablesYour staff authored the code, so you can assign copyright in fullAssign all rights you hold, “if any” for machine-generated parts, and warrant your process instead of originalityCopyright in machine-generated code may be limited depending on the jurisdiction. Local form rules still apply: under Poland’s Copyright Act a transfer must be in writing and name each field of exploitation. Vendor terms assign output rights only “if any” (Anthropic Commercial Terms, checked 2026-09-26)
Open-source and third-party licensesEngineers know what they copiedWarrant that deliverables pass a license scan with an agreed allowlist, and ship an SBOM with each releaseAgents can reproduce licensed code or add dependencies nobody chose; a CI scan makes the warranty testable
IP indemnityYour indemnity is backed by your own conductCap your indemnity, and do not promise to pass a vendor’s indemnity throughAnthropic’s defense of paid use excludes outputs the customer modified and combinations with non-Anthropic technology, which describes most delivered code (Commercial Terms, checked 2026-09-26)
Confidentiality and dataData stays on your laptops and serversName model vendors as subprocessors, list permitted data classes and routes, forbid personal accountsClient code leaves your premises whenever an agent reads it
Warranty“Free from defects” for N daysWarrant conformance to the signed acceptance criteria for N days; defects are failures against them, everything else is a change requestMakes warranty claims decidable against the evidence you delivered
AcceptanceThe client tests manually and signsAcceptance is decided on the evidence bundle and the agreed tests, with a deemed-acceptance periodThe client accepts work without reading every line

How do you prove quality to a client who does not read the code?

Section titled “How do you prove quality to a client who does not read the code?”

Nobody reads a thousand agent-written diffs, so quality is shown by evidence the client can check, signed by named people.

  1. Before the build, the client’s product owner signs the Given/When/Then tests for each unit. They live where the agent cannot edit them (executable acceptance criteria).
  2. Every change passes the same gates in CI: acceptance tests, type and lint checks, a security scan, and a license scan with an agreed allowlist (dependency verification has a tested CI gate for it). A review agent comments on each pull request, and a human reviews every high-risk change and a sample of the rest (agent pull request review).
  3. Each milestone ships with an evidence bundle that your tech lead signs.
  4. The client accepts against the bundle, within the deemed-acceptance period in the contract.
  5. Report three numbers per quarter to the client’s sponsor: accepted changes, warranty defects, and changes reverted within 30 days (metrics frameworks).

In-house teams make the same shift in verifying evidence instead of diffs.

What goes wrong when a software house goes AI-native?

Section titled “What goes wrong when a software house goes AI-native?”

A time-and-materials client learns about agent use and demands a discount. Recovery: disclose first, offer fixed or unit pricing for the next phase, and show the review hours that remain.

A fixed-price bid loses money on verification because the criteria were vague and the legacy code had no tests. Recovery: move unclear scope back to time and materials through a change request, and make paid discovery with signed acceptance tests a condition of the next fixed price.

An engineer runs client code through a personal account. The data now sits under consumer terms that the contract does not cover. Recovery: tell the client under the confidentiality clause, and revoke and rotate anything the session could see. Then enforce commercial access through managed settings that match how the engagement authenticates; the admin note in the Claude Code tab above says which setup applies.

The sales team promises a vendor indemnity the vendor never gave. Recovery: remove pass-through language from the templates, cap your own indemnity, and have counsel read each vendor’s current exclusions.

Margin and the warranty defect rate rise together because pods ship faster than the gates can check. Recovery: run fewer agents per pod and fund stronger acceptance tests before the next fixed-price engagement.

The client is locked out at handover because the deliverables depend on your private skills, prompts, and CI. Recovery: deliver the harness with the code, as clause 7 of the disclosure appendix requires; avoiding lock-in lists which assets travel between tools.

Where to go next for software houses and agencies

Section titled “Where to go next for software houses and agencies”

The executive track starts at the executive start page. From here, put a number on your delivery cost, then take the contract questions to counsel.