AI-native delivery for software houses and agencies
An AI-native software house sells verified outcomes, not hours. When coding agents do most of the build, time-and-materials billing shrinks revenue, so the firm moves to fixed or unit pricing backed by executable acceptance criteria, discloses agent use in the contract, and rewrites its IP, data, and warranty clauses around the evidence it hands over.
Your sales team is quoting a 1,200-hour time-and-materials engagement. Your delivery lead expects agents to do most of the build in a fraction of that, and the client’s procurement questionnaire asks which AI tools touch its code. Bill the hours you no longer spend and you have a disclosure problem; bill the hours you do spend and revenue falls while review, verification, and model costs rise.
This page is for the owner, managing director, and CTO of a software house, agency, or small studio, including the solo freelancer. It covers the commercial side; the engineering side lives on the linked pages.
What this page gives a software house or agency
Section titled “What this page gives a software house or agency”- A decision table that matches five pricing models to the engagements they fit.
- A fixed-price bid worksheet that prices specification, verification, and warranty explicitly.
- Three margin metrics for a delivery business where hours no longer track output.
- A client disclosure appendix and a contract clause checklist on IP, data, and warranties, with vendor terms checked on 2026-09-26.
- A per-client setup for Claude Code, Codex, and Cursor that keeps client data, credentials, and spend separate.
Why does time and materials break when agents write the code?
Section titled “Why does time and materials break when agents write the code?”Time and materials pays for effort. Agents cut the effort in the build, so the same scope earns less, and the remaining hours move from writing code to specifying, verifying, and reviewing it.
Faros AI’s AI Engineering Report 2026 (April 2026, 22,000 developers, vendor data) found task throughput per developer up 33.7%, bugs per pull request up 28.7% and median time in review up 441.5%, so the gain and the verification bill arrive together. No speed-up figure is safe to promise a client either: METR’s randomized trial found experienced open-source developers took 19% longer to complete issues with AI tools (METR, 2025-07-10), and METR calls its 2026 follow-up “an unreliable signal” (METR, 2026-02-24).
So the firm that keeps billing hours hands the agent gain to the client and carries the verification cost itself. The firm that prices the outcome keeps the gain, provided it can prove the outcome was delivered.
Which pricing model fits which engagement?
Section titled “Which pricing model fits which engagement?”The deciding question: can you write down, before the work starts, a test that says the work is done?
| Model | Fits when | Who carries estimation risk | Where the agent gain goes | What must exist first | The trap |
|---|---|---|---|---|---|
| Time and materials | Discovery, research spikes, legacy code with no tests, unclear scope | Client | Client (fewer billed hours) | Nothing new | Revenue falls as delivery speeds up |
| Capped T&M or a team retainer | Ongoing product work for one client, backlog changes weekly | Shared: the cap moves overrun risk to you | Split, depending on how much the team ships inside the cap | A backlog shaped into verifiable tickets | Judged on visible throughput, so it invites ticket splitting |
| Fixed price per milestone | Scope fits executable acceptance criteria agreed before the build | You | You, after verification and warranty costs | Executable acceptance criteria the client signs, a baseline of your own cost per accepted change | Underbidding because specification, verification, and warranty were not priced |
| Unit price per accepted change | High-volume, similar work: migrations, integrations, test backfill, UI screens from a design system | You, per unit | You | A unit definition, a rate card, a shared acceptance gate | Disputes over what counts as one unit; define it in the contract |
| Outcome-based | You control enough of the system to move a business metric, such as conversion or latency | You, plus attribution risk | You, if the outcome arrives | An agreed baseline, a measurement window, a way to separate your effect from theirs | Paying for outcomes you do not control; keep a fixed floor fee |
Most firms mix them: time and materials for discovery, where the acceptance criteria get written, fixed or unit price for the build, and a retainer for maintenance.
How do you price a fixed-price bid when agents do the build?
Section titled “How do you price a fixed-price bid when agents do the build?”The build becomes the cheapest line in the bid. Specification, verification, acceptance, and warranty used to hide inside “development hours”; now they decide your margin, so price each on its own.
- Baseline your own unit cost. Compute your cost per accepted change on two or three recent agent-delivered projects with the economics method. Without this number a fixed price is a guess.
- Slice the scope into verifiable units during paid discovery. Each unit gets an acceptance criterion the client signs. Units you cannot write a test for stay on time and materials.
- Price discovery, specification, and acceptance as hours. Agents shorten this work with the client least.
- Price the build as units × your unit cost, plus a contingency that rises as the test suite weakens. Legacy code with no tests needs billable characterization tests first.
- Add a warranty reserve from your own history of defects found during warranty, not from a hopeful percentage.
- Publish a price per unit for change requests, which ends the most common fixed-price argument.
Fixed-price bid worksheet
Section titled “Fixed-price bid worksheet”Paste the block into cell A1 of an empty Google Sheets or Excel sheet; the columns are tab-separated and the formulas calculate. The values are illustrative assumptions, not benchmarks: replace every number in column B with your own.
Line Value Source or ownerAccepted changes in scope 120 Scope sliced into verifiable units during discovery — tech leadCost per accepted change on comparable past work (USD) 180 Your last two or three agent-delivered projects — finance and deliveryBuild, verification, and review cost (USD) =B2*B3Discovery and specification hours 160 Estimate before signature, actuals afterAcceptance, demo, and handover hours 60 Include client workshops and the handover of the harnessLoaded cost per hour (USD) 70 FinanceDiscovery, acceptance, and handover cost (USD) =(B5+B6)*B7Warranty reserve rate 0.08 Cost of fixing warranty defects ÷ delivery cost on past projectsWarranty reserve (USD) =(B4+B8)*B9Contingency rate 0.15 Raise it for legacy code with weak testsContingency (USD) =(B4+B8)*B11Total delivery cost (USD) =B4+B8+B10+B12Target gross margin 0.35 ManagementFixed price (USD) =B13/(1-B14)Price per unit for change requests (USD) =B15/B2 Goes into the contract's change-request clauseThe agent finds ambiguity well but does not know your unit cost, so do not take hour estimates from it.
How do margin and staffing change in an AI-native delivery firm?
Section titled “How do margin and staffing change in an AI-native delivery firm?”Agents take over much of the junior work a services pyramid used to bill, so leverage moves from people to the harness: the rules, skills, test oracles, and review agents reused on every engagement. The unit of staffing becomes a small pod.
| Pyramid firm | AI-native pod | |
|---|---|---|
| Shape | One senior, several mid-level, many juniors | A tech lead, one or two engineers who specify and verify, agents doing the build |
| What seniors do | Review, unblock, estimate | Write acceptance criteria with the client, design verification, own the evidence |
| What the firm reuses | People on the bench | Harness assets: AGENTS.md or CLAUDE.md templates, skills, eval suites, CI gates |
| Margin driver | Utilisation × rate × leverage | Price per outcome − (people + usage + verification + warranty) |
| Main risk | Bench time | Underpriced verification, and no juniors growing into leads |
Utilisation stops running the business. Track these three per engagement and per quarter:
- Delivery margin = (fee − people cost − agent usage − CI and verification − warranty cost) ÷ fee. Bill agent usage through or price it in.
- Cost per accepted change per client, by the economics method: the input to your next bid.
- Warranty defect rate = defects the client reports during warranty ÷ accepted changes delivered. If it rises with margin, you are borrowing against future disputes.
What changes for freelancers and two-person studios?
Section titled “What changes for freelancers and two-person studios?”Sell packages, not hours: a fixed price for a defined deliverable (“a Stripe checkout with passing end-to-end tests and a staging deploy”) keeps the speed gain that an hourly rate gives away. Keep one commercial account or API key per client, and hand over the tests, acceptance criteria, and rules file with the code.
How do you keep client data, credentials, and spend separate per tool?
Section titled “How do you keep client data, credentials, and spend separate per tool?”Every client needs its own credentials, data route, and cost attribution.
- Credentials per client: keep one settings file per client outside the repository and start sessions with it, for example
claude --settings ~/clients/acme/settings.json. ItsapiKeyHelperfetches the client’s API key from your secret manager, so the key never sits in the repository or the model’s context. - Tools per client: load only that client’s MCP servers with
--mcp-config ~/clients/acme/mcp.json --strict-mcp-config, which ignores every other configured server. - Spend per client: set
CLAUDE_CODE_ENABLE_TELEMETRY=1and a client tag such asOTEL_RESOURCE_ATTRIBUTES="client=acme,engagement=acme-2026-q4". Claude Code ignores the OpenTelemetry exporter variables, includingCLAUDE_CODE_ENABLE_TELEMETRY, in a repository’s.claude/settings.jsonand.claude/settings.local.json(from v2.1.282, thelatestchannel; onstable2.1.274 a project file can still set them), so set telemetry in the shell, user settings, managed settings, or the per-client--settingsfile. The metricclaude_code.cost.usagethen carries the client tag. Cost governance has the full telemetry setup. - Unattended runs: cap each CI run with
claude -p --max-budget-usd 5 "…"; the flag works only with--print(Claude Code 2.1.283). - For the admin: on machines that sign in to your own Claude for Enterprise or Team organization, set
forceLoginMethodandforceLoginOrgUUIDin managed settings. Those keys blockANTHROPIC_API_KEY,ANTHROPIC_AUTH_TOKEN, andapiKeyHelperat startup (Claude Code 2.1.283), so leave them off machines that run on client-owned keys. There, the per-client--settingsfile with itsapiKeyHelperis the setup, andANTHROPIC_API_KEYstays unset because it outranksapiKeyHelperin authentication precedence.
- Configuration per client: put a profile file at
$CODEX_HOME/acme.config.tomland start withcodex -p acme, which layers that file on top of your base configuration (Codex CLI 0.157.1). It holds the client’s MCP servers and model settings. On engineer machines, give each client its own CODEX_HOME (for exampleCODEX_HOME=~/clients/acme/codex codex), because a profile layers config but shares the stored login:codex loginwritesauth.jsoninto CODEX_HOME (Codex CLI 0.157.1). - Credentials in CI: log in with the client’s own API key from the CI secret store,
printenv ACME_OPENAI_API_KEY | codex login --with-api-key, so usage lands on the client’s API project or on a project you rebill. - Usage and evidence per run:
codex exec --jsonprints each run’s events, including token usage, as JSONL; store them with the pull request number so every run is attributable to a client and a change.
- Separation per client: open each client in its own window from its own repository, never in a multi-root workspace that mixes clients. Where a client’s terms require it, engineers use a separate Cursor account or team for that client.
- MCP servers per client: keep the client’s MCP servers in the repository’s
.cursor/mcp.json, tracked in git so they travel at handover, not in your user-global configuration, where one client’s servers would load in every other client’s sessions. - Retention: Cursor’s plans, team administration, and data-retention settings are not covered here (not re-checked on 2026-09-26). Before opening client code in Cursor, confirm on cursor.com which retention setting applies to each account that will touch it, and record it in the disclosure appendix.
- Rules per client: keep the client’s project rules in its repository, so they travel with the code at handover (context patterns in Cursor).
- Spend: take usage from your plan’s admin reporting, and record the plan and billing mode of each engagement.
What should you disclose to clients about agent use?
Section titled “What should you disclose to clients about agent use?”Disclose enough for the client to make its own data and risk decisions, and no less than its procurement questionnaire asks. A client who finds agent use in a commit trailer or an invoice line talks about trust, not price.
Data terms differ by vendor, plan, and model, so the appendix names models and routes, not only tools. On 2026-09-26, Anthropic’s Commercial Terms of Service (effective 2025-06-17) state that “Anthropic may not train models on Customer Content from Services”. Zero data retention (ZDR) for Claude Code is available to qualified Claude for Enterprise accounts, enabled per organization by Anthropic’s account team; API keys from a commercial organization already under ZDR are covered too, and ZDR never covers data processed by MCP servers or other third-party integrations; where the model runs has the per-model and regional detail. OpenAI’s and Cursor’s current terms were not checked for this page on 2026-09-26; read them on the vendor’s site before you name them.
Client disclosure appendix template
Section titled “Client disclosure appendix template”Adapt this text with counsel, fill the bracketed fields per engagement, and attach it to the statement of work. Confirm the training clause in item 2 against each named vendor’s current commercial terms before you sign, and delete it for any vendor whose terms you could not check.
Appendix C — Use of AI coding agents
1. Tools and models. The Supplier uses the following AI coding agents and models to deliver the Services: [Claude Code with Claude Opus 5.5 via the Anthropic API; Codex with GPT-6 Astra]. The Supplier will notify the Client in writing at least [14] days before adding a tool or model whose data terms are less protective.2. Data route. Client code and data are processed by [vendor] under the Supplier's commercial agreement, in [region], with [zero data retention | retention of N days]. [The vendor's commercial terms, checked on (date), do not permit training on this data.] Personal accounts and consumer plans are never used for Client work.3. Data classes. Agents may read: source code, test data, [synthetic data only]. Agents may not read: production personal data, secrets, [list]. Secrets are injected at runtime and never placed in prompts or context files.4. Human accountability. A named engineer of the Supplier is accountable for every change delivered, whoever or whatever wrote it. Agents do not approve their own work.5. Verification. Each milestone is delivered with an evidence bundle: acceptance tests agreed with the Client, CI results, security and license scans, and a summary of what was not verified. Acceptance is decided against that evidence.6. Client choices. The Client may (a) exclude named repositories or data from agent processing, (b) require use of Client-owned accounts or API keys, and (c) request the agent run logs for any change.7. Handover. At the end of the engagement the Supplier delivers the rules files, skills, prompts, acceptance tests, and CI configuration used to build the deliverables.Name the model your sessions actually run (check /status). Opus 5.5 needs Claude Code v2.1.280+, which on 2026-09-26 is the latest channel only.
The privacy and data-handling policy has the data classification behind clause 3; for EU clients, see the EU AI Act page.
Which contract clauses on IP, data, and warranties need rewriting?
Section titled “Which contract clauses on IP, data, and warranties need rewriting?”Six clauses carry most of the risk once agents write the code. This is a checklist to take to counsel, not legal advice, and the legal and IP page goes deeper on each question.
| Clause | The old wording assumes | What to write instead | Why |
|---|---|---|---|
| Ownership of deliverables | Your staff authored the code, so you can assign copyright in full | Assign all rights you hold, “if any” for machine-generated parts, and warrant your process instead of originality | Copyright in machine-generated code may be limited depending on the jurisdiction. Local form rules still apply: under Poland’s Copyright Act a transfer must be in writing and name each field of exploitation. Vendor terms assign output rights only “if any” (Anthropic Commercial Terms, checked 2026-09-26) |
| Open-source and third-party licenses | Engineers know what they copied | Warrant that deliverables pass a license scan with an agreed allowlist, and ship an SBOM with each release | Agents can reproduce licensed code or add dependencies nobody chose; a CI scan makes the warranty testable |
| IP indemnity | Your indemnity is backed by your own conduct | Cap your indemnity, and do not promise to pass a vendor’s indemnity through | Anthropic’s defense of paid use excludes outputs the customer modified and combinations with non-Anthropic technology, which describes most delivered code (Commercial Terms, checked 2026-09-26) |
| Confidentiality and data | Data stays on your laptops and servers | Name model vendors as subprocessors, list permitted data classes and routes, forbid personal accounts | Client code leaves your premises whenever an agent reads it |
| Warranty | “Free from defects” for N days | Warrant conformance to the signed acceptance criteria for N days; defects are failures against them, everything else is a change request | Makes warranty claims decidable against the evidence you delivered |
| Acceptance | The client tests manually and signs | Acceptance is decided on the evidence bundle and the agreed tests, with a deemed-acceptance period | The client accepts work without reading every line |
How do you prove quality to a client who does not read the code?
Section titled “How do you prove quality to a client who does not read the code?”Nobody reads a thousand agent-written diffs, so quality is shown by evidence the client can check, signed by named people.
- Before the build, the client’s product owner signs the Given/When/Then tests for each unit. They live where the agent cannot edit them (executable acceptance criteria).
- Every change passes the same gates in CI: acceptance tests, type and lint checks, a security scan, and a license scan with an agreed allowlist (dependency verification has a tested CI gate for it). A review agent comments on each pull request, and a human reviews every high-risk change and a sample of the rest (agent pull request review).
- Each milestone ships with an evidence bundle that your tech lead signs.
- The client accepts against the bundle, within the deemed-acceptance period in the contract.
- Report three numbers per quarter to the client’s sponsor: accepted changes, warranty defects, and changes reverted within 30 days (metrics frameworks).
In-house teams make the same shift in verifying evidence instead of diffs.
What goes wrong when a software house goes AI-native?
Section titled “What goes wrong when a software house goes AI-native?”A time-and-materials client learns about agent use and demands a discount. Recovery: disclose first, offer fixed or unit pricing for the next phase, and show the review hours that remain.
A fixed-price bid loses money on verification because the criteria were vague and the legacy code had no tests. Recovery: move unclear scope back to time and materials through a change request, and make paid discovery with signed acceptance tests a condition of the next fixed price.
An engineer runs client code through a personal account. The data now sits under consumer terms that the contract does not cover. Recovery: tell the client under the confidentiality clause, and revoke and rotate anything the session could see. Then enforce commercial access through managed settings that match how the engagement authenticates; the admin note in the Claude Code tab above says which setup applies.
The sales team promises a vendor indemnity the vendor never gave. Recovery: remove pass-through language from the templates, cap your own indemnity, and have counsel read each vendor’s current exclusions.
Margin and the warranty defect rate rise together because pods ship faster than the gates can check. Recovery: run fewer agents per pod and fund stronger acceptance tests before the next fixed-price engagement.
The client is locked out at handover because the deliverables depend on your private skills, prompts, and CI. Recovery: deliver the harness with the code, as clause 7 of the disclosure appendix requires; avoiding lock-in lists which assets travel between tools.
Where to go next for software houses and agencies
Section titled “Where to go next for software houses and agencies”The executive track starts at the executive start page. From here, put a number on your delivery cost, then take the contract questions to counsel.