Software built outside engineering
Software built outside engineering is the internal tools, reports, and prototypes that analysts, operations staff, and domain experts now build themselves with coding agents and app builders such as Lovable, Replit, v0, and Bolt. It pays off for disposable prototypes and single-team tools on approved data. Anything with real users, sensitive data, or external effects moves into engineering’s verified delivery.
This page is for executives and CTOs. The situation it answers: your head of operations demos a scheduling app on Monday that she built over the weekend in an app builder. It already has 30 colleagues using it, it reads the HR system through her personal API token, and nobody in engineering has seen the code. You want more of this energy, not less. You also need to know what happens when she goes on holiday and the app breaks.
What you can adopt for builder-made software
Section titled “What you can adopt for builder-made software”- A payoff table that tells you which kinds of builder-made software are worth encouraging and which belong to engineering from day one.
- A four-lane model (disposable, controlled pilot, production, forbidden) that matches the prototype policy in the CTO scorecard.
- An intake record you can paste into a form, so every builder-made tool has an owner, a lane, and an expiry on the day it is created.
- A graduation procedure with copy-paste prompts, which turns a working prototype into acceptance criteria that engineering can verify without reading every line.
- The red lines that no lane crosses, the metrics that show whether the programme works, and the questions to put to your CTO or platform lead.
This page builds on the economics of agent-built software; read it first for how to count the cost of graduation.
Who builds software outside engineering now?
Section titled “Who builds software outside engineering now?”Three groups build software outside engineering, and they need different controls.
| Group | Typical tools | What they build | What they usually lack |
|---|---|---|---|
| App builders | Lovable, Replit, v0, Bolt: a chat prompt produces a running web app | Dashboards, intake forms, small workflow apps, clickable prototypes | A repository your company owns, tests, a named maintainer |
| Agent users outside engineering | The Claude Code desktop app, the Codex desktop app, Cursor | Scripts, data pipelines, integrations, spreadsheet replacements | Sandboxing, secret handling, code review |
| Analysts | The same agents plus SQL and notebooks | Reports, one-off analyses, recurring data jobs | Version control, reproducibility, data classification |
Bolt’s own repository describes the category’s loop plainly: “Prompt, run, edit, and deploy full-stack web applications” (stackblitz/bolt.new, read 2026-09-26). Plans, prices, and enterprise controls for Lovable, Replit, v0, and Bolt change often; confirm them in each vendor’s current documentation (this page does not state them, as of 2026-09-26) before you sanction one, and use the questions in how to choose a sanctioned builder below.
Andrej Karpathy draws the line this page builds on: “Vibe coding raises the floor. Agentic engineering is about extrapolating the ceiling” (Sequoia Ascent 2026 summary, 2026-04-30). Earlier in the same passage he says: “You are still responsible for your software, just as before.” Builders outside engineering raise the floor. Your job is to decide where the floor ends and who owns what sits above it.
Where does software built outside engineering pay off?
Section titled “Where does software built outside engineering pay off?”Builder-made software pays off where the person who understands the problem can also judge whether the result is right, and where a mistake stays small. The table ranks the common cases.
| Use case | Payoff | Why | Default lane |
|---|---|---|---|
| Clickable prototype to settle what a feature should do | High | Replaces weeks of requirement documents with something users react to; the prototype becomes the spec | Disposable |
| One-off analysis or report on approved data | High | The analyst can check the numbers against a known source | Disposable |
| Single-team internal tool on non-sensitive data | Medium to high | Removes a queue for engineering; the team can tell when it is wrong | Controlled pilot |
| Recurring data job feeding other teams’ decisions | Medium | Useful, but others now depend on it and cannot see its errors | Controlled pilot, graduation review at the first new consumer |
| Automation that writes to a system of record (CRM, ERP, HR) | Low without engineering | A silent error corrupts shared data; rollback is hard | Production from day one |
| Anything customer-facing, with logins, payments, or personal data | Negative without engineering | Security, privacy, and liability exceed what the builder can verify | Production or forbidden |
Two things make the payoff real. First, the builder must be able to verify the output against something they already trust: a known total, a manual process, a colleague’s judgment. Second, the tool must be cheap to throw away. When either condition fails, the value moves to engineering, because only engineering has the tests, review, and rollback that make the software trustworthy.
Google Cloud’s announcement of the 2025 DORA report (2025-09-23) puts the organisational side this way: “AI doesn’t fix a team; it amplifies what’s already there.” Builders outside engineering amplify whatever data discipline and ownership habits your company already has.
Which lane does each builder-made tool belong in?
Section titled “Which lane does each builder-made tool belong in?”Use the four lanes from the prototype policy. The lane follows exposure and risk, never the tool that produced the code. The table adds who owns each lane and what the builder may do in it.
| Lane | Boundary | Builder may | Owner | Evidence before anyone relies on it |
|---|---|---|---|---|
| Disposable exploration | Synthetic or approved data, no external users, no writes to shared systems | Build and share a demo; expiry of 30 days or less | The builder | Purpose, owner, and expiry in the inventory |
| Controlled pilot | Named users inside one team, approved data, reversible effects | Run it for the team on the sanctioned platform | The builder, plus a named engineering sponsor | Acceptance criteria, a test run against them, monitoring, a rollback note |
| Production system | Other teams, customers, persistent data, or material effects | Contribute the prototype and the acceptance criteria | An engineering team | The full lifecycle loop: spec, tests, review, evidence bundle, named approver |
| Forbidden path | Secrets, unapproved personal or regulated data, payments, destructive access | Nothing; stop and use the approved route | Security | Not applicable |
The 30-day expiry is a starting value, not a benchmark. Pick yours and write it down. What matters is that every disposable tool has one, so an expired tool is either renewed deliberately or deleted.
What should the intake record capture?
Section titled “What should the intake record capture?”Record every builder-made tool when it is created, not when it breaks. The template below fits in any form tool. Keep it to one screen, or builders will skip it.
Tool name:Builder (name, team):Engineering sponsor (required for controlled pilot and above):Purpose, in one sentence:Lane: disposable | controlled pilot | production | forbiddenUsers today (count, teams):Data it reads (system, classification):Data it writes (system, or "none"):External effects (email, payments, API calls to third parties, or "none"):Credentials used (service account name, never a personal token):Where the code lives (company repository URL, or builder workspace URL):Expiry date:How the builder checks it is right (the trusted source or manual check):The last field is the most important one. A builder who cannot say how they check the output is running a tool nobody verifies.
How does a prototype graduate to production?
Section titled “How does a prototype graduate to production?”A prototype graduates when a trigger fires: new users outside the team, sensitive data, writes to shared systems, or someone else depending on it. Graduation turns the prototype into a specification that engineering implements and verifies, instead of a codebase engineering inherits blind.
-
Freeze the behaviour. The builder records what the tool does today: screens, inputs, outputs, and three to five real examples with the expected result. The examples become test cases.
-
Extract acceptance criteria from the prototype. An engineer, or the builder with an engineer reviewing, runs an agent over the prototype’s code and the builder’s examples to draft acceptance criteria. The builder confirms each criterion matches what users actually need.
-
Classify the risk. Place the future system in the four risk tiers your engineering organisation uses. Authentication, payments, or production data make it Tier 3 no matter how small the code is.
-
Decide: harden in place or rebuild. Use the decision table below. Neither answer is the default.
-
Implement under engineering’s gates. The owning team works in a company repository with CI, tests derived from the acceptance criteria, security review, and an evidence bundle. The builder signs off that the behaviour matches; the engineering owner signs off that it is safe to run.
-
Retire the prototype. Revoke its credentials, delete its data copies, take down its domain, and cancel its builder workspace. Record the retirement in the inventory.
| Signal | Harden in place | Rebuild in the production stack |
|---|---|---|
| Code can move to a company repository | Yes | No, or only as an export nobody can run |
| Stack matches something your platform already runs | Yes | No |
| Data model is sound for the real volume and access rules | Yes | No, or unknown |
| Security review finds isolated issues | Yes | Systemic issues (auth, secrets in code, no access control) |
| The prototype’s main value is the behaviour it demonstrates | Either | Usually rebuild: keep the spec, drop the code |
Rebuilding does not waste the prototype. The expensive part of most internal tools is discovering what they should do, and the prototype has already paid for that.
The dependency check in the second prompt matters more for builder-made code than for engineering code. Researchers found that 19.7% of 2.23 million code samples from 16 models referenced at least one package that does not exist (USENIX Security 2025, Spracklen et al., as reported second-hand by Aikido and the Cloud Security Alliance). A builder who has never run a package manager cannot spot one.
How does engineering run the graduation review in each tool?
Section titled “How does engineering run the graduation review in each tool?”The prompts above work in any of the three agents. What differs is how you keep the review read-only and where the second opinion comes from.
Start the session in plan mode so the agent reads and reports without editing: claude --permission-mode plan. Paste the graduation prompt. Then run /security-review for a dedicated security pass, and /code-review on the pull request once the hardened version exists. For a builder who should only read and analyse, claude --restricted removes the built-in tools that run commands or code, as well as WebFetch, and ignores user, project, and local settings files, while managed settings still apply (CLI flag, checked on v2.1.283).
One default matters for builders outside engineering: from v2.1.283 (the latest release channel), interactive terminal and VS Code sessions start in auto mode on supported models, where a classifier approves actions instead of the person (claude -p, the Agent SDK, and sessions on unsupported models such as Haiku still start in Manual). That suits an engineer who can read a shell command. For builders who cannot, decide deliberately, and switch it off for the whole organisation with permissions.disableAutoMode: "disable" in managed settings.
Import the prototype into a company repository first. Run the graduation prompt in the Codex CLI or desktop app with a read-only permission profile: codex -c default_permissions=":read-only" (permission profiles are beta, checked on 0.157.1). After hardening, codex review --base main reviews the branch non-interactively, which fits as a CI step. Admins can pin permission profiles and restrict MCP servers for everyone through requirements.toml.
Open the imported repository and start in Plan Mode, which “creates detailed implementation plans before writing any code” (Cursor docs, checked 2026-08-28), then paste the graduation prompt. Once the hardened version is a pull request, Bugbot reviews it for bugs, security issues, and code-quality problems. Cursor’s current admin controls for non-engineering seats could not be verified from cursor.com on 2026-09-26; confirm them with your account team before you hand Cursor to builders.
Where does governance draw the line?
Section titled “Where does governance draw the line?”Some lines hold in every lane, for every builder, in every tool. Write them into your AI usage policy and enforce them technically, because a policy document does not stop a pasted token.
| Red line | Why | How to enforce it |
|---|---|---|
| No personal API tokens or production credentials in builder tools | A personal token gives the app its builder’s full access, and it leaves with them | Service accounts with narrow scopes, issued through the agent identity and secrets process |
| No personal or regulated data without a data-protection decision | The builder vendor becomes a processor; retention and location matter | Data classification plus the vendor review in privacy and data handling |
| No payments, no authentication for external users | These are the failures that become incidents and liability | Forbidden lane; engineering builds them |
| No writes to a system of record without an engineering owner | Silent data corruption spreads to every downstream team | API gateways that allow writes only from registered service accounts |
| No public exposure without review | A public URL turns an internal toy into an attack surface | SSO in front of every sanctioned builder; public publishing disabled by default |
| No tool without an owner and an expiry | Orphaned tools become unowned production | Inventory with automatic expiry reminders |
Regulated firms add their own lines. If a builder-made tool touches decisions about people, such as hiring, credit, or access to services, check it against the EU AI Act before it leaves the disposable lane.
How do you choose a sanctioned app builder?
Section titled “How do you choose a sanctioned app builder?”Offer one sanctioned route per group, so people stop using whatever they found. DORA’s guidance on AI policy explains why a sanctioned route beats a ban: “Ambiguity creates risk. A clear policy provides the psychological safety developers need to experiment effectively” (Google Cloud blog, 2025-12-10). The same holds for builders outside engineering.
Ask every vendor these questions, and put the answers in the procurement record:
- Can we enforce company SSO and disable personal accounts on the company domain?
- Can the code sync to a repository our organisation owns, and does the export run outside the vendor?
- Can we disable public publishing, or require SSO in front of every published app?
- Where are app data and prompts stored, for how long, and are they used for training?
- Can we set a spending cap per workspace and see usage per user?
- Can admins list every app in the organisation, with its owner and last activity?
A builder that fails questions 1 or 2 can still serve the disposable lane on synthetic data. It cannot host a controlled pilot.
For the agent-user and analyst groups, the sanctioned route is usually the same agent engineering already runs, set up with tighter managed settings and without production credentials. That keeps one set of controls and lets graduation happen inside the same toolchain.
How do you know the programme is working?
Section titled “How do you know the programme is working?”Measure the programme, not the builders. Each metric has an owner and a definition you can adopt as written.
| Metric | Definition | Owner | Healthy direction |
|---|---|---|---|
| Inventory coverage | Registered builder-made tools ÷ tools found by discovery (SSO logs, builder admin consoles, expense reports), per quarter | CTO’s platform or security lead | Rising towards all of them |
| Owner-and-expiry rate | Registered tools with a named owner and a future expiry date ÷ all registered tools | Platform lead | Near all |
| Graduation lead time | Days from a graduation trigger firing to the tool running under engineering’s gates, or being retired | Engineering sponsor | Stable or falling |
| Unowned production | Tools with users outside the builder’s team and no engineering owner | CTO | Zero |
| Incidents from builder-made tools | Security or data incidents whose root cause is a builder-made tool, per quarter | Security | Zero, and every one reviewed |
| Retirement completeness | Retired tools whose credentials, data, and domains were verified removed ÷ all retired tools | Platform lead | All |
Do not report “hours saved by citizen developers” to the board. Nobody measures the counterfactual, and time saved is not cash saved, as the economics of agent-built software explains. Report the metrics above, plus the named tools that graduated and what they replaced.
What goes wrong with software built outside engineering?
Section titled “What goes wrong with software built outside engineering?”A disposable tool becomes production without anyone deciding it. It gains users, then data, then a dependency. Recovery: run discovery monthly, and treat any tool with users outside its team as a graduation trigger that fires automatically.
A personal token powers a shared app. The builder leaves, the token is revoked, and a team’s workflow stops overnight. Recovery: rotate to a service account before anything else, then register the tool and assign a sponsor.
Engineering becomes the gate everyone routes around. If graduation takes a quarter, builders stop registering tools. Recovery: publish the graduation lead time, staff a small rotation of engineering sponsors, and make the disposable lane need no approval at all.
Engineering rebuilds and loses what the prototype knew. The rebuilt tool is cleaner and wrong. Recovery: make the builder’s examples the acceptance tests, and require the builder’s sign-off on behaviour before release.
Builder sprawl. Five teams use five builders, each with its own data copy. Recovery: sanction one builder per group, and migrate the rest at their next graduation or expiry.
Cost surprises. Usage-based builder plans grow with enthusiasm, not value. Recovery: a spending cap per workspace, reviewed alongside cost governance for engineering’s own agents.
Questions to put to your CTO or platform lead about builder-made tools
Section titled “Questions to put to your CTO or platform lead about builder-made tools”A CTO can use the same list as a self-check before the board asks.
- How many builder-made tools do we run today, and how do we know that number is complete?
- Which of them have users outside the builder’s team, and who owns those?
- Which builder tools are sanctioned, and do they meet the six procurement questions?
- How long does graduation take, and how many tools graduated or retired last quarter?
- Which builder-made tools use personal tokens or touch personal data?
- What would we see first if one of them caused an incident?
Where to go next with builder-made software
Section titled “Where to go next with builder-made software”When a builder-made tool graduates, the owning team follows acceptance criteria and the evidence bundle like any other change.