Engineering organization: running coding agents at company scale
The Engineering organization section shows a CTO or VP Engineering what to standardize once coding agents write much of the code: who owns the shared harness, how much autonomy each delivery loop earns, how policy is enforced on every agent, and how outcomes, spend and hiring are measured and signed off.
Three teams run three agents, the AI line doubled, and nobody owns the shared rules or the merge gate. Each page below settles one decision, so owners, enforced controls and auditable numbers exist before a fourth team starts.
Where is each organization-level question answered?
Section titled “Where is each organization-level question answered?”Each question has its canonical pages and the artifact you adopt from them. The pace of the rollout (stages, stop/go gates, a 12-month plan) lives in the organization-wide transformation roadmap.
| Your question | Canonical page | What you adopt |
|---|---|---|
| Who owns the rules, skills, hooks, MCP servers and evals? | The operating model, the agent platform team | A RACI, a charter and service levels |
| How much may an agent do unattended, and who signs off? | Governance and autonomy | An autonomy policy with risk classes |
| How do we enforce one policy on every agent we run? | Managed policy | Versioned policy files per tool |
| What can an agent be tricked into, and what can it reach? | The agent threat model, agent identity and secrets, security standards | A threat register and an agent identity register |
| What data leaves, and who owns agent-written code? | Privacy and data handling, legal and IP, corporate environment | A data-flow register and an IP position |
| What do we tell engineers is allowed? | An AI usage policy | A policy engineers can follow |
| Which regulations and standards apply? | The EU AI Act, CRA, NIS2 and DORA-EU, regulated industries, ISO 42001, NIST AI RMF, SOC 2 | An AI system register and an applicability register |
| How do we measure outcomes and run a pilot? | Metrics frameworks, DORA AI capabilities, pilot design | Metric definitions and a pilot decision rule |
| What may we spend, and with which vendor? | Cost governance, procurement, model hosting, lock-in and portability | Budgets and alerts, a vendor questionnaire, an exit test |
| What happens when an agent causes an incident? | Agent incidents | An incident runbook |
| How do hiring and career ladders change? | Hiring, career ladders | A four-stage interview loop and a competency matrix |
To find your gaps first, take the free C-Level Scorecard; the CTO track orders these pages into a reading path.
What must exist before a second team gets agents?
Section titled “What must exist before a second team gets agents?”Scaling multiplies whatever controls you have, including the gaps. Each item names its evidence, so an auditor can check it without asking an engineer.
How does leadership know the controls work without reading the code?
Section titled “How does leadership know the controls work without reading the code?”Google Cloud’s announcement of the 2025 DORA report (23 September 2025) states that AI adoption relates positively to delivery throughput and negatively to delivery stability, and explains why: “Without robust control systems, like strong automated testing, mature version control practices, and fast feedback loops, an increase in change volume leads to instability.”
So proof comes from checks a machine runs and numbers a second person can recompute:
- Policy is enforced by the tools. Managed policy overrides developer settings, and a scheduled check compares each machine with git.
- Autonomy is earned per loop. A loop moves up only when its tests, checks and evidence attached to each change meet the gate recorded in the autonomy register.
- Throughput is never reported without stability. Every report pairs merged changes with change failure rate and time in review, as defined in metrics frameworks.
- Sign-off is split. The CTO owns metric definitions, finance signs the cost line, and a named engineering owner approves each autonomy increase.
Where does each tool enforce organization policy?
Section titled “Where does each tool enforce organization policy?”The controls are the same for every tool; the enforcement file differs. Example files, checked on 2026-09-26 against Claude Code 2.1.283 and Codex 0.157.1, are in managed policy.
Organization policy is delivered as server-managed settings from the claude.ai admin console (Team and Enterprise), through MDM managed preferences, or as the system file managed-settings.json (/etc/claude-code/managed-settings.json on Linux). Developers cannot override it, so the MCP allowlist and permission floors belong there.
Admin constraints go in requirements.toml, separate from the config.toml defaults users edit (/etc/codex/requirements.toml on Linux, also deliverable through a workspace layer or macOS MDM). For example, allow_managed_hooks_only = true ignores user and project hooks.
Cursor has team-level and MDM setting layers (the @cursor/sdk 1.0.32 package lists team and mdm setting sources). Which controls each layer enforces is in managed policy; confirm each in Cursor’s enterprise documentation and record the Cursor version.
What goes wrong when a company scales coding agents?
Section titled “What goes wrong when a company scales coding agents?”- Seats arrive before owners. Recover by pausing onboarding until baseline items 1 and 2 exist.
- Policy lives in a document. Recover by moving each rule into the managed policy files, versioned in git.
- Throughput is reported alone. Faros AI’s 2026 report (April 2026, vendor telemetry across 22,000 developers) measured incidents per pull request up 242.7%. Recover by adding stability and review time to every dashboard before the next report.
- Autonomy follows the calendar. Recover by tying each increase to a gate in the autonomy register and rolling a loop back when its change failure rate worsens.
- Agents run on personal tokens. After an incident nobody can tell what the agent could reach. Recover with per-loop identities from agent identity and secrets.