All work

An AI-Native SDLC for Critical Infrastructure: Six Files, Two Tracks

Agentic SDLC · Governance    Live: from the first idea to the agent that watches production

Coding agents made writing code cheap. That did not make delivery fast; it moved the slow part. A request still has to become clear, someone still has to read what the agent wrote, approve the release and react when production breaks — and at a critical-infrastructure energy operator, every one of those steps has to survive an audit. I designed and led the rollout of Anthropic's AI-native SDLC playbook there: every phase writes one file into Git, agents do the routine work between the phases, and people decide at named gates.

The problem & constraints

Anthropic's AI-native SDLC playbook starts from the same observation: once code generation is fast, the bottleneck moves to planning, design alignment, review, approval and incident response. It answers with six phases in a loop, each ending in a committed file that the next phase reads. For a grid operator, three constraints come on top:

  • Regulation. KRITIS and NIS2 duties apply to grid-relevant software. Every change needs an accountable person, and an auditor must be able to see who asked for what and who approved it.
  • Data. No customer data and no grid data in a prompt, and no model call outside the EU.
  • Untrusted prototypes. Business teams can now build a working prototype with an AI tool in an afternoon. It is a useful starting point and a dangerous deployment.

The approach: six phases, one file each

The playbook's loop, as we run it. Each phase ends with a commit; the next phase starts by reading it.

  1. Plan → intent.md. A business owner describes the problem with an agent. The product owner approves the result.
  2. Design → spec.md. The agent turns the intent into a spec. Company rules load as skills, so security and architecture rules apply while it writes, not in review.
  3. Build → plan.md, then code. The agent plans in plan mode, and the engineer corrects the plan before any code is written. AGENTS.md holds the rules of each codebase.
  4. Test → code and evidence. The agent runs the tests and fixes its own failures before a person looks. The evidence is exported as a file.
  5. Deploy → the pull request and the release. An agent reviews every pull request, a code owner approves, and delivery runs through the existing GitOps gates.
  6. Maintain → the next intent.md. Monitoring detects a breach; a read-only agent diagnoses it and proposes a fix through the same gates, written as a new intent.

Every hand-over is a committed file, not a meeting. The chain intent → spec → plan → code → PR → deployment is the audit trail: who asked for what, what the agent produced and who approved it, straight from the Git history.

Architecture

ONE REQUEST, FROM IDEA TO PRODUCTIONINTAKEGIT · AUDIT TRAILBUILDCHECK & GATEDELIVER & RUNMANAGED SETTINGSStakeholderidea · problem · demoChallenger agentasks until it is clearProduct ownerapproves intent, specStatus dashboardevery step visibleintent.mdone per requestspec.md · plan.mdrules applied as writtenCode + testsand exported evidencePull requestthe commit chainEngineercorrects the planCoding agentplan mode, then codeSkills + agentsrules while it writesHooksdeterministic stopsHuman gateA: 1 owner · B: 2Agent PR reviewbugs · security · specScannersIaC · secrets · CVEsPipelinetests, evidence fileGitOps promotiondev → stage → prodKubernetespolicy-enforcedMonitoringmetrics · logs · alertsAgent diagnosisread-only → new intentACROSS EVERY STEPPolicy on every installmanaged settingsshared plugindeny rulesEU model gateway: every model callLiteLLM proxykey per team + repoVertex AI · EUTwo tracks, one flowA office systems, reporting, tools3 people decide: intent, spec, releaseB grid-relevant systems6 people decide, 2 approvers, no auto-deploy

Stakeholders never touch Git, engineers never leave the managed settings, and every model call takes the EU route. Red outline: where a person decides on Track A — the product owner approves the intent and the spec, a code owner approves the release.

  • Intake without Git. Stakeholders use one chat per app, before and after go-live, and follow every request in a status dashboard. A challenger agent pushes back first: problem before solution, users, constraints, success criteria. It refuses to finish until the track and the data classes are answered. Uploaded prototypes are scanned for secrets, vulnerable dependencies and real personal or grid data, and they enter a repository only as a pull request.
  • Policy as code on every install. Managed settings that engineers cannot override, plus one shared, versioned plugin: eight skills (among them the KRITIS/NIS2 coding rules, data classification and a docs updater), three read-only sub-agents (a security reviewer, a test runner on a small model and a spec checker), and hooks that block secrets in edits, deny deploy commands against grid systems, lint after every edit and refuse to finish with red tests.
  • One model route. The managed settings allow a single model endpoint: a LiteLLM gateway with a virtual key per team and repository, budgets and cost per merged change. It routes to Vertex AI in europe-west4, inside the EU.
  • Delivery reuses the gates that already existed. checkov, trivy, conftest with Rego policies, gitleaks, tflint and actionlint; branch protection and CODEOWNERS; GitOps promotion from dev to stage to production with owner approval; and an OpenTofu plan gate with manual, approved applies for infrastructure. New on top: an agent review of every pull request against a REVIEW.md and the spec.
  • Grid systems stay outside the loop. Agents are read-only there, and a named person releases, with two approvers.

What I changed in the playbook

Most of the playbook works as written. Six changes made it fit an operator of critical infrastructure — each one small, none of them optional:

PhaseThe playbookMy change
PlanA business person describes the problem with the agent; the product owner approves intent.md.One required field: does this touch a grid system? The answer sets the track before design starts. Intents are written in German, the business's own language.
DesignThe agent writes spec.md; company rules load as skills.Two skills must exist before the first spec: grid security duties and data classification. Every spec names the data classes it touches.
BuildThe engineer approves plan.md from plan mode; AGENTS.md holds the rules.No real customer or grid data in prompts or test data. Test data is generated, and the rule is in every AGENTS.md.
TestThe agent tests and fixes its own errors; evaluations run in the pipeline.Test evidence is exported as a file. A pipeline log that is deleted after 30 days is not evidence.
DeployAn agent reviews every pull request; a code owner approves; a release manager authorises production.Grid-relevant code: two approvers and no automatic deployment. Office systems: one code owner is enough. The biggest single change.
MaintainMonitoring detects a breach and calls the agent, which diagnoses read-only and proposes a fix.Agents are read-only in production by default. Write access to a grid system stays with a named person, always.

The two-track rule

The playbook says to keep people in the loop for regulated code. It does not say where regulated code starts — and in critical infrastructure, that line decides everything else. So we drew it once, at the start of every request:

Track ATrack B
ScopeOffice systems, reporting, internal toolsGrid-relevant systems
People decide3 times: intent, spec, release6 times: at every phase
ReleaseOne code owner, then GitOpsTwo approvers, no automatic deployment
Agents in productionRead-onlyRead-only, never deploy

Both tracks use the same six phases and the same files; only the number of human gates differs. The track question is answered in intent.md, and a script checks it again on the changed paths — a rule an auditor can read, not a model's judgement. That is why one playbook serves both.

One request, end to end

  1. Ask. A stakeholder explains the idea, the problem or their own solution, by text or voice, and can upload a prototype.
  2. Challenge. The challenger agent sharpens it into a draft intent. Uploaded code is scanned and kept as seed code.
  3. Approve. The product owner edits and approves the draft. Name and time go into the commit.
  4. Submit. The portal commits intent.md. Seed code arrives as a pull request, never as a direct commit.
  5. Build. An engineer works with the coding agent under the managed settings: spec.md, plan.md in plan mode, then code and tests.
  6. Check. The pipeline runs the tests and exports the evidence; the scanners and the agent review run on the pull request.
  7. Gate. Track A: one code owner. Track B: two approvers. Branch protection and CODEOWNERS make the gate unskippable.
  8. Deliver. GitOps promotes the same version from dev to stage to production, with owner approval for each promotion.
  9. Observe. Stakeholders follow their request in the dashboard. Engineers see metrics, logs and alerts, plus agent usage and cost per team through OpenTelemetry.
  10. Loop. An alert starts a read-only agent diagnosis, and its fix becomes a new intent. A live app's next feature starts the same way, in the same chat: sized S/M/L and kept inside the app's quarterly budget.

Leading the team

I designed the target process and the architecture, mapped the playbook onto the platforms that already ran, led the team that built it, and implemented parts of it myself. Two decisions shaped the rollout. First, reuse before build: the scanners, the GitOps promotion and the infrastructure gates already existed, so the work went into what was missing — intake, policy as code, the model route and the loop back from production. Second, one source for the rules: the shared plugin and the managed settings live in one versioned repository with one named owner, so every engineer works with the same rules, and changing a rule is itself a reviewed pull request.

The hardest part & what I learned

The hardest part was not technical. It was drawing the line between code that may flow and code that must stop: the playbook leaves it open, and every later discussion depends on it. Once the two-track rule was a script instead of an opinion, most arguments ended before they started.

The other lesson: put governance where the agent works, not only where the reviewer works. Rules that load as skills while the agent writes, hooks that make some actions impossible, and one model route that cannot be bypassed do more than any review checklist — and they leave the people at the gates free to judge what only people can judge. I wrote about that split in Rules for the Blast Radius and Review Before You Push.

Source: The AI-native SDLC playbook (Anthropic) and the Claude Academy course on it.