In August, Anthropic published its playbook for an AI-native software lifecycle. The sentence at its centre is one I have been circling on this blog for months: "Build is no longer the constraint — the human-speed steps around it are." I have since led its rollout at a critical-infrastructure energy operator, where some code may ship on a Tuesday afternoon and some code must never deploy itself. The playbook carried most of the weight. This post is about the part it leaves to you: where the line is, and how to make it hold.
The playbook in one paragraph
Six phases (plan, design, build, test, deploy, maintain) run in a loop. Each phase ends with a committed file, and the next phase starts by reading it: intent.md, spec.md, plan.md, code with its test evidence, the pull request, and, when production breaches a control, the next intent.md. Agents do the work between the files. People approve at named gates, because, in the playbook's words, "Humans remain accountable for every decision that requires judgment." The full architecture is in the case study.
For a regulated company the appeal is not speed. It is that the audit trail writes itself. Every hand-over is a commit, so who asked for what, what the agent produced and who approved it sit in the Git history, not in someone's memory of a meeting.
File one: the field that decides everything
The cheapest place to add governance is the first file. Our intent.md is short, written in German by the person who has the problem, and drafted together with a challenger agent that will not finish until two fields are filled. Here is a made-up example, translated:
# Intent: Repeat failures from outage reports
Requested by: grid operations, planning team
Approved by: <product owner>, <date> ← written by the portal on approval
## Problem
Planners read every outage report by hand to spot repeat failures.
## Users, and what they must never see
Planners. No customer names, no meter IDs.
## Success looks like
Every Monday: repeat failures listed per asset, checked by one planner.
## Does this touch a grid system? ← required
No. It reads exported reports only. → Track A
## Data classes ← required
Internal operational data. No personal data.
The two required fields are the whole trick. Everything downstream keys off them: which skills load, how deep the review goes, how many people approve. Ask at the end and it is too late; by then someone has built the thing.
The line the playbook leaves to you
The playbook tells you to keep people in the loop for regulated code. It does not tell you where regulated code starts. In critical infrastructure that line decides everything else, so we drew it once:
| Track A | Track B | |
|---|---|---|
| Scope | Office systems, reporting, internal tools | Grid-relevant systems |
| People decide | 3 times: intent, spec, release | 6 times: at every phase |
| Release | One code owner, then GitOps | Two approvers, no automatic deployment |
| Agents in production | Read-only | Read-only, never deploy |
Both tracks use the same six phases and the same files. Only the number of gates changes, which is why one playbook can carry an office tool and a grid system at the same time.
Let a script decide, not a model
The track is answered in intent.md, and checked again on the changed paths by a script:
# classify.py: the track comes from the changed paths, never from a model's judgement.
import fnmatch
import subprocess
GRID_PATHS = ["grid/*", "scada-*/*", "deploy/grid-*"] # example patterns
changed = subprocess.run(
["git", "diff", "--name-only", "origin/main...HEAD"],
capture_output=True, text=True, check=True,
).stdout.split()
track_b = any(fnmatch.fnmatch(path, pattern) for path in changed for pattern in GRID_PATHS)
print("TRACK_B" if track_b else "TRACK_A")
The case for a model is real: it would catch what a path list misses, such as a helper module that only grid code imports. That is exactly why it is the wrong tool here. An auditor cannot read a model's judgement, and a judgement that shifts with the prompt is not a control. A path list is boring, reviewable and wrong in predictable ways. When it is wrong, the fix is a one-line pull request that a person approves.
Guardrails where the agent works
Review catches problems after they exist. The cheaper move is to make some of them impossible while the agent is still typing. That happens in three places, all of them outside the developer's reach.
1. Managed settings that nobody can override
Claude Code reads a managed-settings.json from a system directory (/Library/Application Support/ClaudeCode/ on macOS, /etc/claude-code/ on Linux) and ranks it above every user, project and command-line setting (docs). Here is the shape of such a file, trimmed to the parts that matter:
{
"permissions": {
"deny": ["Read(./.env)", "Read(./secrets/**)"],
"disableBypassPermissionsMode": "disable"
},
"allowManagedPermissionRulesOnly": true,
"allowManagedHooksOnly": true,
"allowedProviders": ["customEndpoint"],
"env": {
"ANTHROPIC_BASE_URL": "https://llm-gateway.example.internal",
"CLAUDE_CODE_ENABLE_TELEMETRY": "1",
"OTEL_METRICS_EXPORTER": "otlp",
"OTEL_EXPORTER_OTLP_PROTOCOL": "grpc",
"OTEL_EXPORTER_OTLP_ENDPOINT": "https://otel.example.internal:4317"
},
"hooks": {
"PreToolUse": [{
"matcher": "Bash",
"hooks": [{
"type": "command",
"command": "/opt/sdlc/hooks/deny-grid-deploy.sh"
}]
}]
}
}
Three keys do most of the work. allowManagedPermissionRulesOnly makes this file the only source of permission rules, and allowManagedHooksOnly does the same for hooks. allowedProviders: ["customEndpoint"], together with the gateway URL in env, makes the gateway the only place a session may send a prompt: Claude Code refuses a session pointed anywhere else (docs). The telemetry variables send usage and cost per team to the monitoring stack everyone already reads (docs).
2. Hooks: rules that cannot be argued with
A rule in AGENTS.md is a request. A hook is a fact: the tool call does not happen. Four of them carry most of the load:
| Hook | Runs | Stops |
|---|---|---|
PreToolUse on edits | before a file is written | secrets in code |
PreToolUse on Bash | before a command runs | deploy commands against grid systems |
PostToolUse on Edit|Write | after every edit | unformatted, unlinted code (it fixes rather than blocks) |
Stop | when the agent wants to finish | finishing with red tests |
The one that matters most in our world, simplified:
#!/usr/bin/env bash
# PreToolUse hook: no agent deploys to a grid system. Exit code 2 blocks the call.
cmd=$(jq -r '.tool_input.command // ""')
if grep -Eq '(kubectl|helm|flux)\b.*\bgrid-' <<<"$cmd"; then
echo "Blocked: grid systems are Track B. A named person releases them." >&2
exit 2
fi
exit 0
3. Skills while it writes, read-only reviewers after
The KRITIS and NIS2 coding rules and the data classification are skills: the agent loads them when the task matches, so the rules apply while it writes, and they cost nothing when they do not apply. Reviewers are sub-agents with their own context and read-only tools (docs):
---
name: security-reviewer
description: Reviews the current diff for security issues before a pull request. Read-only.
tools: Read, Grep, Glob
model: sonnet
---
You review; you never edit. For every finding, give the file, the line,
the risk and the fix. Check the diff against spec.md and the security-rules skill.
The test runner works the same way on a small model and returns only the failures, so the main session stays small. None of this replaces the person at the gate. It changes what reaches them: fewer obvious problems, more decisions that need a human.
Evidence, not logs
One rule looks like paperwork, but it is the one that makes the rest provable: test evidence is exported as a file and kept with the release. A pipeline log that is deleted after 30 days is not evidence. When an auditor arrives in month seven, the question is not "did the tests run?" but "show me", and a link to an expired CI run is not an answer.
What it costs
I try to name the cost of every recommendation on this blog. Here it is:
- Track B is slower, on purpose. Six gates and two approvers mean a grid-relevant change waits for people. The agent's speed does not shorten a release that needs two named humans; plan the calendar around that.
- The intake asks before it builds. The challenger agent adds a conversation before any code exists. A stakeholder with a finished prototype still goes through it, and the prototype enters the repository only as a pull request.
- Rules need owners. Eight skills, three sub-agents and four hooks are code. They live in one versioned plugin with one named owner and change through reviewed pull requests, like everything else.
- The EU route costs extra. On Google Cloud, regional and multi-region endpoints carry a 10% premium over the global endpoint (Anthropic docs). For a grid operator that premium is not up for debate, but it belongs in the budget.
- The gateway is now your infrastructure. It has to keep up with Claude Code releases, or new features break without warning (docs).
The Bottom Line: The playbook gives you six files and a loop. Your job is to decide, once and in writing, which code may flow through it unattended, and then to turn every other rule into something the agent cannot skip rather than something a reviewer has to remember.
Related: the full architecture is in the case study An AI-Native SDLC for Critical Infrastructure. The thinking behind the guardrails is in Rules for the Blast Radius and Review Before You Push. Sources: the playbook and the Claude Academy course.