Building the guardrails
How to turn “assume it will fail” into a working set of controls
The last post in this series ended with three controls most implementations skip: a hard action budget, a real audit trail, and a kill switch. That’s the right list, but naming controls isn’t the same as building them. This post is about the part that comes after you agree the risk is real: what a working set of guardrails actually looks like, and why no single one of them is enough on its own.
Treat anything the model consumes as a potential threat
Start from one rule, because most of what follows is just this rule applied in different places: anything consumed by the model has to be treated as untrusted until it’s been checked. That includes first and foremost “user input” but also the ones teams forget. A document your RAG pipeline retrieves. A result that comes back from a tool call. A sensor reading from a device on the network. Even the model’s own prior output, if it loops back into context on the next turn.
None of that is exotic; it’s the same principle from earlier in this series applied consistently. Prompt injection works because the model can’t reliably tell instructions from data. Every guardrail in this post exists to enforce that distinction somewhere the model itself can’t be trusted to enforce it.
Assume one control will fail
The mistake that undoes most AI security efforts is picking one defence and trusting it. A hardened system prompt helps, until someone finds the phrasing that gets around it. An input filter catches known attack patterns, until it meets one it hasn’t seen. Treat every single control as something that will eventually fail, and design the next layer to catch what gets past it.
A useful way to organise that thinking is in five layers: prevent the bad thing from being possible, detect it if it happens anyway, contain the damage if detection is too slow, recover cleanly, and prove afterwards what actually occurred. Most teams have something in the first layer and nothing in the other four.
Prevent: keep the boundary structural
Prevention starts with a hardened system prompt and a clear scope boundary, but the prompt is the weakest layer, not the strongest. It communicates intent; it doesn’t enforce it. The stronger version of prevention is a grounding mandate: the model is required to call a tool and cite what it returned, rather than being free to answer from its own guess. Pair that with tool-result isolation where the output of a tool call is treated as data, structurally separated from instructions, the same way a database query result would never be treated as SQL. If the answer isn’t tied to something the tool actually returned, it doesn’t ship.
Detect: assume the input is hostile until checked
An input firewall, a pattern check that runs before anything reaches the model, catches the attacks someone has already seen before. It won’t catch a novel one, and it shouldn’t be sold as if it will. Treat it as one detection layer among several, not proof of safety. Identity and authentication sit next to it: verify who’s making the request, decide what that identity is allowed to do, and log the decision. Without a verified identity, there’s no reliable permission decision, and without a reliable permission decision, there’s no trustworthy record of what happened.
Contain: bound the damage before it starts
Containment is about blast radius, not intent. An autonomy budget turns a runaway loop into a bounded failure instead of an open-ended one. Least privilege does the same job from the access side: an agent that can only reach the tools and data required for its task can’t be talked into misusing permissions it never had. Neither control requires the model to behave correctly. That’s the point.
Recover and prove: build the evidence before you need it
A veto engine provides a second check, separate from the model that produced the answer, screening for policy violations before anything reaches a user or a downstream system, while it adds a layer of review that doesn’t depend on the same model being right twice. It’s not a substitute for human oversight on high-stakes actions; it’s a way to catch what a human reviewer would otherwise have to catch by hand, every time.
None of that matters if you can’t reconstruct what happened afterwards. Signed, hash-linked receipts for each action an agent takes, including what tool was called, what it returned and what decision followed, make tampering visible and give an investigation something to work from. And a break-glass switch, wired to bypass the normal release path entirely, needs to exist before the day someone needs it, not get built during the incident.
Structural, not cosmetic
Every control above lives in code, infrastructure, or the orchestration layer and not in the words of a prompt. That’s consistent with where this series started: the model isn’t a line of defence. It’s the thing the defences are built around. OWASP’s Top 10 for LLM Applications puts prompt injection and excessive agency at the top of its list for exactly this reason: the failure modes are structural, so the fixes have to be too.
None of these eleven controls, taken individually, will stop everything. Layered together, they turn “we hope the model behaves” into a system that fails safely, visibly, and recoverably when it doesn’t.