Guardrails
AI Agent Guardrails: What They Catch and What They Miss
AI agent guardrails block bad inputs and outputs, but miss unbacked claims and broken business rules. Here’s what a complete guardrail stack needs.

Senior Growth Manager

In this article — 8 sections
AI agent guardrails are usually the first control a team adds, and for good reason. They’re also where many teams stop, which is how an agent that passes every filter still makes the wrong call.
This post covers what guardrails catch well, the four things they routinely miss, and how to close those gaps without ripping out what you already run.
What AI agent guardrails are
Guardrails are checks that sit around an agent and block or flag behavior outside safe limits. Most fall into three groups:
- Input guardrails catch prompt injection, jailbreak attempts and sensitive data coming in.
- Output guardrails catch toxic content, personal data leaks, off-topic answers and broken formats going out.
- Action guardrails control which tools the agent may call, with what permissions and within what limits.
What guardrails catch well
Guardrails are good at content and access problems: a leaked card number, an injection hidden inside a customer email, an agent trying to call an API it has no business touching. They’re fast, cheap and often deterministic. Every production agent should have them.
What they miss
- Your business rules. A filter knows a refund is a number. It doesn’t know your policy caps refunds at $500 unless an SLA outage applies.
- Unbacked claims. An agent says "a replacement bearing is in stock" when the stock lookup actually timed out. The answer is polite, on topic and contains no personal data, so every content guardrail passes it. The claim is still false.
- Errors in the middle. Most guardrails check the final output, but mistakes happen at step four of twelve, and by the end they look reasonable.
- What your experts already fixed. Static rules don’t learn. When a reviewer corrects an agent, that correction usually lives in one ticket and the next run makes the same mistake.
Add reasoning-level guardrails
The missing layer checks what the agent believed before it acts. At AISquare we do this with RML, our Reasoning Markup Language. It breaks each decision into claims, assumptions and evidence, checks each claim against your systems, and flags anything with no evidence behind it before it becomes a decision.
Take a maintenance agent assessing a compressor. Its claim that vibration exceeds the alert threshold is backed by telemetry. Its claim that a replacement bearing is in stock is unbacked, because the inventory lookup returned nothing. Instead of scheduling a repair with parts you don’t have, the run goes to a person for review.
Enforce business rules at runtime
Write your rules in plain English, version them, and enforce them on every run as pass or fail policy gates. If a support agent proposes a $1,250 refund against a $500 cap, the decision is blocked and sent for manager approval, and the record shows exactly which rule fired.
This matters to auditors as much as to operations. A policy that’s enforced at runtime, with a record of each check, is evidence. A policy in a PDF is a hope. Our guide to agentic governance covers how this fits the wider control set.
Make guardrails learn
When an expert overrides an agent, the fix shouldn’t stay in one ticket. AISquare’s Learning Layer drafts a new rule from the correction and routes it to a person for approval. Once approved, every agent it applies to follows it. Rules stay human-owned, and the same mistake doesn’t happen twice.
A complete guardrail stack
| Layer | What it checks | What it catches |
|---|---|---|
| Input | Injection, sensitive data | Attacks and leaks coming in |
| Output | Toxicity, personal data, format | Unsafe or broken responses |
| Action | Tool permissions, limits | Unauthorized actions |
| Reasoning | Claims against evidence | Confident but unbacked decisions |
| Policy | Your business rules, every run | Decisions that break policy |
| Learning | Expert corrections become rules | The same mistake twice |
Where to start
Keep the guardrails you have. Pick one agent, write down its five most important business rules, turn on reasoning capture, and count the unbacked claims over two weeks. That number usually makes the case for the rest of the stack.
Book a demo to see reasoning-level guardrails on a live agent run.
Read next: Agentic governance: how to get agents into production and Agent observability: why traces are not enough.
Ready to build?
Turn your ideas into interactive AI experiences with the AISquare Creator Studio.
About the author

Senior Growth Manager
Senior Growth Manager at AISquare Studios, writing about what it takes to get AI agents into production: governance, guardrails and observability.




