Guardrails

    AI Agent Guardrails: What They Catch and What They Miss

    AI agent guardrails block bad inputs and outputs, but miss unbacked claims and broken business rules. Here’s what a complete guardrail stack needs.

    ·3 min read
    Omkar Shendge
    Omkar Shendge

    Senior Growth Manager

    In this article — 8 sections

    AI agent guardrails are usually the first control a team adds, and for good reason. They’re also where many teams stop, which is how an agent that passes every filter still makes the wrong call.

    This post covers what guardrails catch well, the four things they routinely miss, and how to close those gaps without ripping out what you already run.

    What AI agent guardrails are

    Guardrails are checks that sit around an agent and block or flag behavior outside safe limits. Most fall into three groups:

    • Input guardrails catch prompt injection, jailbreak attempts and sensitive data coming in.
    • Output guardrails catch toxic content, personal data leaks, off-topic answers and broken formats going out.
    • Action guardrails control which tools the agent may call, with what permissions and within what limits.

    What guardrails catch well

    Guardrails are good at content and access problems: a leaked card number, an injection hidden inside a customer email, an agent trying to call an API it has no business touching. They’re fast, cheap and often deterministic. Every production agent should have them.

    What they miss

    1. Your business rules. A filter knows a refund is a number. It doesn’t know your policy caps refunds at $500 unless an SLA outage applies.
    2. Unbacked claims. An agent says "a replacement bearing is in stock" when the stock lookup actually timed out. The answer is polite, on topic and contains no personal data, so every content guardrail passes it. The claim is still false.
    3. Errors in the middle. Most guardrails check the final output, but mistakes happen at step four of twelve, and by the end they look reasonable.
    4. What your experts already fixed. Static rules don’t learn. When a reviewer corrects an agent, that correction usually lives in one ticket and the next run makes the same mistake.

    Add reasoning-level guardrails

    The missing layer checks what the agent believed before it acts. At AISquare we do this with RML, our Reasoning Markup Language. It breaks each decision into claims, assumptions and evidence, checks each claim against your systems, and flags anything with no evidence behind it before it becomes a decision.

    Take a maintenance agent assessing a compressor. Its claim that vibration exceeds the alert threshold is backed by telemetry. Its claim that a replacement bearing is in stock is unbacked, because the inventory lookup returned nothing. Instead of scheduling a repair with parts you don’t have, the run goes to a person for review.

    Enforce business rules at runtime

    Write your rules in plain English, version them, and enforce them on every run as pass or fail policy gates. If a support agent proposes a $1,250 refund against a $500 cap, the decision is blocked and sent for manager approval, and the record shows exactly which rule fired.

    This matters to auditors as much as to operations. A policy that’s enforced at runtime, with a record of each check, is evidence. A policy in a PDF is a hope. Our guide to agentic governance covers how this fits the wider control set.

    Make guardrails learn

    When an expert overrides an agent, the fix shouldn’t stay in one ticket. AISquare’s Learning Layer drafts a new rule from the correction and routes it to a person for approval. Once approved, every agent it applies to follows it. Rules stay human-owned, and the same mistake doesn’t happen twice.

    A complete guardrail stack

    LayerWhat it checksWhat it catches
    InputInjection, sensitive dataAttacks and leaks coming in
    OutputToxicity, personal data, formatUnsafe or broken responses
    ActionTool permissions, limitsUnauthorized actions
    ReasoningClaims against evidenceConfident but unbacked decisions
    PolicyYour business rules, every runDecisions that break policy
    LearningExpert corrections become rulesThe same mistake twice

    Where to start

    Keep the guardrails you have. Pick one agent, write down its five most important business rules, turn on reasoning capture, and count the unbacked claims over two weeks. That number usually makes the case for the rest of the stack.

    Book a demo to see reasoning-level guardrails on a live agent run.

    Read next: Agentic governance: how to get agents into production and Agent observability: why traces are not enough.

    ai-agent-guardrailsruntime-policy-enforcementguardrails-for-ai-agents

    Ready to build?

    Turn your ideas into interactive AI experiences with the AISquare Creator Studio.

    Book a demo

    About the author

    Omkar Shendge
    Omkar Shendge

    Senior Growth Manager

    Senior Growth Manager at AISquare Studios, writing about what it takes to get AI agents into production: governance, guardrails and observability.

    More from the blog