Reasoning Markup Language (RML)

    See why your agent decided, not just what it did.

    AISquare records every run in RML: the claims the agent made, the evidence behind each one, and the assumptions it filled in.

    AISquare reasoning graph showing claims, evidence and assumptions

    The problem

    A transcript tells you what was said.
    It doesn't tell you why.

    The why disappears

    The evidence and shortcuts behind an answer are gone the moment it appears.

    Silent assumptions

    Agents fill gaps you never see, until one of them is wrong in production.

    No answer for auditors

    When someone asks why the agent did that, all you have is a wall of transcript.

    The trust loop

    You can't trust an agent you don't understand.

    Trust is earned in a loop. Understand why the agent decided, prevent the bad action next time, and fix what already happened. It starts with seeing the reasoning.

    Here's what understanding looks like in practice.

    How it works

    One language for how agents reason.

    Every model explains itself differently. RML gives every agent one format for its reasoning, so other agents, reviewers, auditors and your own tools all read it the same way.

    ANY AGENT WRITES IT

    OpenAI agent
    Claude agent
    Gemini agent
    LangChain
    Agno
    Your own agent
    RML record
    1Run header · agent, model, steps
    2Claims
    3Evidence, with confidence
    4Assumptions
    5Inference chain
    6Unbacked claims flagged

    ANYONE CAN READ IT

    Other agents

    Reuse past reasoning as precedent.

    Reviewers

    Read it in plain language.

    Auditors

    Find it in the signed record.

    Your tools

    Pull it over the API.

    Built on OpenTelemetry, so it fits the tracing you already run.

    We're working with the Trust & Safety Institute to publish RML as an open standard.

    The catch

    The claim sounded right. Nothing backed it.

    Agents state numbers with total confidence. Without the reasoning, you can't tell a real figure from a made-up one.

    1 · The agent says it

    “The Global Equity Index Fund has returned an average of 10.5% a year over the last ten years.”

    Sounds right. Is it?

    2 · AISquare checks what backs it

    Evidence for this claim

    Fund factsheet · retrieved

    covers 2019 to 2024 only

    Performance database · never queried
    No source in this run says 10.5%
    Unbacked claim

    3 · You make it a rule

    Every return figure must cite a fund source.

    Every agent checks it from the next run on.

    See the Policy Engine

    Illustrative example.

    What else RML catches

    Three mistakes traces don't show you.

    Each one is pinned to the exact step that caused it, so you fix the cause, not the symptom.

    RUN 4821 · CLAIMS AGENTHallucination risk · High
    Unbacked claim

    Quoted a 10.5% average return.

    No source retrieved in this run contains that figure.

    Pinned to · Step 6 · model call
    Skipped step

    Never called the eligibility lookup.

    The instructions require it before any decision. The agent decided without it.

    Pinned to · Step 4 · tool call missing
    Error that spread

    A wrong customer match in step 2 flowed into the decision.

    The lookup matched the wrong record, and every later step built on it.

    Pinned to · Step 2 · tool call

    Illustrative example.

    Reasoning markup (RML)

    Claims, evidence, assumptions.

    Every RML record breaks a run into these parts, step by step.

    Claim

    The refund is within policy limits.

    EVIDENCE · 94%

    Order #4821, $39, placed 6 days ago

    EVIDENCE · 91%

    30-day refund window

    ASSUMPTION · 0.72

    Customer tier resolved to standard · inferred, not confirmed

    1. 1Order total $39 is under the $50 auto-approve ceiling.
    2. 2Placed 6 days ago, inside the 30-day window.
    3. 3No prior refund on this order.
    4. 4Therefore auto-approvable without review.

    Claims

    What the agent asserted and acted on.

    Evidence

    What each claim rests on, each with a confidence score.

    Assumptions

    The gaps the agent filled in for itself.

    86%

    Extraction confidence

    How faithfully the reasoning was captured. Low means read the raw trace, not the summary.

    The reasoning goes into the signed record for every run. See Audit

    Observability and reasoning

    Everything you see on every run.

    This is the run view your engineers and reviewers open. Here's what every part of it tells you.

    I'm 32, medium-high risk tolerance, 15-year horizon...

    21.1s · 5 spansPrompt injection: 1
    FlowGraph
    6 steps · click any block
    INTAKE
    Customer request received
    AGENT3.95s
    FinanceGuidanceAdvisor
    LLM1.32s
    gpt-4o-mini
    TOOL CALL2ms
    Recommend Specific Product
    LLM2.57s
    gpt-4o-mini
    DECISION
    DecisionIMPROVED
    Cost across providers2 cheaper options7 pricier optionsactual $0.0006
    ● SuccessDuration 21.1sCost $0.0006Tokens 2,318
    DetailsTraceRMLPoliciesAttestations
    By LLM nodeWhole run
    Analysis confidence62%

    5

    Reasoning

    5

    Assumptions

    8

    Claims

    1

    Evidence used

    3

    Unbacked

    1 · AGENTFinanceGuidanceAdvisorno evidence
    2 · AGENTFinanceGuidanceAdvisor2 unbacked
    REASONED

    The user has a medium-high risk tolerance and a long horizon.

    ↓ concluded A global equity fund is appropriate.

    ASSUMED

    A 15-year horizon is long enough for global equity exposure.

    CLAIMED

    I recommend a specific ESG global fund.

    UNBACKED

    This fund is ESG-screened and low-cost.

    DERIVED

    PROMPT INJECTION

    1 attempt to override the agent was detected and reviewed.

    Delimiter ConfusionHIGH
    1

    Every step, in order

    Each LLM call, tool call and agent hand-off as it ran.

    2

    Latency per step

    See which step slowed the run down.

    3

    Tool calls

    What the agent called, with inputs and timing.

    4

    Fixed before it shipped

    Decisions your rules corrected, marked on the run.

    5

    Cheaper model options

    What the same run would cost on other models.

    6

    Attacks caught

    Attempts to override the agent, flagged with the pattern and severity.

    7

    Reasoning, rules and the signed record

    RML sits next to Policies and Attestations on every run.

    8

    Per step or whole run

    Read the reasoning step by step, or for the run.

    9

    Analysis confidence

    How faithfully the reasoning was captured.

    10

    Claims, assumptions, evidence

    Counted for every run, so risk is visible at a glance.

    11

    Steps with no evidence

    Risky steps are marked, so you know where to look.

    12

    Unbacked claims

    Claims with nothing behind them, flagged in red.

    See the full list
    OBSERVABILITY

    Full trace

    Every LLM call, tool call, retrieval and decision, in order.

    Latency per step

    See which step slowed the run down.

    Tokens and cost

    Per step, per run, per agent, per provider and model.

    Errors and status

    Failed steps flagged where they happened.

    Policy verdicts

    Pass or block, with the reason and the rule it cites.

    Sensitive data screening

    PII, PHI and PCI detected before anything is saved.

    Search and filter runs

    By agent, status, errors, policy hits, cost or duration.

    Prompt injection detection

    Override attempts flagged with pattern and severity.

    Cost across providers

    Cheaper and pricier model options for the same run.

    REASONING

    Claims

    What the agent asserted and acted on, per step.

    Evidence with confidence

    What each claim rests on, scored item by item.

    Assumptions

    The gaps the agent filled in on its own.

    Inference chain

    How it got from evidence to decision, step by step.

    Unbacked claims and hallucination risk

    Claims with nothing behind them, and a risk level for the run.

    Extraction confidence

    How faithfully the reasoning was captured.

    Plain-English summary

    The whole run, explained in a paragraph.

    Suggested fixes

    Each finding comes with a plain-language fix.

    Sent for review

    Low-confidence runs can be routed to a person.

    Every item is also available over the API. Read the API guide

    In the product

    One run, read by an engineer or a reviewer.

    Engineers follow the steps. Reviewers read what happened in plain English. Same run, same record.

    AISquare · Run trace
    Flow
    Graph

    7 steps · click any block

    INTAKE

    Customer request received

    LLM

    gpt-4o-mini · 802ms

    TOOL CALL

    Customer Profile · 2ms

    LLM

    gpt-4o-mini · 1.41s

    TOOL CALL

    AISquare · Policy Enforcement · 6ms

    LLM

    gpt-4o-mini · 1.79s

    DECISION

    Decision: REJECT

    Trace overview

    Spans

    8

    Artifacts

    12

    Duration

    15.8s

    Cost

    $0.0009

    Tokens

    5,715

    NarrativeCustomer account · claim 15.8s

    A customer asked to be reimbursed for cataract surgery. The agent pulled the customer profile, checked the claim against the policy, and ran the policy check. The policy has a 24-month waiting period for cataract surgery and only 10 months have passed, so the agent rejected the claim and cited the waiting-period rule.

    Policy check · passedDecision · Reject
    1

    Every step, with latency and cost

    2

    The policy check is its own step

    3

    The same run in plain English

    4

    Same run, same signed record

    Both are reading the same RML record.

    For developers

    Read the reasoning over the API.

    request
    curl -s -H "X-API-KEY: $EXPLAINABILITY_API_KEY" \
      "$EXPLAINABILITY_GATEWAY_URL/v1/studios/$STUDIO_ID/runs/$RUN_ID/rml"
    response
    {
      "extraction_confidence": 0.86,
      "claims": [
        "The refund is within policy limits."
      ],
      "assumptions": [
        {
          "proposition": "Customer tier resolved to standard",
          "depends_on": "tier was inferred, not confirmed"
        }
      ]
    }

    Every run's RML record is one API call away. Prefer tools? The learnings MCP server wraps every read for Claude Code, Cursor and any MCP client. Read the API guide

    FAQ

    Questions about AI agent explainability and reasoning graphs.

    What is RML?

    RML, or reasoning markup, is how AISquare structures an agent's thinking on each run: the claims it made, the evidence each claim rests on with a confidence score, the assumptions it filled in, and the inference chain that connects them.

    How is RML different from OpenTelemetry?

    OpenTelemetry records what happened: the calls, timings and costs of a run. RML adds the reasoning on top: what the agent claimed, what backed each claim and what it assumed. AISquare's SDK is OpenTelemetry-native, so RML fits the tracing you already run.

    Is RML open?

    We're working with the Trust & Safety Institute to publish RML as an open standard, so any agent or tool can write and read it. Today, every RML record from your agents is yours to read and export over the API.

    Isn't this just the model's own reasoning?

    No. A model's reasoning is its own account of what it did. RML checks each claim against the evidence the agent actually retrieved and the steps it actually ran, so a confident claim with nothing behind it still gets flagged.

    How is a reasoning graph different from a knowledge graph?

    A knowledge graph stores what your organization knows. A reasoning graph records how an agent used that knowledge on one specific run: which claims it committed to, what evidence it weighed and what it assumed. They work together, and a knowledge lookup shows up as a step on the reasoning graph.

    How do you catch a claim with nothing behind it?

    Every claim is linked to the evidence the agent actually retrieved, with a confidence per item. A claim with weak or missing evidence stands out on the run, so a reviewer sees it and can turn it into a rule for every agent.

    How accurate is the flagging?

    Every analysis carries an extraction confidence, and a low score tells you to read the raw trace. Reviewers can accept or dismiss each finding, and those calls sharpen how each rule is graded over time.

    How are assumptions detected?

    Assumptions are facts the agent inferred rather than received, like a customer tier it resolved on its own. They carry a lower confidence and are marked separately, so you can challenge them.

    What is extraction confidence?

    It's how sure AISquare is that it captured the agent's reasoning faithfully. A low score is a prompt to read the raw trace rather than rely on the structured view.

    Does capturing reasoning slow my agent down?

    No. Reasoning is extracted asynchronously after a run lands, so the agent runs at its normal speed.

    Do our traces leave our environment?

    Traces go only to your own AISquare workspace, over TLS. PII, PHI and PCI are screened before anything is saved and can be masked. Talk to us if you have stricter residency needs.

    Can it catch a step the agent skipped?

    Yes. Because every run is recorded step by step, a required tool call that never happened shows up as a gap on the run, next to the decision that was made without it.

    Which frameworks and models does it work with?

    Any agent connected through the proxy or the SDK, including Anthropic, OpenAI, Azure OpenAI and Gemini models, Agno, LangChain, custom agents, and existing OpenTelemetry setups.

    Can I read the reasoning programmatically?

    Yes. List runs, fetch the graph and fetch the reasoning over the API, or use the learnings MCP server from Claude Code, Cursor or any MCP client.

    Stop trusting outputs you cannot explain.

    Start your journey with AISquare

    Connect one agent and see its reasoning, assumptions and all, in your own workspace.