1 · The agent says it
“The Global Equity Index Fund has returned an average of 10.5% a year over the last ten years.”
Sounds right. Is it?
Reasoning Markup Language (RML)
AISquare records every run in RML: the claims the agent made, the evidence behind each one, and the assumptions it filled in.

The problem
A transcript tells you what was said.
It doesn't tell you why.
The why disappears
The evidence and shortcuts behind an answer are gone the moment it appears.
Silent assumptions
Agents fill gaps you never see, until one of them is wrong in production.
No answer for auditors
When someone asks why the agent did that, all you have is a wall of transcript.
The trust loop
Trust is earned in a loop. Understand why the agent decided, prevent the bad action next time, and fix what already happened. It starts with seeing the reasoning.
You are here
See why it decided, in RML.
Stop a bad action before it runs.
Correct a decision where it happened.
Carry the fix to every run.
Here's what understanding looks like in practice.
How it works
Every model explains itself differently. RML gives every agent one format for its reasoning, so other agents, reviewers, auditors and your own tools all read it the same way.
ANY AGENT WRITES IT
ANYONE CAN READ IT
Reuse past reasoning as precedent.
Read it in plain language.
Find it in the signed record.
Pull it over the API.
Built on OpenTelemetry, so it fits the tracing you already run.
We're working with the Trust & Safety Institute to publish RML as an open standard.
The catch
Agents state numbers with total confidence. Without the reasoning, you can't tell a real figure from a made-up one.
1 · The agent says it
“The Global Equity Index Fund has returned an average of 10.5% a year over the last ten years.”
Sounds right. Is it?
2 · AISquare checks what backs it
Evidence for this claim
Fund factsheet · retrieved
covers 2019 to 2024 only
3 · You make it a rule
Every return figure must cite a fund source.
Every agent checks it from the next run on.
See the Policy EngineIllustrative example.
What else RML catches
Each one is pinned to the exact step that caused it, so you fix the cause, not the symptom.
No source retrieved in this run contains that figure.
The instructions require it before any decision. The agent decided without it.
The lookup matched the wrong record, and every later step built on it.
Illustrative example.
Reasoning markup (RML)
Every RML record breaks a run into these parts, step by step.
Claim
The refund is within policy limits.
EVIDENCE · 94%
Order #4821, $39, placed 6 days ago
EVIDENCE · 91%
30-day refund window
ASSUMPTION · 0.72
Customer tier resolved to standard · inferred, not confirmed
What the agent asserted and acted on.
What each claim rests on, each with a confidence score.
The gaps the agent filled in for itself.
Extraction confidence
How faithfully the reasoning was captured. Low means read the raw trace, not the summary.
The reasoning goes into the signed record for every run. See Audit
Observability and reasoning
This is the run view your engineers and reviewers open. Here's what every part of it tells you.
I'm 32, medium-high risk tolerance, 15-year horizon...
21.1s · 5 spansPrompt injection: 15
Reasoning
5
Assumptions
8
Claims
1
Evidence used
3
Unbacked
The user has a medium-high risk tolerance and a long horizon.
↓ concluded A global equity fund is appropriate.
A 15-year horizon is long enough for global equity exposure.
I recommend a specific ESG global fund.
UNBACKEDThis fund is ESG-screened and low-cost.
DERIVEDPROMPT INJECTION
1 attempt to override the agent was detected and reviewed.
Every step, in order
Each LLM call, tool call and agent hand-off as it ran.
Latency per step
See which step slowed the run down.
Tool calls
What the agent called, with inputs and timing.
Fixed before it shipped
Decisions your rules corrected, marked on the run.
Cheaper model options
What the same run would cost on other models.
Attacks caught
Attempts to override the agent, flagged with the pattern and severity.
Reasoning, rules and the signed record
RML sits next to Policies and Attestations on every run.
Per step or whole run
Read the reasoning step by step, or for the run.
Analysis confidence
How faithfully the reasoning was captured.
Claims, assumptions, evidence
Counted for every run, so risk is visible at a glance.
Steps with no evidence
Risky steps are marked, so you know where to look.
Unbacked claims
Claims with nothing behind them, flagged in red.
Every LLM call, tool call, retrieval and decision, in order.
See which step slowed the run down.
Per step, per run, per agent, per provider and model.
Failed steps flagged where they happened.
Pass or block, with the reason and the rule it cites.
PII, PHI and PCI detected before anything is saved.
By agent, status, errors, policy hits, cost or duration.
Override attempts flagged with pattern and severity.
Cheaper and pricier model options for the same run.
What the agent asserted and acted on, per step.
What each claim rests on, scored item by item.
The gaps the agent filled in on its own.
How it got from evidence to decision, step by step.
Claims with nothing behind them, and a risk level for the run.
How faithfully the reasoning was captured.
The whole run, explained in a paragraph.
Each finding comes with a plain-language fix.
Low-confidence runs can be routed to a person.
Every item is also available over the API. Read the API guide
In the product
Engineers follow the steps. Reviewers read what happened in plain English. Same run, same record.
7 steps · click any block
INTAKE
Customer request received
LLM
gpt-4o-mini · 802ms
TOOL CALL
Customer Profile · 2ms
LLM
gpt-4o-mini · 1.41s
TOOL CALL
AISquare · Policy Enforcement · 6ms
LLM
gpt-4o-mini · 1.79s
DECISION
Decision: REJECT
Trace overview
Spans
8
Artifacts
12
Duration
15.8s
Cost
$0.0009
Tokens
5,715
A customer asked to be reimbursed for cataract surgery. The agent pulled the customer profile, checked the claim against the policy, and ran the policy check. The policy has a 24-month waiting period for cataract surgery and only 10 months have passed, so the agent rejected the claim and cited the waiting-period rule.
Every step, with latency and cost
The policy check is its own step
The same run in plain English
Same run, same signed record
Both are reading the same RML record.
For developers
curl -s -H "X-API-KEY: $EXPLAINABILITY_API_KEY" \
"$EXPLAINABILITY_GATEWAY_URL/v1/studios/$STUDIO_ID/runs/$RUN_ID/rml"{
"extraction_confidence": 0.86,
"claims": [
"The refund is within policy limits."
],
"assumptions": [
{
"proposition": "Customer tier resolved to standard",
"depends_on": "tier was inferred, not confirmed"
}
]
}Every run's RML record is one API call away. Prefer tools? The learnings MCP server wraps every read for Claude Code, Cursor and any MCP client. Read the API guide
FAQ
RML, or reasoning markup, is how AISquare structures an agent's thinking on each run: the claims it made, the evidence each claim rests on with a confidence score, the assumptions it filled in, and the inference chain that connects them.
OpenTelemetry records what happened: the calls, timings and costs of a run. RML adds the reasoning on top: what the agent claimed, what backed each claim and what it assumed. AISquare's SDK is OpenTelemetry-native, so RML fits the tracing you already run.
We're working with the Trust & Safety Institute to publish RML as an open standard, so any agent or tool can write and read it. Today, every RML record from your agents is yours to read and export over the API.
No. A model's reasoning is its own account of what it did. RML checks each claim against the evidence the agent actually retrieved and the steps it actually ran, so a confident claim with nothing behind it still gets flagged.
A knowledge graph stores what your organization knows. A reasoning graph records how an agent used that knowledge on one specific run: which claims it committed to, what evidence it weighed and what it assumed. They work together, and a knowledge lookup shows up as a step on the reasoning graph.
Every claim is linked to the evidence the agent actually retrieved, with a confidence per item. A claim with weak or missing evidence stands out on the run, so a reviewer sees it and can turn it into a rule for every agent.
Every analysis carries an extraction confidence, and a low score tells you to read the raw trace. Reviewers can accept or dismiss each finding, and those calls sharpen how each rule is graded over time.
Assumptions are facts the agent inferred rather than received, like a customer tier it resolved on its own. They carry a lower confidence and are marked separately, so you can challenge them.
It's how sure AISquare is that it captured the agent's reasoning faithfully. A low score is a prompt to read the raw trace rather than rely on the structured view.
No. Reasoning is extracted asynchronously after a run lands, so the agent runs at its normal speed.
Traces go only to your own AISquare workspace, over TLS. PII, PHI and PCI are screened before anything is saved and can be masked. Talk to us if you have stricter residency needs.
Yes. Because every run is recorded step by step, a required tool call that never happened shows up as a gap on the run, next to the decision that was made without it.
Any agent connected through the proxy or the SDK, including Anthropic, OpenAI, Azure OpenAI and Gemini models, Agno, LangChain, custom agents, and existing OpenTelemetry setups.
Yes. List runs, fetch the graph and fetch the reasoning over the API, or use the learnings MCP server from Claude Code, Cursor or any MCP client.
Stop trusting outputs you cannot explain.
Connect one agent and see its reasoning, assumptions and all, in your own workspace.