CALIBRATEDAGENTS Playground
PRODUCT 1 · LIVE

The Gate.

Before your model answers, know whether the answer is even in the context. One call returns a calibrated probability — so "don't answer" becomes a decision you can set, not a guess.

LIVE GATE
 
 
answer absentanswer present
Try it in the playground →
The problem

Confidently wrong is the failure mode.

RAG answers even when the context doesn't contain the answer — fluently, and wrong. Watch it happen, then watch the gate stop it.

WITHOUT A GATE
“What's the early-closure penalty after year five?”
📄 Retrieved: clauses covering years one to three only.
⚠ fabricated — no clause supports this
WITH THE GATE
“What's the early-closure penalty after year five?”
📄 Retrieved: clauses covering years one to three only.
p(answerable) 0.12
Held — the answer isn't in the retrieved clauses. Routed to a specialist rather than guessing.
In the real world

What the gate catches.

Banking

Loan servicing assistant

“What penalty applies if I close after year five?”
Not in the retrieved clauses — held, routed to a specialist.

Support

KB-backed deflection bot

“Why was I double-charged in March?”
Account-specific — the KB can't answer. Escalated, not guessed.

Compliance

Policy copilot

“What's our data-retention window for KYC?”
Stated in the SOP — answered, with the citation attached.
Runs in your infrastructure

Nothing leaves your network.

The gate self-hosts. Your documents, your customers' questions, and the scores never touch our servers — same calibrated behavior as the hosted tier, deployed entirely inside your perimeter. Built for teams where data residency isn't optional.

Talk to us about self-hosting →

YOUR INFRASTRUCTURE documents the gatelocal p(answerable) no egress
The numbers

Measured, not asserted.

Benchmarks against LLM-judge baselines, calibration behavior, and self-hosted results — in the technical report.

Read the Gate technical report →