When institutions evaluate AI for regulated work, the question they ask is almost always the same: how accurate is it?
It is the wrong question — or at least an incomplete one. Accuracy is a single number standing in for a much larger judgment: should this institution rely on this answer, defend it to a regulator, and reproduce it a year from now? No single metric can carry that weight. A system can score well on accuracy and still fail the review it was built for.
This is not a hypothetical concern. A system can be highly accurate and still cite the wrong part of a document. It can be accurate and still produce a different answer to the same question on a different run. It can be accurate and still take a reviewer twice as long to trust, because the answer was rephrased rather than quoted. Accuracy does not measure any of these failures — which means a system can look excellent on the one number an institution asked for, while quietly failing the institution in a way that number was never built to catch.
Trust, in a regulated process, is not one property. It is at least seven.
Why one dimension is not enough
Any single dimension, used alone, becomes a blind spot the moment it is optimised for in isolation. A system tuned purely for accuracy has no incentive to remain stable across runs, no incentive to make its reasoning inspectable, and no incentive to know when it is out of its depth.
The questions below are the ones MultiplAI cares about.
Is the answer right?
Necessary, and still not sufficient. A correct answer can be the wrong thing to hand a reviewer if it fails any of the six dimensions that follow.
Correctness is the floor, not the ceiling.
Is it traceable to the exact source?
The literal sentence, in the literal document, that the answer rests on. A synthesised, paraphrased answer can drift from what its citation actually supports without anyone noticing until an auditor asks. At MultiplAI, an answer is only shown if it can be matched verbatim to a snippet in the source.
Is it inspectable without explanation?
A correct, well-evidenced answer can still create work if a reviewer has to dig to understand why the system landed on it. A trustworthy system produces output a second person can evaluate on its own, without a conversation with whoever built it.
Is it stable across runs?
If the same input can produce two completely different answers depending on which run processed it, that difference is hard to explain and impossible to defend when a regulator later asks why.
Is it consistent across similar cases?
Distinct from stability. Stability asks whether the same input produces the same output. Consistency asks whether two similar cases — a near-identical question, a comparable document from a different client — are treated the same way, or whether small, irrelevant differences produce materially different answers.
Does it know what it doesn't know?
A system that always produces an answer, even when the evidence is thin or missing, is more dangerous than one that says so. The safest failure in a compliance workflow is not a wrong answer — it is a system that recognises when it lacks enough support and stops, rather than filling the gap with something plausible.
Is the underlying evidence still current?
Every answer records the age of the evidence it was built on at the moment it was produced. That record is what makes the question answerable in the future — a system that does not track evidence age now will never be able to answer this question later, regardless of how sophisticated it becomes.
A single accuracy score cannot represent seven independent dimensions of risk without hiding which ones it is failing to measure. An institution evaluating AI for a regulated process should ask which of these seven a system is being held to — not just how well it performs on the ones being reported.
A system that is strong on several of these and honest about the rest is more trustworthy than one that offers a single confident number and stays silent on what that number leaves out.
An unsupported answer is a bigger risk than no answer at all.