XemaSConstitutional Document

Why XemaS Refuses Certainty It Doesn't Have

If you've ever seen a security tool label something “safe” and felt uncertain about why you were supposed to believe that, this is written for you.

I  ·  What we observed

Over the past year, we made four decisions that seemed, at the time, to be unrelated.

Removing the word SAFE

We removed SAFE from every output the platform produces. Not because the platform became less accurate. Because we realised that in some situations we simply didn't know enough to use that word. Coverage was incomplete. Analysis was partial. The evidence hadn't been gathered. But the output said SAFE anyway, as though our silence about what we hadn't assessed was equivalent to verification that everything was fine.

It wasn't. We changed the word to CLEAR, which means only that we didn't detect a specific threat. Not that none exists. We introduced coverage states: explicit indicators that tell you not just what we found, but what we assessed and what we didn't. FULL. PARTIAL. UNVERIFIED. These aren't confidence signals. They're admissions about the limits of the analysis.

Preserving UNKNOWN

When we couldn't determine risk level, when the evidence was genuinely absent or ambiguous, the simpler choice was to return a low-risk score. Something had to fill the field. LOW was defensible.

We decided not to. LOW means we assessed the risk and found it to be low. UNKNOWN means we assessed the risk and couldn't determine it. Those aren't the same thing. Collapsing them would have been cleaner. It would also have been false.

Distinguishing "not observed" from "verified absent"

We started requiring assessments to distinguish between two phrases. Not observed means we looked and didn't find it, which might mean it isn't there, or might mean our analysis couldn't reach it, or that the data wasn't available. Verified absent means we specifically checked and confirmed its nonexistence.

The first is silence. The second is evidence. They aren't interchangeable, and we stopped writing as though they were.

Constraining what AI is allowed to produce

The platform uses AI to help interpret findings and explain what the evidence shows. We felt the temptation to let the AI fill interpretive gaps with plausible inferences, to produce conclusions that felt complete even where the data wasn't.

We introduced a constraint: the AI is permitted to explain what evidence exists. It is not permitted to manufacture observations the platform didn't make. If the data is incomplete, the AI must say so.

II  ·  What they had in common

We spent some time treating these as four separate decisions, each with its own justification. The word change was about accuracy. The UNKNOWN preservation was about precision. The language distinction was about rigor. The AI constraint was about reliability.

It took longer than it should have to notice that they all had the same shape.

Each one was a moment where we could have given users more certainty than our evidence justified, and chose not to. Each one required refusing a simpler, more marketable outcome in favor of a more honest one. Individually they looked like engineering choices. Together they weren't. They were the same choice, made repeatedly, in different contexts.

III  ·  What that decision is

We refuse to give you certainty we don't have.

Eventually we realised they weren't four decisions. They were one.

  • Coverage states exist because we refuse to give you certainty we don't have.
  • UNKNOWN is preserved because we refuse to give you certainty we don't have.
  • “Not observed” is kept distinct from “verified absent” because we refuse to give you certainty we don't have.
  • The AI explains evidence rather than inventing it because we refuse to give you certainty we don't have.

IV  ·  What this means

The refusal is a constraint, not a feature. It applies to everything the platform produces.

It means we publish coverage states even when they reveal the limits of our analysis, especially then. It means we preserve distinctions that make the platform harder to understand than a single score. It means we hold new capabilities to the same standard: before something ships, the question is not whether we can build it, but what evidence would make its conclusions trustworthy.

It also means accepting a specific kind of cost. Users sometimes want a cleaner answer than the evidence supports. The platform sometimes appears less confident than competitors who are willing to fill uncertainty with assertion. We've accepted that friction. A score you cannot inspect is just an opinion. We're not willing to give you opinions we can't show the reasoning behind.

It changes the relationship between us. Our job isn't to convince you that we're right. Our job is to show you enough that you can judge whether we are.


V  ·  The standard

This document is not a promise about the future. It's an explanation of decisions we've already made.

If you've used XemaS for a while, you may have already recognised the pattern without knowing its name. The coverage states that admitted partial analysis, the UNKNOWN preserved where a low score would have been more reassuring, the AI outputs that stopped at the boundary of the evidence. Now you know why.

But a description of past decisions is not a guarantee of future ones. That trust has to be earned, not declared. So we'll offer the only accountability that's consistent with how we think about evidence.

If you find us giving certainty we don't have - hiding what we couldn't assess, collapsing unknowns into false confidence, letting AI invent observations it couldn't have made - don't measure us against a marketing promise. Measure us against this document.

We wrote it as an explanation of what we've already done.
That's the only kind of explanation worth giving.

Everything else the platform does should make this document feel increasingly obvious.