Why Ethen Shows What It Knows—and What It Doesn't
AI products have a built-in honesty problem: their output sounds equally confident whether it is right, wrong, estimated or invented. Ethen treats that as a design problem, not only a model problem. Across the product, we try to keep four states of knowledge apart — known, estimated, unknown and not checked — and to show each for what it is. When a model fact lacks provenance, Ethen Model Intelligence shows Unknown rather than a guess. When an agent's action times out without confirmation, Ethen's mission system records the outcome as unknown rather than as success or failure. When Ethen Research Lab publishes a proposal, it says the proposal is untested. We do this because people rely on AI outputs more than the outputs deserve when uncertainty is hidden, and because a gap shown honestly is more useful than a gap filled with something plausible.
AI products have a built-in honesty problem: their output sounds equally confident whether it is right, wrong, estimated or invented. Ethen treats that as a design problem, not only a model problem. Across the product, we try to keep four states of knowledge apart — known, estimated, unknown and not checked — and to show each for what it is. When a model fact lacks provenance, Ethen Model Intelligence shows Unknown rather than a guess. When an agent's action times out without confirmation, Ethen's mission system records the outcome as unknown rather than as success or failure. When Ethen Research Lab publishes a proposal, it says the proposal is untested. We do this because people rely on AI outputs more than the outputs deserve when uncertainty is hidden, and because a gap shown honestly is more useful than a gap filled with something plausible.
Key takeaways
- Fluency is not confidence. Well-written output says nothing about whether it is correct.
- Four states, kept apart. Known, estimated, unknown and not checked are different, and collapsing them misleads.
- Show the gap, with a reason. Unknown is a legitimate answer when it says why and what would resolve it.
- Never fill a gap with a plausible guess. Interpolated, carried-forward or borrowed values look like facts and are not.
- Fewer, better warnings. Uncertainty belongs where a decision depends on it, not as boilerplate on every line.
Why don't AI tools say when they don't know?
AI tools rarely say "I don't know" because the systems that generate their output produce fluent text by default, and fluent text sounds sure. A language model asked a question will usually produce an answer-shaped response whether or not it has a reliable basis for one. Products then compound the problem by presenting every answer, number and status in the same confident visual style.
Two kinds of research explain why this matters. First, people over-rely on automated systems: human-factors research has described complacency and automation bias — a tendency to accept automated output and to monitor it less carefully — especially when the system usually works (Parasuraman & Manzey, 2010). Second, AI systems' own signals about their reliability are weaker than they look. Research on long-running agents, for example, has found that uncertainty signals early in a long task are poor predictors of whether the task will eventually fail (Li et al., 2026). An agent cannot always notice its own mistakes in time, and a product that relies on the agent to volunteer doubt will often get none.
Human-AI interaction guidelines have recommended for years that AI products make clear not only what they can do but how well they can do it (Amershi et al., 2019). The hard part is turning that advice into specific design decisions. The rest of this article describes ours.
What are the four states of knowledge?
The four states of knowledge are the distinctions Ethen tries to keep visible wherever the product shows a fact, a result or a status.
Known means supported by evidence that can be traced: a source and a date for a model fact, an observation for an action's outcome, a check for a completed task.
Estimated means a reasoned approximation that has not been measured: a cost quoted before a job runs, an expected duration, a predicted model suitability. Estimates are useful as long as they are labeled.
Unknown means the evidence is missing or contradictory. The honest display says so, gives the reason, and — when something is being done to resolve it — says what.
Not checked means nobody has looked yet. It is different from unknown: an unchecked claim might be easy to confirm. The danger is that "not checked" is silently presented as "passed".
Why is "unknown" different from "failed"?
Unknown and failed lead to opposite actions, which is why rounding one into the other is dangerous. If a refund failed, the right next step is to try again. If a refund's outcome is unknown, trying again may pay the customer twice; the right next step is to find out what happened. If a research claim is refuted, you stop relying on it. If it is unchecked, you may only need to look it up.
The same logic applies to information. A model whose benchmark result is unknown is not a model that performed badly; it is a model nobody has measured on that test in a way that can be traced. Treating missing values as low scores penalizes models for gaps in the record. Treating them as average scores rewards the gap. Neither is honest. The only honest option is to show the gap and let the person decide how much it matters for their task — which is also why our guide to making AI model choice less confusing recommends testing on your own examples when the evidence you need is missing.
Distrust of "unknown" usually comes from systems that use it lazily, as a default for anything difficult. The fix is not to hide unknowns but to make them precise: say exactly what is unknown, why, and what would settle it.
Where does Ethen apply this today?
Ethen applies this principle in several places that are already described in our engineering and research writing.
Model facts: Unknown instead of a guess
In Ethen Model Intelligence, a benchmark value is shown as known only if it has a source URL, a valid retrieval date, a methodology statement and adequate confidence. If any of those is missing, the cell shows Unknown and keeps the reason. The system does not interpolate, carry forward an old value, or borrow a number from a similar benchmark. The engineering details are in Showing Unknowns in Ethen Model Intelligence. A table with visible gaps looks less finished than a full one. It is also the only kind that does not quietly mislead someone choosing a model.
Agent outcomes: unknown is a real state
When an agent's action is interrupted — a timeout after the request was sent, a crash before the result was recorded — Ethen's mission system records the action's effect as unknown. It does not round it to failure, which would invite a duplicate retry, or to success, which would let later steps build on something that may not exist. The unknown stays visible until an observation resolves it, and a mission cannot complete while it remains. See When an Agent Action's Outcome Is Unknown.
Completion: the agent's word is not evidence
An agent reporting that a task is done is a claim, not a known fact. Ethen's mission system requires a separate check with evidence before a task counts as succeeded. We explore what "done" should mean in What "Done" Should Mean for an AI Agent.
Research: every claim carries its evidence status
Ethen Research Lab labels every publication with its type and evidence status, and labels individual numbers inside papers as measured, external, illustrative, proposed targets, hypotheses or unknown. A research proposal never appears as a finding. We explain why in Why Ethen Keeps Research Separate From Product Claims.
Releases: scope, not slogans
Ethen's release certificates record which checks ran, against what, on which date — including checks marked partial or not performed. A passing certificate for one scope says nothing about another. That is the "not checked" state, made explicit.
How should Unknown be shown?
Unknown is only useful if it is shown well. An "Unknown" with no explanation feels like a bug; an "Unknown" with a reason and a next step feels like a trustworthy system doing its job.
Good Unknowns share four properties:
- They say why. "No retrieval date on the source", "the payment provider did not confirm", "this source could not be opened".
- They say what would resolve it. "Re-checking the provider's records", "awaiting your decision", "no further action planned".
- They do not block what does not depend on them. An unknown benchmark value should not hide the known ones next to it.
- They are rare enough to read. Uncertainty that appears on every line becomes wallpaper.
How do we decide when uncertainty is worth showing?
Not every uncertain detail deserves the reader's attention. We use four questions to decide whether to surface uncertainty prominently, mention it quietly, or leave it in the record for anyone who looks.
Does a decision depend on it? If a person is about to approve, spend, send or rely on something, uncertainty about that thing goes in front of them. If nothing depends on it, it can stay in the details.
Is it reversible? Uncertainty about a reversible step matters less than uncertainty about an irreversible one. A draft can be fixed later; a payment cannot easily be recalled. We discuss reversibility in How Ethen Thinks About AI Actions That Can't Be Undone.
Can the person do anything about it? Uncertainty that comes with an action — check this source, confirm this figure, answer this question — is useful. Uncertainty with no available action is often better recorded than displayed.
Who is looking? An engineer debugging a run wants every unknown in the log. A manager reading a summary wants the two that affect the conclusion. The same facts can be shown at different levels of detail for different readers, as long as the summary never claims more certainty than the detail supports.
Why not just add disclaimers?
Blanket disclaimers fail for the same reason over-reliance happens: people learn to ignore signals that are always present. A banner saying "AI can make mistakes" on every screen is true and useless; it gives no information about which output deserves a second look. The design goal is targeted uncertainty: show it where a decision depends on it, attach it to the specific value or step it concerns, and keep it out of places where it changes nothing.
This is also why we prefer structural signals to hedged language. A cell that says Unknown is clearer than a paragraph full of "may", "might" and "possibly". A status that says "awaiting confirmation from the payment provider" is clearer than "the refund was probably issued".
What about the answers Ethen writes?
Generated answers are the hardest place to apply this principle, because the uncertainty lives inside fluent prose. Our direction for research-style answers in Ethen Chat is to make it easy to see which source supports which claim, to show where sources disagree, and to mark figures that could not be verified — while keeping ordinary conversation uncluttered. Research on evaluation offers a useful pattern: systems that can abstain when uncertain, and escalate to a stronger check, can be more trustworthy than systems that always answer (Jung et al., 2024). For a conversational product, abstaining sometimes looks like "I couldn't confirm that figure; here is where I looked".
Ethen Research Lab's methods paper Evaluating the Evaluators: Reward Integrity for AI Agents applies the same idea to the systems that judge AI work: measure how often a checker accepts bad work or rejects good work, measure its calibration, and let it decline to judge rather than guess. It is a research synthesis; it reports no measurements of Ethen's own verifiers.
An example
Illustrative example — describes the intended behavior, not a specific shipped screen.
A procurement analyst asks Ethen to compare three cloud storage vendors' prices and contract terms. The result is a table. Two vendors' prices are shown with links to their pricing pages and the date they were retrieved. The third vendor's price cell says Unknown — pricing page requires a sales contact; no public figure found. One contract term is marked Estimated from a 2025 sample contract; confirm with vendor. A note at the bottom says Not checked: data-residency options for vendor C. The analyst knows exactly which two cells to chase before the comparison goes to their manager. A fully filled-in table would have looked better and been worse.
Tradeoffs
Showing uncertainty has costs. Products with visible gaps can look less capable than products that fill them, especially in side-by-side demos. Some people want an answer, not a caveat, and will be frustrated by Unknowns when a decision has to be made anyway. Too much uncertainty display turns into noise. And tracking provenance and status takes engineering work that a confident-looking product can skip. We accept these costs because the alternative — hiding uncertainty until it surfaces as a mistake someone relied on — is worse for the people who use Ethen.
Frequently asked questions
Why does Ethen show "Unknown" instead of a number? Because the number would imply evidence the system does not have — for example, a benchmark value without a source, date or method. Unknown, with its reason, is more useful than a guess.
Does showing uncertainty mean Ethen is less accurate? No. It means Ethen is showing where its accuracy has not been established. Accuracy and honesty about accuracy are different properties.
Can AI models reliably report their own uncertainty? Not reliably on their own. That is why Ethen relies on structural evidence — sources, observations, independent checks — rather than only on a model's stated confidence.
Will every Ethen answer include confidence levels? No. Uncertainty is shown where a decision depends on it. Blanket confidence labels on every sentence would be ignored.
Related reading
- Showing Unknowns in Ethen Model Intelligence
- Agent Verification and Evaluation
- Evaluating the Evaluators — Ethen Research Lab methods paper (research synthesis).
References
- Amershi, S. et al. (2019). Guidelines for Human-AI Interaction. Proceedings of CHI 2019. https://www.microsoft.com/en-us/research/publication/guidelines-for-human-ai-interaction/
- Parasuraman, R., Manzey, D. H. (2010). Complacency and Bias in Human Use of Automation: An Attentional Integration. Human Factors 52(3):381–410. https://doi.org/10.1177/0018720810376055
- Li, Z. et al. (2026). Last Step Matters: Early Uncertainty Cannot Predict Failure in Long-Horizon Agents. arXiv:2608.29685. https://arxiv.org/abs/2608.29685
- Jung, J. et al. (2024). Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement. arXiv:2407.18370. https://arxiv.org/abs/2407.18370
- Ethen Blog (2026). Showing Unknowns in Ethen Model Intelligence. https://upcube.ai/blog/showing-unknowns-in-ethen-model-intelligence
- Ethen Blog (2026). When an Agent Action's Outcome Is Unknown. https://upcube.ai/blog/when-an-agent-actions-outcome-is-unknown
- Ethen Research Lab (2026). Evaluating the Evaluators: Reward Integrity for AI Agents. Methods paper; research synthesis. https://upcube.ai/resources/research/reward-integrity