Why Ethen Is Treating AI Safety as a Product Experience
Good AI safety user experience starts from a simple observation: most of what keeps an AI product safe in daily use is something the user can see or do. They see what the AI is allowed to touch in a task. They see a preview of a consequential action and approve that exact action. They watch progress and can stop or take over. They are told honestly when an outcome is unknown rather than given a guess. And afterward they can see a record of what happened and recover from mistakes. Safety measures that users cannot see or use tend to be ignored, misunderstood or worked around. So Ethen treats safety as part of the product experience, designed with the same care as any feature, while keeping some protections — such as credential handling — deliberately invisible. This article explains that approach and where it shows up in Ethen today.
Good AI safety user experience starts from a simple observation: most of what keeps an AI product safe in daily use is something the user can see or do. They see what the AI is allowed to touch in a task. They see a preview of a consequential action and approve that exact action. They watch progress and can stop or take over. They are told honestly when an outcome is unknown rather than given a guess. And afterward they can see a record of what happened and recover from mistakes. Safety measures that users cannot see or use tend to be ignored, misunderstood or worked around. So Ethen treats safety as part of the product experience, designed with the same care as any feature, while keeping some protections — such as credential handling — deliberately invisible. This article explains that approach and where it shows up in Ethen today.
Key takeaways
- Safety is experienced, not just enforced. Users meet it as scopes, previews, approvals, stop controls and records.
- Annoying safety gets bypassed. Friction must be spent on decisions that matter.
- Some guardrails should be invisible. Users should never have to think about where credentials live.
- Anything that needs judgment should be visible. Consequential actions, refusals, uncertainty and data use.
- High automation and high control can coexist. The goal is not less AI, but AI people can steer.
What does AI safety look like inside a product?
AI safety research covers a wide range: model behavior, misuse, robustness, long-term risks. Much of that work happens far from any user interface. But for people using an AI product to do their work, safety is mostly experienced in a handful of moments. Figure 1 shows five of them.
Before the task. Does the user know what the AI can see and do? Which files, which accounts, which tools, which budget?
At the decision. When the AI is about to do something consequential, does the user see exactly what will happen and approve that specific action?
During the task. Can the user see progress, stop the AI, or take over?
When unsure. When the AI cannot tell whether something worked, does it say so, or does it guess — or quietly try again?
After the task. Is there a record of what was done and approved, and a way to undo or recover?
Each of these is a design problem as much as a safety problem. A safety mechanism that exists in the system but is invisible or confusing at these moments does far less than it could.
Why safety that users cannot see falls short
There are three common failure modes when safety is treated purely as back-end enforcement.
Users cannot calibrate trust. If people cannot see what an AI is allowed to do and what it actually did, they cannot judge how much to rely on it. They either over-trust and stop checking, or under-trust and stop using it.
Users work around friction they do not understand. A safety step that feels arbitrary — a confirmation that appears for trivial actions, a refusal with no explanation — trains people to click through or route around it. Guardrails that are too strict make a product feel paternalistic; guardrails that are inconsistent make it unpredictable. Either way, users lose the safety benefit.
Errors are discovered too late. If the user's first view of a consequential action is its result, the chance to catch a mistake has passed.
Ben Shneiderman, writing about human-centered AI, argued that designers have treated automation and human control as a single trade-off — more of one means less of the other — and that this is a mistake. The goal should often be high automation and high human control at the same time: systems that do a lot, which people can still understand and steer. Treating safety as a product experience is one way to pursue that.
Visible or invisible?
Not every protection should be visible. Users should never need to think about some things, and showing them would only add noise. Figure 2 sorts common guardrails.
Usually invisible. Abuse filtering, rate limiting and credential handling. For example, when Ethen Chat starts a voice session, the server checks sign-in, rate limits and configuration, and the browser receives only a short-lived pass; long-lived keys never reach it. Users benefit without having to know. We describe that mechanism in How Ethen Chat Starts a Voice Session.
Usually visible. Consequential actions, refusals, uncertain outcomes and data use. These either change what the user can rely on or need the user's judgment, so hiding them removes the user's ability to steer.
The rule of thumb: if a protection changes what the user can rely on, or needs their decision, make it visible. If it does neither, keep it out of the way.
The consequential action, designed for safety
The clearest example of safety as experience is how a consequential action is handled — sending a message, submitting a form, spending money, deleting something. Figure 3 shows the pattern.
Preview. Show exactly what will happen: the recipient, the wording, the amount, the file.
Approve. One decision for one action. Ethen's computer-use approvals are bound to the exact action, run and attempt they were requested for, and can be used once; an approval for one action cannot authorize a different one. See Binding Computer-Use Approvals to Specific Actions.
Act once. No silent repeats. If something goes wrong in transit, the system should not quietly try again.
Confirm. Check and show the result.
Record. Keep who approved what and when.
We explain why Ethen keeps people in the loop for these decisions in Why Ethen Keeps Human Approval in the Loop, and how we think about actions that cannot be undone in How Ethen Thinks About AI Actions That Can't Be Undone.
A worked example
The following example is illustrative. An operations lead asks an AI agent to update shipping addresses for twelve customers from a spreadsheet and email each customer a confirmation.
Before. The task screen shows the scope: read access to one spreadsheet, write access to the customer records, permission to draft emails, and no permission to send without approval.
At the decision. The agent updates the records — a reversible change within scope — and prepares twelve emails. Rather than sending them, it shows a preview: the recipients, the message and the new address in each. The lead notices that two addresses in the spreadsheet look truncated, edits them, and approves the remaining ten as a batch of exact messages.
During. Partway through sending, the email service stops responding.
When unsure. The agent does not retry. It reports that seven emails were confirmed sent, two failed, and one has an unknown outcome because the service timed out after the request.
After. The record shows every address change, every approval and every send status. The lead checks the outbox, confirms the uncertain email did go out, and sends the two failed ones.
Nothing about this example is dramatic, and that is the point. Every moment where something could have gone wrong — a bad address, a duplicate email, a silent failure — was visible to a person at the time it mattered.
Anti-patterns to avoid
A few patterns undermine safety even when the underlying controls are sound.
Approve-all buttons. A single "approve everything" control turns a sequence of decisions into one blind decision.
Confirmations without content. "Are you sure?" with no preview of what will happen tests patience, not judgment.
Silent retries. Retrying a consequential action after an unclear failure can double its effect.
Narration instead of evidence. An agent saying "I have sent the email" is not the same as showing the sent message.
Unexplained refusals. A refusal with no reason or alternative teaches users to rephrase until something gets through.
Where this shows up in Ethen today
Several published Ethen mechanisms reflect this approach. Each is described in its own engineering post, with its limits stated.
Precise approvals for computer use. Approvals bind to one exact action and are single-use. The same post is explicit that this does not make an agent resistant to manipulation by web content; precise approvals make each decision clear, not each proposal trustworthy.
Voice sessions start with no tools. A new voice session in Ethen Chat begins with no tool grants, and microphone access is allowed only on chat pages. Capabilities must be granted deliberately.
Unknown outcomes stay unknown. When an agent action's result cannot be determined, Ethen's mission design keeps the effect recorded as unknown until evidence resolves it, rather than guessing. See When an Agent Action's Outcome Is Unknown.
Local model changes need approval. In Ethen Desktop, changes to local models — installing or removing them — go through approval in the app's privileged process, as described in How Ethen Desktop Talks to Local Models.
History stays with its owner. Ethen Code scopes local run history to the signed-in account and clears it on sign-out, as described in Keeping Code History Scoped to the Signed-In User.
Model facts show unknowns. Model information that lacks complete provenance shows "Unknown" rather than a guessed value.
None of these is a claim that Ethen is "safe" in general. Each is a specific mechanism with a specific scope.
Design principles we apply
Spend friction on what matters. Confirmations for trivial actions teach people to click through confirmations. Routine, reversible steps should flow; consequential ones should stop.
Make refusals useful. When Ethen declines to do something, it should say what it will not do and, briefly, why — and, where possible, what the user can do instead.
Show scope before work starts. What the AI can access in a task should be visible up front, not discovered afterward.
Prefer honest uncertainty to confident error. "I could not confirm this" is a safety feature.
Keep the stop button real. Stopping should take effect immediately, not after the current batch of actions.
Make records readable. A log that only engineers can read does not help a user understand what happened.
These principles echo widely cited guidelines for human-AI interaction — making clear what a system can do, supporting efficient correction and dismissal, and explaining behavior — and practices proposed for governing agentic AI systems, such as constraining the action space, keeping systems interruptible and making their actions legible.
Safety experience and research
Some of these ideas are still research questions. Ethen Research Lab's Mandates research note explores how a user's intent could be turned into bounded authority, so approvals can focus on what falls outside it. Unknown Effects in Autonomous AI Systems, a research note, argues why uncertain outcomes should be reconciled before retrying. And the Lab's one published system card, AgentTrustBench, reports how one pinned build behaved against its own stated autonomy boundaries — on that build and that set of conditions only.
How we judge whether safety UX works
The test of safety as experience is behavioral, not cosmetic. The questions we want to be able to answer are practical. Do users read previews before approving, or approve reflexively? Do they use the stop control when something looks wrong? Do they understand what an "unknown" status means and act on it? Do refusals lead to a sensible next step, or to frustration? Are there actions users repeatedly approve without changes — a sign they may not need approval at all? We are not publishing measurements of these here; they describe how we intend to evaluate the experience, not results.
Tradeoffs and limitations
Visible safety adds friction. Previews and approvals take time. The answer is to make them rare and meaningful, not to remove them.
Too many prompts cause fatigue. If everything needs approval, nothing gets real attention. Deciding which actions are consequential is an ongoing design problem.
Visibility is not protection by itself. Showing a user an action does not guarantee they read it carefully.
Some risks are outside the interface. Model-level risks, adversarial content and infrastructure security need protections users never see. Product experience complements them; it does not replace them.
FAQ
How do you design safe AI products? Make scope visible before a task, preview and precisely approve consequential actions, let users stop or take over, show uncertainty honestly, keep readable records, and keep protections such as credential handling out of the user's way.
Should AI guardrails be visible to users? Those that change what users can rely on or need their judgment — consequential actions, refusals, uncertainty and data use — should be visible. Purely protective ones, like abuse filtering and credential handling, usually should not.
What does AI safety look like in an app? Mostly in five moments: knowing the scope, approving consequential actions, watching and stopping work, seeing honest unknowns, and reviewing what happened afterward.
Does more human control mean less automation? Not necessarily. Human-centered AI research argues for designing systems with both high automation and high human control.
Is Ethen safe? We do not make blanket safety claims. Ethen publishes specific mechanisms, each with its scope and limits stated.
Related reading
- Why Ethen Keeps Human Approval in the Loop
- What We Mean by Responsible Autonomy
- How Ethen Thinks About AI Actions That Can't Be Undone
- Binding Computer-Use Approvals to Specific Actions
- When an Agent Action's Outcome Is Unknown
References
- Shneiderman, B. (2020). Human-Centered Artificial Intelligence: Reliable, Safe & Trustworthy. International Journal of Human–Computer Interaction, 36(6), 495–504. https://doi.org/10.1080/10447318.2020.1741118
- Amershi, S., et al. (2019). Guidelines for Human-AI Interaction. Proceedings of CHI 2019. https://doi.org/10.1145/3290605.3300233
- Shavit, Y., Agarwal, S., Brundage, M., et al. (2023). Practices for Governing Agentic AI Systems. OpenAI. https://openai.com/index/practices-for-governing-agentic-ai-systems/
- Ethen Research Lab (2026). Mandates: Compiling Human Intent Into Bounded Agent Authority. Research note. https://upcube.ai/resources/research/agent-mandates
- Ethen Research Lab (2026). AgentTrustBench system card. https://upcube.ai/resources/research/agent-trust-boundaries