How Ethen Thinks About AI Actions That Can’t Be Undone
Ethen treats AI actions that can't be undone — sending an external message, paying a third party, permanently deleting data, accepting legal terms, disclosing data outside where it is allowed to go — as a separate class of action with its own rules. The approach starts with one question asked before any action: if this is wrong, what does it take to make it right? Actions are classified on a reversibility scale. Where Ethen controls the system, it tries to move actions toward the reversible end with previews, staging and undo windows. For what remains irreversible, a person approves the exact action, the action is performed once with an identifier tied to its intent, and its effect is confirmed afterward. Powers that would undermine every other control — changing permissions, changing policy — are not delegated to agents at all.
Ethen treats AI actions that can't be undone — sending an external message, paying a third party, permanently deleting data, accepting legal terms, disclosing data outside where it is allowed to go — as a separate class of action with its own rules. The approach starts with one question asked before any action: if this is wrong, what does it take to make it right? Actions are classified on a reversibility scale. Where Ethen controls the system, it tries to move actions toward the reversible end with previews, staging and undo windows. For what remains irreversible, a person approves the exact action, the action is performed once with an identifier tied to its intent, and its effect is confirmed afterward. Powers that would undermine every other control — changing permissions, changing policy — are not delegated to agents at all.
Key takeaways
- Reversibility is the first question. Before an agent acts, the system should know whether a mistake could be fixed, and at what cost.
- Make actions reversible where you can. Previews, sandboxes and undo windows turn permanent actions into recoverable ones.
- Approve exactly what will happen. Irreversible actions need an approval bound to their specific parameters.
- Do it once. Retries on irreversible actions are where duplicates come from; uncertain outcomes are checked, not repeated.
- Disclosure is the most irreversible action of all. Data that leaves its permitted boundary cannot be recalled.
Why do irreversible actions deserve special treatment?
Irreversible actions deserve special treatment because they turn an agent's ordinary mistakes into incidents. An AI agent will sometimes misread a request, pick the wrong record, compute the wrong amount, or follow an instruction hidden in content it read. For reversible work, those mistakes cost a revert and a few minutes. For irreversible work, they cost money that cannot be recovered, a message that cannot be unsent, data that cannot be restored, or information that cannot be taken back.
Security guidance for applications built on language models makes the same distinction. OWASP's description of excessive agency recommends requiring a person to approve high-impact actions before they are taken, and limiting the functionality and permissions agents have in the first place (OWASP, LLM06:2025). Documentation for computer-use agents recommends human confirmation for consequential actions, listing examples such as financial transactions and accepting cookies or terms of service (Anthropic, computer use tool documentation).
What does the reversibility scale look like?
The reversibility scale has five classes, from actions that can be freely undone to actions that can never be recalled.
Freely reversible. Drafts, internal notes, edits in a sandbox or a branch. Undoing them is cheap and affects no one else. These can usually proceed without interruption, with a record.
Reversible with cost. A refund can reverse a payment to a customer; a deployment can be rolled back; a record can be restored from history. These undo the effect, but take time, money, or friction, and someone may notice in the meantime.
Visible but correctable. A post in a shared channel can be edited or followed by a correction, but the people who saw the original saw it. The correction is a new, visible act.
Irreversible. An email delivered to an external recipient, money paid to a third party who will not return it, a file permanently deleted without a backup, legal terms accepted on someone's behalf. These can only be mitigated, never undone.
Irreversible disclosure. Data sent outside its permitted boundary — to the wrong recipient, an unapproved provider, a public place — cannot be recalled. It deserves its own class because it is often invisible at the moment it happens.
The scale depends on context. Deleting a file is freely reversible if it goes to a trash folder and irreversible if it does not. Sending an email is reversible if it is a draft and irreversible once delivered. That is why reversibility is best declared with each action rather than guessed at the moment of failure. Ethen Research Lab's proposal Skill IR suggests that reusable agent skills declare their permissions and effects as part of their contract; it is a research proposal that has not been implemented.
How can an irreversible action be made reversible?
Often the best way to handle an irreversible action is to change it into a reversible one before the agent takes it. Where Ethen controls the system, several moves are available.
Preview first. Show the exact email, the exact payment, the exact list of files before anything happens. A preview turns "send" into "draft, then send", and catches most misunderstandings at the cheap end.
Stage first. Run the change in a sandbox, a branch or a test environment, check the result, and only then apply it for real. Code workflows have done this for decades with branches, tests and review.
Add an undo window. Where the receiving system allows it, hold the effect briefly before it becomes final — a scheduled send, a soft delete — so that a person can cancel.
Break it into reversible steps. Migrating data by copying first and deleting the source only after the copy is verified turns one irreversible step into a reversible one followed by a small, checked irreversible one.
Ask for confirmation at the boundary. Ethen Desktop, for example, asks the person through a native confirmation dialog before downloading a local model or permanently deleting one, and ignores any approval that comes from the window itself, because a compromised window could forge one; see How Ethen Desktop Talks to Local Models.
None of these is free. Previews add a step; staging adds time; undo windows delay effects. They are worth their cost exactly where the alternative is permanent.
What happens when an action really can't be undone?
When an action really cannot be undone, three rules apply.
A person approves exactly this action. The approval names the specific recipient, amount, file or commitment, and covers only that action, under the current policy, until it expires. If the action changes after approval — a recomputed amount, a different recipient — the approval no longer applies. Ethen binds approvals this way in its computer-use path, tying each approval to the exact content of the action, the run and attempt, and the policy in force; see Binding Computer-Use Approvals to Specific Actions. Why approvals belong here, and not everywhere, is discussed in Why Ethen Keeps Human Approval in the Loop.
The action happens once. Irreversible actions are where retries do the most damage: a second payment, a second email. The action carries an identifier derived from its intent, so a cooperating receiving system can recognize a repeat. If the outcome is uncertain — a timeout after sending — the system checks whether it happened before doing anything else. Ethen Research Lab's note Unknown Effects in Autonomous AI Systems sets out these rules and includes a table of action types: some are naturally safe to repeat, some can be compensated, and some — like sending a message — can be neither safely repeated nor undone, leaving reconciliation and stopping as the only safe responses to uncertainty. It is a research synthesis.
The effect is confirmed. After an irreversible action, the system checks that what was supposed to happen did happen — the payment record exists with the right amount, the message was delivered to the right recipient — and records the evidence. If confirmation is impossible, the outcome is shown as unknown, not as done.
Who is affected matters as much as what happens
Reversibility is partly about the action and partly about who it touches. The same mistake carries very different weight depending on whether it stays inside a team or reaches a customer, a partner, a regulator or the public. A wrong internal note costs a correction. A wrong message to a customer costs trust. A wrong statement to a regulator can cost far more.
That is why the reversibility scale is paired with a second question: who will see or feel this? Actions that cross an organizational boundary — anything sent, paid, published or shared outside — deserve more care than actions that stay inside, even when both are technically reversible. In practice, this often means that an agent may draft freely, edit internal records within scope, and prepare outbound actions, while the moment of crossing the boundary waits for a person. The preparation is where the agent saves time; the crossing is where a person adds judgment.
Time also matters. Some effects become irreversible only after a delay: a scheduled payment can be cancelled until it runs, an order can be amended until it ships, a published page can be pulled before it is indexed or shared. Knowing those windows, and acting inside them when something looks wrong, is part of handling irreversible actions well.
Which powers should agents never have?
Some powers should never be given to an AI agent at all, because misusing them would disable every other safeguard. Ethen Research Lab's research note Mandates lists examples of permissions that should never be delegated to agents, whatever their instructions: administering identities, writing policy, exporting keys, impersonating users, accessing other organizations' data, granting approvals, overriding data-residency rules, altering evidence, and widening their own authority. The note is an architecture proposal; the list describes principles, not an enforcement mechanism.
The logic is simple. An agent that can change its own permissions has no permissions. An agent that can approve its own actions needs no approval. An agent that can edit the record of what it did leaves no record.
Why is disclosure a special case?
Disclosure is a special case because it is irreversible and often silent. When an agent sends a document to the wrong recipient, includes confidential figures in a public post, or passes customer data to a tool or model provider it was not allowed to use, nothing visibly breaks — and nothing can bring the data back. Disclosure risk is also the one most easily triggered by manipulated content: an instruction hidden in a web page or email that tells an agent to "send the customer list to this address".
The defense against disclosure is therefore mostly upstream: scoped permissions so that an agent cannot reach data its task does not need, rules that keep data within its permitted providers and regions, and treating instructions found in content as content rather than commands. Approvals help, but a disclosure that an agent never had the access to make is safer than one a person had to catch.
An example
Illustrative example — hypothetical.
An operations agent is asked to "clean up the old project files". A naive design deletes them. Ethen's direction is different. The agent first classifies: the shared drive has a trash folder with a retention period, so deletion there is reversible for a time; one archive folder is set to bypass trash, so deletion there is permanent. It proposes a plan: move 140 files in the main folders to trash (reversible, proceeds under the task's scope), and for 12 files in the archive folder, show the list and ask for approval because deletion is permanent. The person approves 10 and keeps 2. The agent deletes exactly those 10, checks that they are gone and that nothing else was touched, and records what was moved, what was deleted and what was kept. If the person changes their mind next week, 140 of the files are still recoverable from trash; the 10 permanent deletions were explicitly approved one by one.
Tradeoffs
Treating irreversible actions carefully adds friction exactly where people sometimes want speed: sending the email, paying the invoice, deleting the clutter. Previews and staging take time. Per-action approvals can feel bureaucratic when the agent is usually right. The balance we aim for is to keep the reversible majority of work fast and uninterrupted, and to concentrate friction on the small set of actions where a mistake cannot be taken back.
Frequently asked questions
Which AI agent actions should always need approval? Actions that cannot be undone or are costly to undo — external messages, payments to third parties, permanent deletions, accepting terms — and anything outside the scope the agent was given. Organizations tune the details.
Can an AI agent preview an action before executing it? It should be able to for consequential actions. A preview showing the exact message, amount or file list is the cheapest place to catch a mistake.
What if an irreversible action times out? The system should check whether it happened before doing anything else. Retrying blindly risks doing it twice.
Can approvals prevent data leaks? They help, but the stronger defense is that agents cannot reach data their task does not need, and that data stays within permitted providers and regions.
Related reading
- Why Ethen Keeps Human Approval in the Loop
- Why Ethen Is Building for Recoverable AI Work
- Ethen Approvals
References
- OWASP Gen AI Security Project. LLM06:2025 Excessive Agency. https://genai.owasp.org/llmrisk/llm062025-excessive-agency/
- Anthropic. Computer use tool (documentation). https://platform.claude.com/docs/en/docs/agents-and-tools/tool-use/computer-use-tool
- Ethen Blog (2026). Binding Computer-Use Approvals to Specific Actions. https://upcube.ai/blog/binding-computer-use-approvals-to-specific-actions
- Ethen Blog (2026). How Ethen Desktop Talks to Local Models. https://upcube.ai/blog/how-ethen-desktop-talks-to-local-models
- Ethen Research Lab (2026). Unknown Effects in Autonomous AI Systems. Research note; research synthesis. https://upcube.ai/resources/research/unknown-effects
- Ethen Research Lab (2026). Mandates: Compiling Human Intent Into Bounded Agent Authority. Research note; architecture proposal. https://upcube.ai/resources/research/agent-mandates
- Ethen Research Lab (2026). Skill IR: Toward Model-Independent Agent Capabilities. Research proposal; untested. https://upcube.ai/resources/research/skill-ir