Why Long-Running AI Work Needs a Different UX Than Chat
Long-running AI work needs a different user experience than chat because chat is built on four assumptions that stop being true once work lasts longer than a few minutes: that the work takes seconds, that you are watching, that the result is a single reply, and that you can judge that reply on the spot. AI agents that research, code, operate tools or carry out multi-step tasks now routinely run for minutes or hours, often while the person who started them does something else. That work needs a job with its own identity and state, progress shown as phases rather than a spinner, decision points that reach you at the right moment with the context to answer, pause, resume and recovery that never repeat an action twice, and a result delivered with evidence for review — not a confident final message.
Long-running AI work needs a different user experience than chat because chat is built on four assumptions that stop being true once work lasts longer than a few minutes: that the work takes seconds, that you are watching, that the result is a single reply, and that you can judge that reply on the spot. AI agents that research, code, operate tools or carry out multi-step tasks now routinely run for minutes or hours, often while the person who started them does something else. That work needs a job with its own identity and state, progress shown as phases rather than a spinner, decision points that reach you at the right moment with the context to answer, pause, resume and recovery that never repeat an action twice, and a result delivered with evidence for review — not a confident final message.
Key takeaways
- Chat's assumptions break with time. Seconds, attention, a single reply and on-the-spot judgment do not hold for work that runs for an hour.
- Work needs an identity. A long task should be a job you can leave and return to, not a message you must keep open.
- Progress means phases and next steps, not percentages. People need to know what is happening and what comes next.
- The agent should wait for you safely. Decisions arrive with context; nothing consequential proceeds while you are away unless you allowed it.
- "Done" is a review, not a reply. Results should arrive with evidence and with anything still unknown clearly marked.
Why does chat work so well for short tasks?
Chat works so well for short tasks because it matches how people already talk. You ask, it answers, you react. Feedback is immediate; correction is a sentence away; the whole interaction fits in your attention span. Usability research has long distinguished delays that feel instantaneous, delays people notice but tolerate without losing their train of thought, and delays of around ten seconds or more, after which attention drifts and people want progress indicators and the option to do something else (Nielsen, 1993). Most chat replies fall in the first two categories, so the interface can get away with a spinner and a reply.
Agent work has moved far past those limits. Research tracking the length of software tasks that AI systems can complete has found that length growing quickly (Kwa et al., 2025). A task that takes an agent forty minutes is not a slow reply. It is a different kind of thing.
What breaks when AI work takes minutes or hours?
When AI work outlasts a person's attention, every assumption behind chat breaks in a specific way.
The spinner stops meaning anything. A spinner that runs for twenty minutes tells you nothing except that something might still be happening. You cannot tell whether the agent is making progress, stuck, waiting for something or quietly failing.
Closing the tab becomes a risk. If the work lives in the conversation, closing the window, losing a connection or switching devices can lose it. People end up babysitting browser tabs.
Questions arrive at the wrong time. An agent that needs a decision asks in the thread, where it sits unseen while you are in a meeting. When you return, the agent has either waited without telling anyone or — worse — guessed.
Failure becomes ambiguous. In chat, a failed reply is obvious and you ask again. In long work, a step can fail halfway through after some actions already happened. "Ask again" might repeat those actions.
The final message is a story. After an hour of work, the agent writes a summary. The summary is the least reliable evidence of what happened, but it is the only thing chat shows you.
What does long-running work need instead?
Long-running work needs seven things that chat does not provide. None of them are exotic; most come straight from how reliable systems have handled long-lived work for decades. What is new is applying them to AI agents and presenting them clearly to people.
1. A job with its own identity
The first change is structural: long work becomes a job — an object with an identity, an owner and a state — rather than a reply in a conversation. You can close the window, switch devices or hand the job to a colleague, and come back to the same job. Durable execution systems in software engineering are built around exactly this idea: a workflow that keeps running, and can be resumed from where it stopped, even if the infrastructure under it fails (Garcia-Molina and Salem, 1987). Ethen's own durable job service is described in Inside Ethen's Durable Job Service.
2. Immediate acknowledgement
When you start long work, the system should confirm immediately that the job exists, what it understood the goal to be, and how you will be told about progress. A plan, even a short one, is the cheapest place to catch a misunderstanding — before an hour of work is spent on the wrong thing.
3. Progress as phases and next steps
A percentage bar implies a precision that agent work rarely has; the agent itself often does not know how many steps remain. More honest and more useful is progress expressed as phases ("gathering sources", "drafting", "checking figures"), the current step, what has been completed, and what comes next. A timeline of what happened is more informative than a number that jumps from 40% to 95%.
4. Decision points that reach you
Some steps need a person: an approval for a consequential action, a choice between two interpretations, a missing piece of information. Long-running work needs those decisions to reach you where you are — not buried in a thread — with enough context to answer quickly, and with the job waiting safely until you do. Waiting safely means the job does not proceed past the decision point, and does not time out into a default you never chose. We discuss where approvals belong in Why Ethen Keeps Human Approval in the Loop.
5. Pause, resume and recovery
Long work will be interrupted. You will want to pause it; infrastructure will restart; a provider will be unavailable. The job should pause cleanly and resume from its last good point without repeating actions that already happened. That last condition is the hard one. If an action's outcome is uncertain — a request was sent but no confirmation came back — the job should find out what happened before trying again, rather than risk doing it twice. We cover this in What Happens When an AI Task Fails Halfway Through?.
6. Cancel that means something
Cancelling a long job is not the same as closing a chat. Some actions may already have happened and cannot be undone; some may be in flight. A good cancel stops new actions immediately, reports what was already done and what was in progress, and leaves a clear record. "Cancelled" should never mean "we are not sure what happened".
7. A result for review, with evidence
The end of long work should be a review, not a reply. The result arrives with the evidence behind it — sources for claims, records of actions taken, checks that ran — and with anything that remains unknown or unfinished clearly marked. Treating outcomes as states rather than as a single success flag matters here: a result can be complete, partial, pending confirmation, or unknown, and each needs a different response from the person reviewing it. Ethen Research Lab's note From AI Traces to Verified Experience develops this idea of outcomes as versioned states; it is a research synthesis.
A worked example: the same task, two interfaces
Illustrative example — hypothetical.
An operations manager asks an AI agent to audit the company's software subscriptions: find every active subscription across three billing systems, match each to an owner, flag unused seats and draft cancellation requests for anything unused for ninety days. The task takes about fifty minutes.
In a chat interface, the manager types the request and watches a spinner. After ten minutes they switch to email. Twenty minutes in, the agent asks in the thread whether it may access the third billing system; nobody sees the question. At some point the browser tab is closed by accident. When the manager reopens the conversation, they find a partial list, an unanswered question, and no way to tell whether the agent stopped, failed or is still running somewhere. They start again.
In a job-based interface, the request becomes a job with a short plan: three systems, owner matching, an unused-seat rule, drafts but no sending. The manager approves the plan and leaves. Twenty minutes in, a notification arrives: "Access to the third billing system is needed to continue — approve read-only access?" The manager approves from their phone. The job continues from where it waited. A worker restart midway is invisible except as a short pause in the timeline. At the end, the manager receives a review: forty-two subscriptions found, thirty-eight matched to owners, four unmatched and listed for follow-up, nine flagged as unused with the evidence for each, and nine cancellation drafts — not sent. One billing system returned an error for a single page of results, and the review says so instead of quietly omitting it.
The model doing the work could be identical in both versions. The difference is entirely in how the work is held and shown.
What are the anti-patterns to avoid?
Five patterns show up repeatedly in products that bolt long-running agents onto a chat interface.
- The endless spinner. Twenty minutes of animation with no phases, no current step and no way to tell stuck from busy.
- The buried question. An agent asks for permission in a thread and then either waits silently or proceeds on a guess.
- The tab-bound job. Work that stops, or is lost, when the browser closes.
- The blind retry. A failure halfway through is handled by starting again from the top, repeating actions that already happened.
- The victory message. The only output is a confident summary of what the agent says it did.
Each one is a symptom of the same underlying mistake: treating work as a conversation because the conversation was the interface that already existed.
Does this mean chat goes away?
No. Chat remains the best way to start and steer work in natural language. In a good design, you describe the work in conversation, the system turns it into a job, and you talk to the job when you want to adjust it — "skip the third vendor", "use last quarter's figures instead". The difference is that the job, not the conversation, holds the work. We describe the general idea of a workspace that holds the work in What Is an AI Workspace?.
How do you know if a product handles long work well?
Five questions are enough to evaluate any AI product that runs long tasks:
- Can I close the window and come back to the same job?
- Can I see what phase it is in and what comes next?
- If it needs me, will I find out — and will it wait safely?
- If it fails halfway, will it resume without repeating what already happened?
- When it finishes, can I see the evidence, not just the summary?
A product that answers yes to all five is treating long work as work. Our broader checklist for verifiable agent work is in What Makes an AI Agent Job Verifiable.
How is Ethen applying this?
Ethen is building long-running work on durable jobs, explicit states including unknown outcomes, approvals bound to specific actions, and completion that depends on evidence — mechanisms described in our engineering posts with their stated test limits. How those foundations turn into the experience people see — status, notifications, pause and resume, review — is the subject of a companion article, Designing Ethen for Work That Takes Minutes or Hours. The direction for delegated work in particular is described in What Ethen Work Is Meant to Become.
Limitations
The patterns here are general design guidance, not measured findings. Different kinds of long work weight them differently: a nightly data job needs little interaction, while a research task may need several decisions. Some patterns add friction for tasks that turn out to be short, so good products adapt the presentation to the actual duration of the work.
Frequently asked questions
Why is a progress percentage misleading for AI agents? Because agents rarely know in advance how many steps a task will take. Phases, the current step and what comes next are more honest and more useful.
What should happen if an AI agent needs my approval while I'm away? The job should notify you with the context to decide, and wait safely — not proceed on a guess and not time out into a default you never chose.
Can long-running AI work still be steered with chat? Yes. Conversation is a good way to adjust a job. The job, not the conversation, should hold the work and its state.
What does "resume without repeating effects" mean? That after an interruption, the job continues from its last good point, and checks whether an uncertain action already happened before trying it again.
Related reading
- Designing Ethen for Work That Takes Minutes or Hours
- AI Agents and Autonomous Work
- What "Done" Should Mean for an AI Agent
References
- Nielsen, J. (1993). Response Times: The 3 Important Limits. Nielsen Norman Group. https://www.nngroup.com/articles/response-times-3-important-limits/
- Kwa, T. et al. (2025). Measuring AI Ability to Complete Long Software Tasks. arXiv:2503.14499. https://arxiv.org/abs/2503.14499
- Garcia-Molina, H., & Salem, K. (1987). Sagas. ACM SIGMOD Record 16(3). https://doi.org/10.1145/38713.38742
- Ethen Blog (2026). Inside Ethen's Durable Job Service. https://upcube.ai/blog/inside-ethens-durable-job-service
- Ethen Blog (2026). What Makes an AI Agent Job Verifiable. https://upcube.ai/blog/what-makes-an-ai-agent-job-verifiable
- Ethen Research Lab (2026). From AI Traces to Verified Experience. Research note; research synthesis. https://upcube.ai/resources/research/verified-experience