What We Want Ethen Voice to Feel Like
Voice AI latency and turn-taking decide whether talking to an assistant feels like a conversation or like operating a machine. People take turns with remarkably short gaps, and they do it by predicting when the other person will finish rather than waiting for silence. A voice assistant that waits too long feels slow; one that jumps in too early interrupts your thinking; one that keeps talking when you speak feels rude. We want Ethen Voice to feel quick but not hasty, brief by default, easy to interrupt, calm rather than performative, honest when it did not hear you, and always clear that you are talking to an AI. This article explains those principles, the research behind them, and the tensions that make good voice harder than it sounds.
Voice AI latency and turn-taking decide whether talking to an assistant feels like a conversation or like operating a machine. People take turns with remarkably short gaps, and they do it by predicting when the other person will finish rather than waiting for silence. A voice assistant that waits too long feels slow; one that jumps in too early interrupts your thinking; one that keeps talking when you speak feels rude. We want Ethen Voice to feel quick but not hasty, brief by default, easy to interrupt, calm rather than performative, honest when it did not hear you, and always clear that you are talking to an AI. This article explains those principles, the research behind them, and the tensions that make good voice harder than it sounds.
Key takeaways
- Speed matters, but timing matters more. Knowing when you have finished is as important as answering fast.
- Interruption is a feature. You should be able to cut in, and the assistant should stop.
- Short answers respect attention. Detail belongs on a screen, not in a monologue.
- Calm beats charm. No performed emotion, no flattery, no imitation of a real person.
- Honesty includes hearing. "I didn't catch that" is better than a confident guess.
- Voice must never be the only way. Captions, speed controls and text alternatives are part of the design.
Why do AI voice assistants feel awkward?
AI voice assistants usually feel awkward because their timing is wrong, not because their voices sound artificial. Human conversation runs on a precise rhythm, and small departures from it are noticeable.
Research on conversation across languages makes the rhythm concrete. A study of ten languages, from very different cultures, found the same basic pattern everywhere: people avoid talking over each other and minimize the silence between turns. Languages differed in their average gap, but only within a narrow band of about a quarter of a second around the overall mean. A later review of turn-taking research put the typical gap between turns at around 200 milliseconds — while noting that planning even a short utterance takes considerably longer than that. The only way people manage it is by predicting where the other person's turn will end and preparing a response before it does.
That is the core difficulty for voice AI. Many systems have decided that a turn is over by waiting for a stretch of silence. Set the wait short, and the assistant interrupts people who pause to think. Set it long, and every exchange has a dead beat that makes the assistant feel slow. A review of turn-taking in conversational systems describes exactly this problem with silence-based approaches, and the research direction of predicting turn ends from more than silence alone — the words so far, the intonation, the rhythm.
Voice models have become much faster in the last few years, and that helps. But a fast answer at the wrong moment is still the wrong answer. Figure 1 breaks a good turn into its parts.
Principle 1: quick, but not hasty
We want Ethen Voice to respond quickly enough that the conversation keeps its rhythm, without answering before you have finished. In practice that means two things.
First, the first sound should be useful. When a full answer takes time — because it needs to look something up or think through a problem — a short, honest acknowledgment ("Let me check the project notes") keeps the rhythm without pretending the answer is ready. What it should not do is fill time with empty phrases or restate your question at length.
Second, the assistant should not rush to fill a pause. People pause mid-sentence to find a word, check a number or think. Treating every pause as the end of a turn makes the assistant feel impatient, and it forces people to speak in an unnaturally continuous way. Erring slightly toward patience when the sentence sounds unfinished is better than cutting in.
We are deliberately not publishing response-time targets here. Timing depends on the task, the network, the device and the model, and any number we stated would be either a measurement we have not published or a promise we cannot yet support. What we can state is the priority: timing that follows the person, not the system.
Principle 2: easy to interrupt
You should be able to interrupt Ethen Voice at any time, and when you do, it should stop. This is sometimes called barge-in, and it is one of the clearest differences between a voice assistant that feels like a conversation and one that feels like a recording.
People interrupt for good reasons: the answer is already clear, the assistant misunderstood, or something more important has come up. An assistant that keeps talking, or that finishes its sentence before acknowledging you, communicates that its output matters more than your attention.
Interruption also needs to be handled gracefully afterward. If you cut in with a correction, the assistant should treat the correction as the new direction rather than resuming where it left off. If you cut in by accident — a cough, someone else in the room — it should be easy to continue. Getting this right is harder than stopping audio playback, because the assistant has to decide what the interruption meant.
Principle 3: brief by default
Spoken answers should be short. Listening is slower than reading, and you cannot skim audio. A paragraph that takes a few seconds to read can take much longer to hear, and the listener has to hold the beginning in memory while the end arrives.
So Ethen Voice should give the answer first, in a sentence or two, and offer more only if you want it. When the full answer is long or detailed — a comparison, a list of options, a passage of text, any numbers you will need later — the right move is to put it on the screen and say so briefly. We describe how voice and screens divide work in How Voice Fits Into the Ethen Experience.
Brevity is also a matter of tone. Spoken padding — "Great question! I'd be happy to help you with that" — is more tiring to hear than to read. We want Ethen Voice to skip it.
Principle 4: calm presence, not performance
We want Ethen Voice to sound calm, steady and attentive. We do not want it to perform emotion, flatter you, or try to make you feel attached to it.
This is a choice with real tradeoffs. Warm, expressive voices can be pleasant, and modern voice models are very good at expressiveness. But in a work tool, performed enthusiasm quickly becomes noise, and emotional performance can make it harder to judge what the assistant actually knows. A voice that sounds equally delighted about good news and bad news is not communicating anything.
Calm does not mean cold. The voice should be natural, polite and responsive to the situation — quieter and slower when you are clearly concentrating, more concise when you are in a hurry. The aim is the kind of presence you would want from a capable colleague: there when needed, not demanding attention, not putting on a show. The same idea runs through our description of what we want My Ethen to feel like, where calm and user control are explicit design principles.
Figure 2 summarizes the stance.
Principle 5: honest about what it heard
When Ethen Voice is not sure what you said, it should say so. Speech recognition makes mistakes, especially with names, numbers, technical terms, accents and background noise. A confident answer to a misheard question is worse than a short "I didn't catch the name — could you spell it?" because the error may not be noticed until later.
This matters most when the stakes rise. For a casual question, a small mishearing costs little. For anything that will become a record or an action — a task, a message, a date, a recipient — the assistant should read back the important parts, and for consequential actions the exact details should be confirmed on a screen.
Honesty about hearing is part of a broader principle in Ethen: showing what the system does not know rather than hiding it. We wrote about that in Why Ethen Shows What It Knows—and What It Doesn't. Ethen Research Lab has explored a related idea for agent actions — treating an uncertain outcome as unknown rather than guessing success or failure — in Unknown Effects in Autonomous AI Systems, a research note that proposes the idea rather than reporting measurements. The same instinct applies to listening: an unknown should be marked as unknown.
Principle 6: clearly an AI
You should always know you are talking to an AI. Ethen Voice should not claim to be a person, should not imitate a real person's voice, and should not use conversational tricks designed to obscure what it is.
This is partly about trust and partly about good interaction design. Guidelines for human-AI interaction, developed from a large review of design recommendations, stress making clear what a system can do and how well it can do it. A voice that sounds and behaves exactly like a human invites people to assume human-level understanding, memory and judgment. A voice that is natural but unmistakably an assistant sets better expectations.
Principle 7: never voice-only
Voice should be one way to use Ethen, not a requirement. People who are deaf or hard of hearing, people with speech disabilities, people in noisy or shared spaces and people who simply prefer to read should never be at a disadvantage.
That means captions or transcripts for spoken output, control over speaking speed, the ability to switch to text in the middle of a conversation without losing context, and no feature that can only be reached by voice. Accessibility standards such as the Web Content Accessibility Guidelines provide the baseline for web interfaces; for voice, the principle extends naturally to always offering an equivalent visual path.
The tensions behind good voice
Every principle above sits on a balance, and most voice problems come from setting a balance too far one way. Figure 3 lists the main tensions.
End-of-turn detection is highlighted because it is where the most visible failures happen: an assistant that interrupts people while they think, or one that leaves a dead pause after every sentence. Confirmation versus flow is the tension with the highest stakes: asking for confirmation on everything makes voice tedious, while confirming nothing means acting on mishearings.
No single setting is right for everyone. Some people speak in long, considered sentences with pauses; others speak in quick fragments. A meeting room needs different behavior from a quiet car. Over time, the right approach is likely to combine good defaults with simple controls that let people adjust how patient, how brief and how cautious their assistant is.
How these principles connect to trust
The feel of a voice assistant and its trustworthiness are related. An assistant that interrupts, rambles or guesses teaches people not to rely on it. One that listens carefully, answers briefly and admits uncertainty earns reliance gradually.
The trust also has to be built underneath the voice. In Ethen Chat, a voice session starts only after server-side checks, the browser receives a short-lived pass rather than a long-lived key, microphone access is allowed only on chat pages, and a new session starts with no tools. The details are in How Ethen Chat Starts a Voice Session. That engineering post is also candid that its release checks did not include live microphone capture or end-to-end spoken conversations, which is why this article describes design intent rather than measured results.
Tradeoffs and limitations
These are design principles, not measurements. We have not published voice latency, interruption or recognition results for Ethen Voice, and this article does not claim any.
Some principles conflict with what impresses in demos. Highly expressive, emotional voices and long, flowing answers demo well. We are choosing restraint, and some people will prefer the more theatrical style.
Turn-taking remains an open research problem. Predicting when someone has finished speaking is hard for machines, and even good systems will sometimes interrupt or hesitate.
Calm is subjective. What sounds calm to one person sounds flat to another, and preferences vary across cultures and contexts. Defaults will need adjusting with real feedback.
Availability is not announced here. Ethen Voice is the name we use for voice as a product direction. Languages, devices and plans will be stated on product pages when they are true.
FAQ
Why do AI voice assistants feel awkward? Mostly because of timing. People take turns with very short gaps and predict when the other person will finish. Assistants that wait for long silences feel slow, and ones that treat every pause as an ending interrupt people.
How fast should a voice assistant respond? Fast enough to keep the rhythm of conversation, but not before the person has finished. A short, useful first response is better than a long silence followed by a monologue.
Should an AI voice sound human? It should sound natural, but it should always be clear that it is an AI. Ethen Voice should not claim to be a person or imitate a real person's voice.
Can I interrupt Ethen Voice? That is a core principle: you should be able to cut in at any time, and the assistant should stop and follow your correction.
What if Ethen Voice mishears me? It should say when it is unsure, read back important details, and send consequential actions to a screen for confirmation.
Can I use Ethen without voice? Yes. Voice is always optional, with text and visual alternatives for everything.
Related reading
- How Voice Fits Into the Ethen Experience
- How Ethen Chat Starts a Voice Session
- Why Ethen Shows What It Knows—and What It Doesn't
- Why Long-Running AI Work Needs a Different UX Than Chat
- What We Want My Ethen to Feel Like
References
- Stivers, T., Enfield, N. J., Brown, P., Englert, C., Hayashi, M., Heinemann, T., Hoymann, G., Rossano, F., de Ruiter, J. P., Yoon, K.-E., & Levinson, S. C. (2009). Universals and cultural variation in turn-taking in conversation. Proceedings of the National Academy of Sciences, 106(26), 10587–10592. https://doi.org/10.1073/pnas.0903616106
- Levinson, S. C., & Torreira, F. (2015). Timing in turn-taking and its implications for processing models of language. Frontiers in Psychology, 6, 731. https://doi.org/10.3389/fpsyg.2015.00731
- Skantze, G. (2021). Turn-taking in conversational systems and human-robot interaction: A review. Computer Speech & Language, 67, 101178. https://doi.org/10.1016/j.csl.2020.101178
- Amershi, S., et al. (2019). Guidelines for Human-AI Interaction. Proceedings of CHI 2019. https://doi.org/10.1145/3290605.3300233
- W3C. Web Content Accessibility Guidelines (WCAG) 2.2. https://www.w3.org/TR/WCAG22/
- Ethen Research Lab (2026). Unknown Effects in Autonomous AI Systems: Why Timeouts Are Not Permission to Retry. Research note. https://upcube.ai/resources/research/unknown-effects