Ethen Voice
Give intelligence a voice.
Talk with Ethen in real time, generate speech, and build voices, transcripts and voice agents — one voice platform that powers Chat, Studio and the agents you run.
- Live in Chat Realtime conversation
- Studio early access Speech generation
- YouWhat needs me today?
- EthenMorning. Two things need you today: the pricing brief is ready for review, and the launch call moved to three.
- YouRead me the first line of the brief.
Your browser has no built-in speech engine, so this preview is silent. Ethen voices play inside Chat and Studio.
The platform
One voice platform, every kind of speech.
Conversation, speech generation, transcription, voices, agents and localization share one identity, one consent model and one record of what was said.
Generate natural speech
Turn text into speech for narration, voiceovers and audio for video. Speech generation runs as a Studio audio job, so it is quoted, tracked and exported like the rest of your media.
- Text to speech as a tracked job
- Voice chosen per project
- Exports with your Studio media
Every launch starts the same way: a blank page, a deadline, and a team that needs one clear story.
Transcribe spoken audio
Every voice conversation in Chat is saved as text in the same thread, so you can read it back, search it and follow up by typing. Transcription of uploaded recordings is built in Studio and opens when audio upload does.
- Voice turns saved as chat text
- Same history, search and deletion as typed turns
- File transcription: coming with Studio upload
- YouMove the design review to Thursday and tell Maya.voice
- EthenMoved to Thursday at 2. I drafted the note to Maya — want me to send it?voice
- YouSend it, and add the agenda.typed
Talk with Ethen in real time
Start a voice session in Chat and talk. Ethen listens, answers out loud and keeps the conversation in your thread — no separate assistant, no separate history.
- Speech in, speech out
- Single-use session credential
- Text and voice in one thread
Listening…
Build and manage AI voices
Voice identities are versioned and bound to a provider voice before use. Designing a voice from a description is built in Studio; a browsable public library is coming.
- Versioned voice identities
- Explicit binding state, never assumed usable
- Consent checked at creation and at use
- Kind
- Original — designed from a description
- Binding
- Pending — not usable until bound
- Consent
- Not a real person’s voice
Build voice agents
Voice agents that listen, reason, use tools and keep context, with spend caps, transcripts and tool oversight. The Studio route is built; it opens when the realtime media transport is enabled.
- Spend caps per session
- Transcripts and tool oversight
- Reconnect recovery
- CallerCan I move my booking to Friday?
- AgentFriday at 10 is open. I’ll hold it while I confirm.
Localize speech across languages
Dubbing is designed as five tracked stages — transcribe, translate, speaker synthesis, align and mix — each one a job you can inspect. It is not available yet.
- Transcribe → translate → synthesize → align → mix
- Explicit source and target languages
- Speaker map per project
- 01Transcribe
- 02Translate
- 03Speaker synthesis
- 04Align
- 05Mix
Voice library
Discover and keep voices.
A library of original voices to browse, compare and bind to your projects, with your own voices beside them. Voice identities are built in Studio; the public library is coming.
- AsterNarration
- Language
- English (US)
- Tone
- Warm, unhurried
- MarlowConversational
- Language
- English (UK)
- Tone
- Dry, direct
- JunoExplainer
- Language
- English (US)
- Tone
- Bright, clear
- SableDocumentary
- Language
- English (US)
- Tone
- Low, measured
- RenAssistant
- Language
- English (US)
- Tone
- Calm, neutral
- OdetteCharacter
- Language
- French (FR)
- Tone
- Playful
Demo voices — fictional names and descriptions for illustration. Not shipping voices; no real person’s voice.
Realtime
Create realtime voice experiences.
Voice in Ethen Chat is live today: speech in, speech out, and the conversation kept in your thread. This is how a session works.
- 01
You start it
Voice begins only when you choose it. Your browser asks for the microphone, and anonymous voice is refused.
- 02
A single-use session
Chat issues a short-lived credential that expires after ten minutes. The long-lived provider key never leaves the server.
- 03
Direct and encrypted
Your browser connects to the realtime voice service directly, so Upcube's servers neither relay nor record the audio.
- 04
Kept as text
When the exchange ends, what you and Ethen said is saved to the conversation as text, alongside typed turns.
Realtime conversation runs on OpenAI’s realtime voice model, reached through the Vercel AI Gateway. Ethen builds the session, the thread and the safety layer around it.
Voice agents
Build voice agents that listen, reason and act — with permission.
Agents that hold a spoken conversation, keep context, use the tools you allow and stop for approval before anything sensitive. Spend caps, transcripts and tool oversight come built in. The Studio route is built; it opens when the realtime media transport is enabled.
Coming- ListenStreams speech in, with turn-taking.
- ReasonPlans the reply with the context it keeps.
- Use toolsCalls tools you allow, under oversight.
- Ask firstSensitive actions wait for approval.
For developers
Integrate with the Voice API.
Today, audio and voice endpoints are part of the Studio Media API, in alpha: they need a signed-in Studio session and a project. A public, key-based Voice API is coming.
POST /api/studio/v1/audio/projectsCreate a batch audio projectGET /api/studio/v1/audio/projectsRead a project, its stages and transcriptPOST /api/studio/v1/voices/cloneDesign or clone an identity (consent required)GET /api/studio/v1/voices/bindingsList provider bindings
// Studio Media API (alpha) — signed-in Studio session + project scope.
const res = await fetch("https://studio.upcube.ai/api/studio/v1/audio/projects", {
method: "POST",
credentials: "include",
headers: { "content-type": "application/json" },
body: JSON.stringify({
projectId: "<your-project-id>",
kind: "tts", // tts | transcribe | dub | changer
targetLanguage: "en-US" // BCP-47
}),
});Across Ethen
Voice powers experiences across Ethen.
Studio is the creative workspace and Chat is the conversation. Both use Ethen Voice; neither replaces it.
Conversation
Chat
Talk with Ethen and keep the thread. Voice turns live beside typed ones.
Live in ChatEthen ChatCreation
Studio
Narration, voiceovers and audio for video, made and exported as Studio media.
Studio early accessEthen StudioAction
Agents
Voice agents that answer, use tools and hand off — inside the limits you set.
Coming
Safety
Voice safety and consent.
A voice belongs to a person. Ethen Voice is built so that consent comes first, impersonation is out of bounds and you can see who processes your audio.
Consent before creation
A cloned voice requires consent evidence and the rights to the source audio before anything is made.
Consent at every use
Creating a voice never makes it usable by itself. Use stays consent-gated after creation.
What Ethen records, and what it does not verify
Ethen records what you declared about consent and when. It does not verify the speaker's identity or your rights — that responsibility stays with you.
No impersonation
Voices of real people without their permission, and deceptive impersonation, are not allowed.
Disclosure
Synthetic speech should be disclosed where people could mistake it for a real person.
No background listening
No wake word, no voiceprint, no listening outside a session you started. The provider that processes your audio is named in the policy.
Voice cloning is not currently available in Ethen. When it is, the Voice Cloning Policy is updated first.
Research
Research behind a trustworthy voice.
Ethen Research Lab has not published a Voice benchmark, and this page quotes none. The work that shapes Voice is about rights and verification.
- Rights as Infrastructure: Building AI Datasets That Know How They May Be Used
Consent, purpose and revocation as data the system enforces — the model for voice rights.
- How Should We Measure the Reliability of LLM Verifiers?
How Ethen measures whether a check can be trusted — the method behind evaluating spoken answers.
- Ethen Research Lab
All published research, methods and system cards.
From the Ethen Blog
Latest on Ethen Voice.
Questions
What is Ethen Voice?
Ethen Voice is Upcube's platform for speech, voices, audio intelligence and conversational voice AI. It powers voice conversations in Ethen Chat and audio in Ethen Studio, and it is where voice agents, a voice library and a Voice API are being built.
What can I use today?
Realtime voice conversation in Ethen Chat is live for signed-in users, and the exchange is saved as text in your conversation. Speech generation is available in Studio's early access. Voice cloning, dubbing, voice agents, a public voice library and a public Voice API are not available yet.
Which models power Ethen Voice?
Realtime conversation in Chat runs on OpenAI's realtime voice model, reached through the Vercel AI Gateway. Studio audio jobs route to licensed third-party speech models. Ethen builds the product, routing, identity, workflow and safety layers; it does not claim to have trained those speech models.
Does Ethen record my voice?
In Chat, your browser streams audio directly to the realtime voice service only while a session you started is open. Ethen saves the text of the exchange, not an audio recording, does not listen in the background and does not create a voiceprint.
Can I clone a voice?
Not yet. When voice cloning arrives it will require consent evidence and the rights to the source audio before anything is created, and a cloned voice stays consent-gated every time it is used. The Voice Cloning Policy sets out the rules.
Is there a Voice API?
The Studio Media API, in alpha, exposes audio project endpoints for signed-in Studio projects. A public, key-based Voice API is not available yet.
Give intelligence a voice.
Start a realtime voice conversation in Ethen Chat today. Speech generation is open in Studio early access.