audiospeech
Dia TTS
1.6B-parameter Dia TTS with multi-speaker tags, emotion control, and natural nonverbals.
Studio is in early access; availability resolves per project after sign-in.
- Endpoints
- 2
- Eligible
- 2
- Input schemas imported
- 2 / 2
- Delivery
- fal.ai
- Weights
- Open weights
- Developer
- Not stated in source
Overview
Dia TTS transforms text into natural-sounding, studio-quality speech with exceptional clarity and prosody. A 1.6-billion-parameter model trained on extensive voice datasets captures pronunciation, intonation, and emotional nuance. Multi-speaker conversations run on [S1]/[S2] tags with natural nonverbals like laughter; a voice-clone endpoint joins the base.
Capabilities
- 1.6B dialogue model
- Multi-speaker tags
- Emotion conditioning
- Natural nonverbals
Best for
- Realistic dialogue
- Studio-quality speech
- Voice cloning
Use cases
- Conversational synthesis
- Multi-speaker scenes
- Expressive narration
Endpoints
2 eligible. Cataloged is not the same as executable: Studio resolves which endpoints a project can run.
| Endpoint | Task | Catalog status | Inputs |
|---|---|---|---|
fal-ai/dia-tts | text to audio | Eligible | 1 fields · requires text |
fal-ai/dia-tts/voice-clone | text to audio | Eligible | 1 fields · requires text |
Questions
Multi-speaker?
Yes, via [S1] and [S2] tags.
Model size?
1.6B parameters.