Skip to content

EthenEthenEthen

audiospeech

Dia TTS

1.6B-parameter Dia TTS with multi-speaker tags, emotion control, and natural nonverbals.

Studio is in early access; availability resolves per project after sign-in.

Endpoints
2
Eligible
2
Input schemas imported
2 / 2
Delivery
fal.ai
Weights
Open weights
Developer
Not stated in source

Overview

Dia TTS transforms text into natural-sounding, studio-quality speech with exceptional clarity and prosody. A 1.6-billion-parameter model trained on extensive voice datasets captures pronunciation, intonation, and emotional nuance. Multi-speaker conversations run on [S1]/[S2] tags with natural nonverbals like laughter; a voice-clone endpoint joins the base.

Capabilities

  • 1.6B dialogue model
  • Multi-speaker tags
  • Emotion conditioning
  • Natural nonverbals

Best for

  • Realistic dialogue
  • Studio-quality speech
  • Voice cloning

Use cases

  • Conversational synthesis
  • Multi-speaker scenes
  • Expressive narration

Endpoints

2 eligible. Cataloged is not the same as executable: Studio resolves which endpoints a project can run.

EndpointTaskCatalog statusInputs
fal-ai/dia-ttstext to audioEligible1 fields · requires text
fal-ai/dia-tts/voice-clonetext to audioEligible1 fields · requires text

Questions

Multi-speaker?

Yes, via [S1] and [S2] tags.

Model size?

1.6B parameters.