Skip to content

EthenEthenEthen

audiospeech

F5 TTS

F5-TTS clones voices zero-shot from one sample with F5 and E2 architecture options.

Studio is in early access; availability resolves per project after sign-in.

Endpoints
1
Eligible
1
Input schemas imported
1 / 1
Delivery
fal.ai
Weights
Not stated
Developer
Not stated in source

Overview

F5-TTS synthesizes natural speech from a single reference sample with no fine-tuning, the reference defining the output voice. Two architectures, F5-TTS and E2-TTS, select via model_type. Voice cloning for content, multilingual production, and game character voices lead uses.

Capabilities

  • Zero-shot cloning
  • Single-sample input
  • Dual architectures
  • No training needed

Best for

  • Content voice cloning
  • Multilingual audio
  • Game character voices

Use cases

  • Custom voices
  • Reference-driven synthesis
  • Dataset-free cloning

Endpoints

1 eligible. Cataloged is not the same as executable: Studio resolves which endpoints a project can run.

EndpointTaskCatalog statusInputs
fal-ai/f5-ttstext to audioEligible5 fields · requires gen_text, ref_audio_url, model_type

Questions

Training needed?

No: single sample suffices.

Variants?

F5-TTS and E2-TTS.