audiospeech
F5 TTS
F5-TTS clones voices zero-shot from one sample with F5 and E2 architecture options.
Studio is in early access; availability resolves per project after sign-in.
- Endpoints
- 1
- Eligible
- 1
- Input schemas imported
- 1 / 1
- Delivery
- fal.ai
- Weights
- Not stated
- Developer
- Not stated in source
Overview
F5-TTS synthesizes natural speech from a single reference sample with no fine-tuning, the reference defining the output voice. Two architectures, F5-TTS and E2-TTS, select via model_type. Voice cloning for content, multilingual production, and game character voices lead uses.
Capabilities
- Zero-shot cloning
- Single-sample input
- Dual architectures
- No training needed
Best for
- Content voice cloning
- Multilingual audio
- Game character voices
Use cases
- Custom voices
- Reference-driven synthesis
- Dataset-free cloning
Endpoints
1 eligible. Cataloged is not the same as executable: Studio resolves which endpoints a project can run.
| Endpoint | Task | Catalog status | Inputs |
|---|---|---|---|
fal-ai/f5-tts | text to audio | Eligible | 5 fields · requires gen_text, ref_audio_url, model_type |
Questions
Training needed?
No: single sample suffices.
Variants?
F5-TTS and E2-TTS.