Skip to content

EthenEthenEthen

Open Source Model Profile · ResembleAI

chatterbox-turbo

Chatterbox-Turbo is a ResembleAI text-to-speech model described as a streamlined 350M-parameter Turbo variant. It is recorded with MIT licensing and voice-cloning tags.

Publisher
ResembleAI
Task
text-to-speech
Model type
Unknown
License
mit
Library
Unknown
Publication status
Accepted · not indexed

Model overview

Chatterbox-Turbo is published by ResembleAI as a text-to-speech model in the Chatterbox family. The card introduces it as the most efficient model yet on a streamlined 350M-parameter architecture. Hub tags mark speech, speech-generation, voice-cloning, and English, with MIT licensing.

Recorded capabilities

350M streamlined Turbo architecture

According to the model card, Turbo uses a streamlined 350M-parameter architecture intended to need less compute and VRAM than prior models.

Distilled one-step decoder

The card states the speech-token-to-mel decoder was distilled from 10 steps to one while retaining high-fidelity output.

Native paralinguistic tags

The publisher states paralinguistic tags are now native to the Turbo model.

Use cases in the source record

  • Text-to-speech synthesis with voice cloning by supplying an audio prompt, as shown in the card’s generate-with-audio-prompt example.
  • Multilingual synthesis experiments matching the card’s Chinese-text example with a language identifier.

Limitations and unknowns

  • No Safetensors parameter count, architecture class, or model-type record was extracted; the 350M figure is a publisher claim.
  • No evaluation results were extracted from this record.
  • Latency, VRAM, and quality comparisons in the card are publisher descriptions and were not independently measured.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.

Source and provenance

Source: ResembleAI/chatterbox-turbo

Captured: Unknown. Processed: 2026-09-07T19:34:36.310845+00:00.

Chatterbox TTS Made with ❤️ by Chatterbox is a family of three state-of-the-art, open-source text-to-speech models by Resemble AI. We are excited to introduce Chatterbox-Turbo , our most efficient model yet. Built on a streamlined 350M parameter architecture, Turbo delivers high-quality speech with less compute and VRAM than our previous models. We have also distilled the speech-token-to-mel decoder, previously a bottleneck, reducing generation from 10 steps to just one , while retaining high-fidelity audio output. Paralinguistic tags are now native to the Turbo model, allowing you to use [cough] , [laugh] , [chuckle] , and more to…

F001F002F003F004F005F006F007F008F009F011F012