350M streamlined Turbo architecture
According to the model card, Turbo uses a streamlined 350M-parameter architecture intended to need less compute and VRAM than prior models.
Open Source Model Profile · ResembleAI
Chatterbox-Turbo is a ResembleAI text-to-speech model described as a streamlined 350M-parameter Turbo variant. It is recorded with MIT licensing and voice-cloning tags.
Chatterbox-Turbo is published by ResembleAI as a text-to-speech model in the Chatterbox family. The card introduces it as the most efficient model yet on a streamlined 350M-parameter architecture. Hub tags mark speech, speech-generation, voice-cloning, and English, with MIT licensing.
According to the model card, Turbo uses a streamlined 350M-parameter architecture intended to need less compute and VRAM than prior models.
The card states the speech-token-to-mel decoder was distilled from 10 steps to one while retaining high-fidelity output.
The publisher states paralinguistic tags are now native to the Turbo model.
Source: ResembleAI/chatterbox-turbo
Captured: Unknown. Processed: 2026-09-07T19:34:36.310845+00:00.
Chatterbox TTS Made with ❤️ by Chatterbox is a family of three state-of-the-art, open-source text-to-speech models by Resemble AI. We are excited to introduce Chatterbox-Turbo , our most efficient model yet. Built on a streamlined 350M parameter architecture, Turbo delivers high-quality speech with less compute and VRAM than our previous models. We have also distilled the speech-token-to-mel decoder, previously a bottleneck, reducing generation from 10 steps to just one , while retaining high-fidelity audio output. Paralinguistic tags are now native to the Turbo model, allowing you to use [cough] , [laugh] , [chuckle] , and more to…
F001F002F003F004F005F006F007F008F009F011F012