Conversational speech generation
According to the model card, CSM generates RVQ audio codes from text and audio inputs.
Open Source Model Profile · sesame
csm-1b is a Sesame 1.55B-parameter Conversational Speech Model for text-to-speech. The model card documents RVQ audio generation with Transformers support.
csm-1b is published by Sesame as a Conversational Speech Model for text-to-speech. The captured configuration identifies CsmForConditionalGeneration with model type csm, and Safetensors metadata reports 1,552,791,552 parameters. According to the model card, it uses a Llama backbone with an audio decoder producing Mimi codes under Apache-2.0.
According to the model card, CSM generates RVQ audio codes from text and audio inputs.
The model card describes a Llama backbone plus a smaller audio decoder that produces Mimi audio codes.
The model card documents batched inference and says CSM can be fine-tuned with the Transformers Trainer.
According to the model card, the model cannot generate text and is not a general-purpose multimodal LLM.
Source: sesame/csm-1b
Captured: Unknown. Processed: 2026-09-07T19:34:57.720016+00:00.
CSM 1B 2025/05/20 - CSM is availabile natively in Hugging Face Transformers 🤗 as of version 4.52.1 2025/03/13 - We are releasing the 1B CSM variant. The checkpoint is hosted on Hugging Face . CSM (Conversational Speech Model) is a speech generation model from Sesame that generates RVQ audio codes from text and audio inputs. The model architecture employs a Llama backbone and a smaller audio decoder that produces Mimi audio codes. A fine-tuned variant of CSM powers the interactive voice demo shown in our blog post . A hosted HuggingFace space is also available for testing audio generation. Usage Generate a sentence import torch from…
F001F002F003F004F005F006F007F010F012F016F018F019F020