Skip to content

EthenEthenEthen

Open Source Model Profile · sesame

csm-1b

csm-1b is a Sesame 1.55B-parameter Conversational Speech Model for text-to-speech. The model card documents RVQ audio generation with Transformers support.

Publisher
sesame
Task
text-to-speech
Model type
csm
License
apache-2.0
Library
transformers
Publication status
Accepted · not indexed

Model overview

csm-1b is published by Sesame as a Conversational Speech Model for text-to-speech. The captured configuration identifies CsmForConditionalGeneration with model type csm, and Safetensors metadata reports 1,552,791,552 parameters. According to the model card, it uses a Llama backbone with an audio decoder producing Mimi codes under Apache-2.0.

Recorded capabilities

Conversational speech generation

According to the model card, CSM generates RVQ audio codes from text and audio inputs.

Llama backbone with Mimi decoder

The model card describes a Llama backbone plus a smaller audio decoder that produces Mimi audio codes.

Documented inference and tuning path

The model card documents batched inference and says CSM can be fine-tuned with the Transformers Trainer.

Audio-only scope

According to the model card, the model cannot generate text and is not a general-purpose multimodal LLM.

Use cases in the source record

  • Text-conditioned speech generation using the card's documented Transformers processor and batched-inference workflow.
  • Fine-tuning experiments using the Transformers Trainer path documented in the model card.

Limitations and unknowns

  • No evaluation results were extracted from this record.
  • No context-window, hardware requirement, or quantization detail was extracted.
  • Provider state is historical snapshot data, not independently refreshed current availability.
  • Architecture and capability details come from the publisher model card and were not independently verified by Ethen.

Source and provenance

Source: sesame/csm-1b

Captured: Unknown. Processed: 2026-09-07T19:34:57.720016+00:00.

CSM 1B 2025/05/20 - CSM is availabile natively in Hugging Face Transformers 🤗 as of version 4.52.1 2025/03/13 - We are releasing the 1B CSM variant. The checkpoint is hosted on Hugging Face . CSM (Conversational Speech Model) is a speech generation model from Sesame that generates RVQ audio codes from text and audio inputs. The model architecture employs a Llama backbone and a smaller audio decoder that produces Mimi audio codes. A fine-tuned variant of CSM powers the interactive voice demo shown in our blog post . A hosted HuggingFace space is also available for testing audio generation. Usage Generate a sentence import torch from…

F001F002F003F004F005F006F007F010F012F016F018F019F020