Skip to content

EthenEthenEthen

Open Source Model Profile · genmo

mochi-1-preview

Mochi-1-preview is a 10.03B-parameter text-to-video diffusion model from Genmo. According to the model card, it uses an Asymmetric Diffusion Transformer architecture and is released under Apache-2.0.

Publisher
genmo
Task
text-to-video
Model type
Unknown
License
apache-2.0
Library
diffusers
Publication status
Accepted · not indexed

Model overview

Mochi-1-preview is published by Genmo as a text-to-video model. Safetensors metadata reports 10,027,677,744 parameters, and the record lists the diffusers library with an Apache-2.0 license. According to the model card, it is a 10B-parameter diffusion model built on the Asymmetric Diffusion Transformer architecture and trained from scratch.

Recorded capabilities

Open text-to-video generation

According to the model card, Mochi 1 preview generates video from text with high-fidelity motion and strong prompt adherence in preliminary evaluation.

AsymmDiT and AsymmVAE design

According to the model card, the 10B AsymmDiT jointly attends to text and visual tokens with separate MLP layers, paired with a 362M AsymmVAE using 8x8 spatial and 6x temporal compression.

Diffusers workflow

According to the model card, the model runs through MochiPipeline in Diffusers with documented height, width, frame-count, and scheduler settings.

Documented hardware guidance

According to the model card, the highest-quality Diffusers example needs 42GB VRAM, while the repository path needs about 60GB VRAM on a single GPU.

Use cases in the source record

  • Text-to-video synthesis from prompts, following the model card's pipeline and Diffusers examples for frame count, resolution, and export to video.
  • Local or multi-GPU video-generation experiments using the documented repository or Diffusers paths, subject to the publisher's VRAM and H100 guidance.

Limitations and unknowns

  • No evaluation results were extracted from this record beyond publisher-described preliminary evaluation.
  • According to the model card, the models reflect training-data biases and need additional safety protocols before commercial deployment.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: genmo/mochi-1-preview

Captured: Unknown. Processed: 2026-09-07T19:34:45.219934+00:00.

Mochi 1 Blog | Direct Download | Hugging Face Download | Playground | Careers A state of the art video generation model by Genmo . Your browser does not support the video tag. Overview Mochi 1 preview is an open state-of-the-art video generation model with high-fidelity motion and strong prompt adherence in preliminary evaluation. This model dramatically closes the gap between closed and open video generation systems. We’re releasing the model under a permissive Apache 2.0 license. Try this model for free on our playground . Installation Install using uv : git clone https://github.com/genmoai/models cd models pip install uv uv venv…

F001F002F003F004F005F006F008F011F012F015F016F017F018F020F021F022F023