Skip to content

EthenEthenEthen

Open Source Model Profile · meituan-longcat

LongCat-Video

LongCat-Video is a 13.6B-parameter video-generation model from meituan-longcat. According to the model card, one dense model covers text-to-video, image-to-video, and video continuation, including minutes-long output.

Publisher
meituan-longcat
Task
text-to-video
Model type
Unknown
License
mit
Library
diffusers
Publication status
Accepted · not indexed

Model overview

LongCat-Video is published by meituan-longcat as a foundational text-to-video model. Hub metadata records diffusers support with image-to-video, video-continuation, and text-to-video tags. According to the model card, it is a 13.6B-parameter dense model built for efficient high-quality long-video generation.

Recorded capabilities

Unified video-task framework

According to the model card, one model natively supports text-to-video, image-to-video, and video-continuation tasks.

Long-video continuation design

The publisher says native video-continuation pretraining supports minutes-long videos without color drift or quality degradation.

Efficient 720p inference account

According to the model card, 720p 30fps videos are generated within minutes through coarse-to-fine temporal and spatial generation with Block Sparse Attention.

Publisher-reported MOS tables

According to the model card, text-to-video overall quality is 3.38 and image-to-video overall quality is 3.17, with full text, visual, motion, and alignment breakdowns versus Veo3, PixVerse, Wan, Seedance, and Hailuo systems.

Use cases in the source record

  • Text-to-video generation using the documented single- or multi-GPU torchrun demo scripts.
  • Image-to-video and video-continuation generation, including long-video, interactive-video, and Streamlit demo paths.
  • 720p, 30fps video work that uses the publisher-described coarse-to-fine temporal and spatial strategy.

Limitations and unknowns

  • Benchmark and WBench figures are publisher-reported values from the model card and have not been independently verified by Ethen.
  • According to the model card, the model has not been specifically designed or comprehensively evaluated for every possible downstream application.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: meituan-longcat/LongCat-Video

Captured: Unknown. Processed: 2026-09-07T19:34:51.665944+00:00.

LongCat-Video Model Introduction We introduce LongCat-Video, a foundational video generation model with 13.6B parameters, delivering strong performance across Text-to-Video , Image-to-Video , and Video-Continuation generation tasks. It particularly excels in efficient and high-quality long video generation, representing our first step toward world models. Key Features 🌟 Unified architecture for multiple tasks : LongCat-Video unifies Text-to-Video , Image-to-Video , and Video-Continuation tasks within a single video generation framework. It natively supports all these tasks with a single model and consistently delivers strong perfor…

F001F002F003F004F005F006F007F008F009F010F011F012F013F014F015F016F019F022