Skip to content

EthenEthenEthen

Open Source Model Profile · tencent

HunyuanVideo

HunyuanVideo is a text-to-video foundation model from tencent. According to the model card, it provides PyTorch weights and inference code with Diffusers integration, FP8 weights, and multi-GPU support.

Publisher
tencent
Task
text-to-video
Model type
Unknown
License
tencent-hunyuan-community
Library
Unknown
Publication status
Accepted · not indexed

Model overview

HunyuanVideo is published by tencent as a text-to-video foundation model. According to the model card, the repository provides PyTorch model definitions, pretrained weights, and inference code, and describes a video generative model over 13B parameters trained with image-video joint learning. Card data records an other license, with arxiv:2412.03603 and arxiv:2405.07719 tags.

Recorded capabilities

Unified image-video Transformer

According to the model card, HunyuanVideo introduces a Transformer design with Full Attention for unified image and video generation, described as a video foundation model over 13B parameters.

MLLM text encoder

According to the model card, prompts are encoded with a decoder-only multimodal large language model plus a bidirectional token refiner, rather than CLIP plus T5-XXL.

Causal 3D VAE

According to the model card, a 3D VAE with CausalConv3D compresses video length, space, and channel by 4, 8, and 16 to reduce diffusion-transformer tokens.

Diffusers, FP8, and xDiT paths

According to the model card, the model is integrated into Diffusers, FP8 weights save about 10GB of GPU memory, and xDiT sequence parallelism supports multi-GPU inference.

Documented GPU requirements

According to the model card, 720px1280px129f needs 60GB and 544px960px129f needs 45GB minimum, with an 80GB GPU recommended and CUDA required.

Use cases in the source record

  • Text-to-video generation through the documented single-GPU CLI and Gradio paths at listed resolutions and frame counts.
  • Diffusers, ComfyUI, and xDiT multi-GPU workflows using the publisher-documented parallel and FP8 inference options.

Limitations and unknowns

  • No independently verified evaluation results were extracted; comparison figures in the card are publisher-reported.
  • No structured parameter count, architecture config, or context-window value was extracted from this record.
  • Provider state is historical snapshot data, not independently refreshed current availability.
  • Scale, training, and comparison claims come from the publisher model card and were not independently verified by Ethen.

Source and provenance

Source: tencent/HunyuanVideo

Captured: Unknown. Processed: 2026-09-07T19:34:59.469947+00:00.

HunyuanVideo: A Systematic Framework For Large Video Generation Model Training This repo contains PyTorch model definitions, pre-trained weights and inference/sampling code for our paper exploring HunyuanVideo. You can find more visualizations on our project page . HunyuanVideo: A Systematic Framework For Large Video Generation Model Training News!! Jan 13, 2025: 📈 We release the Penguin Video Benchmark . Dec 18, 2024: 🏃‍♂️ We release the FP8 model weights of HunyuanVideo to save more GPU memory. Dec 17, 2024: 🤗 HunyuanVideo has been integrated into Diffusers . Dec 7, 2024: 🚀 We release the parallel inference code for HunyuanVid…

F001F002F003F004F005F006F007F011F014F015F017F018F019F022F023F025F026F028F029F030F032F033F034