Omni-modal video generation
According to the model card, the system handles text, image, video, and audio context and generates video with native stereo audio.
Open Source Model Profile · MiniMaxAI
MiniMax-H3 is a 33.12B-parameter multimodal video system from MiniMaxAI. Its model card documents 4-15 second generation with stereo audio and a 2K regeneration path.
MiniMax-H3 is published by MiniMaxAI as an omni-modal generative video system. Captured Safetensors metadata reports about 33.12B parameters. According to the model card, it unifies text, image, video, and audio context around a 33B-parameter H3-Omni-Transformer for video with native stereo audio.
According to the model card, the system handles text, image, video, and audio context and generates video with native stereo audio.
The publisher documents 4-15 second outputs at 24 FPS, with H3-Regenerate-2K rebuilding 768p results at 2K resolution.
According to the model card, H3-Omni-Transformer is a 33B-parameter dense single-stream Transformer, with sparse attention planned for a later release.
Source: MiniMaxAI/MiniMax-H3
Captured: Unknown. Processed: 2026-09-07T19:34:34.322694+00:00.
MiniMax H3 News Offical skills to improve prompt writing: skills on github Online API Use MiniMax-H3 directly via API. Global: platform.minimax.io | CN: platform.minimaxi.com Online App Use MiniMax-H3 directly via App. WebApp Global: hailuoai.video | CN: hailuoai.com Desktop Global: hub.minimax.io | CN: hub.minimaxi.com System Overview MiniMax H3 is a general-purpose, omni-modal generative system. It supports unified understanding of multimodal contexts composed of text, images, video, and audio, and can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds. Thanks to its task-generalizati…
F001F002F003F004F005F006F007F008F011F012F014F017F023F025F030F031F039