video
MiniMax H3 Max
H3-Max family with camera controls, director, image, lip-sync, reference, and text video modes.
Studio is in early access; availability resolves per project after sign-in.
- Endpoints
- 6
- Eligible
- 4
- Input schemas imported
- 5 / 6
- Delivery
- fal.ai
- Weights
- Not stated
- Developer
- Not stated in source
Overview
The H3-Max family spans six video endpoints: camera-controls, director, image-to-video, lip-sync image-to-video, reference-to-video, and text-to-video. Billing separates output generation from reference inputs, with an included reference-token allowance shared across images, videos, and audio clips. Output bills per second at the selected resolution.
Capabilities
- Six-mode video family
- Camera-controls endpoint
- Director endpoint
- Lip-sync i2v tier
Best for
- Reference-driven video
- Text-prompted clips
- Image-animated scenes
Use cases
- Camera-directed generation
- Lip-synced animation
- Reference-guided video
Endpoints
2 under review · 4 eligible. Cataloged is not the same as executable: Studio resolves which endpoints a project can run.
| Endpoint | Task | Catalog status | Inputs |
|---|---|---|---|
minimax/h3-max/camera-controls | unknown | Under review | 9 fields · requires image_url, prompt_expansion_mode |
minimax/h3-max/director | unknown | Under review | Schema not imported |
minimax/h3-max/image-to-video | image to video | Eligible | 10 fields · requires prompt, prompt_expansion_mode |
minimax/h3-max/lip-sync/image-to-video | image to video | Eligible | 6 fields · requires image_url, audio_url |
minimax/h3-max/reference-to-video | reference to video | Eligible | 15 fields · requires prompt, prompt_expansion_mode |
minimax/h3-max/text-to-video | text to video | Eligible | 9 fields · requires prompt, prompt_expansion_mode |
Questions
Which modes?
Camera, director, image, lip-sync, reference, text.
How does billing split?
Output generation plus reference inputs.