Skip to content

EthenEthenEthen

video

MiniMax H3 Max

H3-Max family with camera controls, director, image, lip-sync, reference, and text video modes.

Studio is in early access; availability resolves per project after sign-in.

Endpoints
6
Eligible
4
Input schemas imported
5 / 6
Delivery
fal.ai
Weights
Not stated
Developer
Not stated in source

Overview

The H3-Max family spans six video endpoints: camera-controls, director, image-to-video, lip-sync image-to-video, reference-to-video, and text-to-video. Billing separates output generation from reference inputs, with an included reference-token allowance shared across images, videos, and audio clips. Output bills per second at the selected resolution.

Capabilities

  • Six-mode video family
  • Camera-controls endpoint
  • Director endpoint
  • Lip-sync i2v tier

Best for

  • Reference-driven video
  • Text-prompted clips
  • Image-animated scenes

Use cases

  • Camera-directed generation
  • Lip-synced animation
  • Reference-guided video

Endpoints

2 under review · 4 eligible. Cataloged is not the same as executable: Studio resolves which endpoints a project can run.

EndpointTaskCatalog statusInputs
minimax/h3-max/camera-controlsunknownUnder review9 fields · requires image_url, prompt_expansion_mode
minimax/h3-max/directorunknownUnder reviewSchema not imported
minimax/h3-max/image-to-videoimage to videoEligible10 fields · requires prompt, prompt_expansion_mode
minimax/h3-max/lip-sync/image-to-videoimage to videoEligible6 fields · requires image_url, audio_url
minimax/h3-max/reference-to-videoreference to videoEligible15 fields · requires prompt, prompt_expansion_mode
minimax/h3-max/text-to-videotext to videoEligible9 fields · requires prompt, prompt_expansion_mode

Questions

Which modes?

Camera, director, image, lip-sync, reference, text.

How does billing split?

Output generation plus reference inputs.