MiniMax modelsvideo generation api

minimax/h3-max/reference-to-video

H3 Max is MiniMax's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.

Input
Text prompt for video generation. Refer to reference assets by their modality and order in the reference lists: Image 1, Image 2, Video 1, Audio 1, and so on.
The aspect ratio of the generated image. Allowed values: 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16.
The native generation resolution of the video. Allowed values: 480P, 768P.
515
The duration of the video in seconds. Range: 5 to 15.
Random seed. A random seed is selected when omitted.
URLs of subject/style reference images, referenced in the prompt as Image 1, Image 2, and so on. Reference images, videos, and audio clips must add up to at most 12 files.
How much effort to spend rewriting the prompt before generation. 'balanced' returns in about a second. 'quality' spends up to ~30s on a richer prompt.
URLs of motion/reference video clips (2-15 seconds each, combined duration at most 15 seconds), referenced in the prompt as Video 1, Video 2, and so on. Reference images, videos, and audio clips must add up to at most 12 files.
URLs of reference audio clips (2-15 seconds each, combined duration at most 15 seconds), referenced in the prompt as Audio 1, Audio 2, and so on. Audio cannot be the only reference input; provide at least one reference image or video with it. Reference images, videos, and audio clips must add up to at most 12 files.
Idle

Example output — click Run to generate your own

API README

MiniMax H3 Max Reference to Video

Generate video from a text prompt plus image, video, and audio references with MiniMax H3 Max. Reference assets can establish subject identity, appearance, motion, timing, dialogue, ambience, and style.

The endpoint supports 5–15 second output at 480P or 768P, adaptive or fixed aspect ratios, and up to 12 reference files in total. Audio cannot be the only reference modality.

Pricing

fal charges $0.08 per generated second. Reference inputs share a 4,096-token free allowance; usage above the allowance costs $0.02 per 1,000 tokens.

Technical Specs

  • Model ID: minimax/h3-max/reference-to-video
  • Required input: prompt
  • Duration: 5–15 seconds; default 5
  • Resolution: 480P or 768P; default 768P
  • References: images, videos, and audio; maximum 12 files combined
  • Reference video/audio duration: 2–15 seconds each; maximum 15 seconds combined per modality
  • Output: downloadable video URL
  • Execution: asynchronous

Related Models