MiniMax modelsvideo generation api

minimax/h3-max/text-to-video

H3 Max by MiniMax - generate cinematic videos from text descriptions with AI. Create high-quality video content with natural motion, camera control, and optional audio generation.

Input
Text prompt for video generation
The aspect ratio of the generated image. Allowed values: 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16.
The native generation resolution of the video. Allowed values: 480P, 768P.
515
The duration of the video in seconds. Range: 5 to 15.
Random seed. A random seed is selected when omitted.
How much effort to spend rewriting the prompt before generation. 'balanced' returns in about a second. 'quality' spends up to ~30s on a richer prompt.
Idle

Example output — click Run to generate your own

API README

MiniMax H3 Max Text to Video

Generate cinematic videos directly from text with MiniMax H3 Max, a fal post-trained H3 variant optimized for stronger prompt adherence, visual quality, and inference throughput.

The endpoint supports 5–15 second video at 480P or 768P, ten aspect ratios, deterministic seeds, and balanced or quality prompt expansion.

Pricing

Current fal promotional rates are $0.0125 per second at 480P and $0.02 per second at 768P through September 7, 2026. Completed-task duration is used when available; otherwise the requested duration is used.

Technical Specs

  • Model ID: minimax/h3-max/text-to-video
  • Required input: prompt
  • Duration: 5–15 seconds; default 5
  • Resolution: 480P or 768P; default 768P
  • Aspect ratios: 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16
  • Prompt expansion: balanced or quality
  • Output: downloadable video URL
  • Execution: asynchronous

Related Models