minimax/h3-max/reference-to-video
H3 Max is MiniMax's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.
Example output — click Run to generate your own
Send your first request
OpenAI-compatible endpoint with unified authentication and usage tracking.
https://api.sandbase.ai/v1/runminimax/h3-max/reference-to-videoInput Schema
9 parameters · 1 required · 8 optional
| Parameter | Type | Required | Description |
|---|---|---|---|
prompt | string | Required | Text prompt for video generation. Refer to reference assets by their modality and order in the reference lists: Image 1, Image 2, Video 1, Audio 1, and so on. · Min length: 1 · Max length: 50000 |
seed | integer | Optional | Random seed. A random seed is selected when omitted. |
duration | integer | Optional | The duration of the video in seconds. · Min: 5 · Max: 15 · Default: 5 |
resolution | string | Optional | The native generation resolution of the video. · Options: 480P, 768P · Default: "768P" 480P768P |
aspect_ratio | string | Optional | The aspect ratio of the generated image. · Options: 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16 21:916:93:24:35:41:14:53:42:39:16 |
reference_audio_urls | string[] | Optional | URLs of reference audio clips (2-15 seconds each, combined duration at most 15 seconds), referenced in the prompt as Audio 1, Audio 2, and so on. Audio cannot be the only reference input; provide at least one reference image or video with it. Reference images, videos, and audio clips must add up to at most 12 files. |
reference_image_urls | string[] | Optional | URLs of subject/style reference images, referenced in the prompt as Image 1, Image 2, and so on. Reference images, videos, and audio clips must add up to at most 12 files. |
reference_video_urls | string[] | Optional | URLs of motion/reference video clips (2-15 seconds each, combined duration at most 15 seconds), referenced in the prompt as Video 1, Video 2, and so on. Reference images, videos, and audio clips must add up to at most 12 files. |
prompt_expansion_mode | string | Optional | How much effort to spend rewriting the prompt before generation. 'balanced' returns in about a second. 'quality' spends up to ~30s on a richer prompt. · Default: "balanced" |
Output Schema
| Field | Type | Description |
|---|---|---|
id | string | Unique identifier for the generation task |
status | string | Task status: pending, running, completed, failed, timeout |
model | string | Model used for the generation |
outputs | array | Array of output items |
outputs[].url | string | URL of the generated artifact |
outputs[].content_type | string | MIME type (e.g. image/png, video/mp4) |
error | object | null | Error details if failed, null on success |
error.type | string | Machine-readable error type code |
error.message | string | Human-readable error description |
Async Workflow
This model uses asynchronous execution. Submit a request and poll for the result.
- Submit — POST to /v1/run, receive an
id - Poll — GET /v1/run/{id} until status is
completed,failed, ortimeout - Retrieve — Read
outputsfrom the completed response
Code Examples
Ready-to-run snippets
# Step 1: Submit
curl -X POST https://api.sandbase.ai/v1/run \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "minimax/h3-max/reference-to-video",
"prompt": "Image 1 is the female protagonist. Image 2 is her small dog. Keep the woman and dog consistent with their respective reference images while they walk together through a sunlit garden.",
"duration": 5,
"resolution": "768P",
"reference_image_urls": [
"https://storage.googleapis.com/falserverless/example_inputs/hailuo23/pro_i2v_in.jpg"
],
"prompt_expansion_mode": "balanced"
}'
# Step 2: Poll result (replace <id>)
curl https://api.sandbase.ai/v1/run/<id> \
-H "Authorization: Bearer YOUR_API_KEY"API README
MiniMax H3 Max Reference to Video
Generate video from a text prompt plus image, video, and audio references with MiniMax H3 Max. Reference assets can establish subject identity, appearance, motion, timing, dialogue, ambience, and style.
The endpoint supports 5–15 second output at 480P or 768P, adaptive or fixed aspect ratios, and up to 12 reference files in total. Audio cannot be the only reference modality.
Pricing
fal charges $0.08 per generated second. Reference inputs share a 4,096-token free allowance; usage above the allowance costs $0.02 per 1,000 tokens.
Technical Specs
- Model ID:
minimax/h3-max/reference-to-video - Required input:
prompt - Duration: 5–15 seconds; default 5
- Resolution:
480Por768P; default768P - References: images, videos, and audio; maximum 12 files combined
- Reference video/audio duration: 2–15 seconds each; maximum 15 seconds combined per modality
- Output: downloadable video URL
- Execution: asynchronous

