SandBase is live — $1 in free credits on signupStart free ›

KwaiVGI modelsvideo generation api

kwaivgi/kling-video/v3/pro/image-to-video

Kling Video V3 Pro is KwaiVGI's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.

Input
Text prompt for video generation. Either prompt or multi_prompt must be provided, but not both.

PNG, JPEG, WebP, or GIF · 20 MiB maximum

URL of the image to be used for the video

PNG, JPEG, WebP, or GIF · 20 MiB maximum

URL of the image to be used for the end of the video
The duration of the generated video in seconds Allowed values: 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15.
Whether to generate native audio for the video. Supports Chinese and English voice output. Other languages are automatically translated to English. For English speech, use lowercase letters; for acronyms or proper nouns, use uppercase.
The type of multi-shot video generation. 'intelligent' lets the model automatically determine shot structure. Allowed values: customize, intelligent.
List of prompts for multi-shot video generation. If provided, divides the video into multiple shots.
Idle

Example output — click Run to generate your own

API README

Kling Video V3 Pro Image-to-Video

Kling Video V3 Pro animates a supplied opening image from a written direction. The image fixes the first frame while the prompt describes action, camera behavior, and changes over time. An optional end image can guide the final frame.

The route produces clips from 3 to 15 seconds and returns a video URL. A request can select custom or intelligent shot structure, and native audio is enabled by default. The opening image determines the frame shape because this route has no separate aspect-ratio field.

Highlights

  • Cinematic visual quality. The exact Pro image-to-video route is documented for cinematic visuals. This is the route's stated visual positioning.
  • Fluid motion. Fluid motion is explicitly identified for this Pro image animation route. Movement is a model capability rather than a request parameter.
  • Native audio generation. The exact route is documented to generate native audio with video. Audio belongs to the generated result rather than a separate post-process.
  • Strong subject and text consistency. The exact Pro page identifies strong subject and text consistency. This capability applies while the opening image is animated.

Pricing

DurationPrice
3 seconds$0.504
4 seconds$0.672
5 seconds$0.840
6 seconds$1.008
7 seconds$1.176
8 seconds$1.344
9 seconds$1.512
10 seconds$1.680
11 seconds$1.848
12 seconds$2.016
13 seconds$2.184
14 seconds$2.352
15 seconds$2.520

When to Use

✅ Good fit❌ Consider alternatives
Add movement to an existing product or character still.Begin without any opening image; use text-to-video.
Direct a camera move around the composition in the first frame.Select frame shape independently of the supplied image.
Guide a transition toward a supplied final frame.Require a clip longer than 15 seconds in one request.
Produce a short scene with generated sound enabled.Require voice behavior outside the documented language handling.
Use the Pro tier for this image-led workflow.Prefer the lower local price of the Standard sibling.

Prompt Guide

Describe how the visible scene should move. Keep the request to the required image and prompt fields first, then add duration, audio, end-frame, or shot settings only when needed.

Subject action: [movement over time]
Camera: [framing and camera path]
Environment: [light, wind, particles, background movement]
Timing: [pace and progression]
Audio: [ambience, effects, or speech when enabled]
Final state: [arrival at the optional end image]
{
  "image": "<start-image-url>",
  "prompt": "The craftsman slowly examines the bowl while warm light shifts and dust drifts through the air.",
  "duration": 5,
  "generate_audio": true
}

Technical Specs

SpecValue
Required inputsimage, prompt
Prompt lengthUp to 2,500 characters
Duration3–15 seconds; default 5
Frame guidanceRequired image; optional end_image
Shot typescustomize, intelligent
Multi-prompt fieldExposed locally, but prompt remains required by the request contract
Native audioOptional; default enabled
OutputVideo URL with optional file metadata

Related

Related Models

kwaivgi/kling-video/v3/pro/text-to-videoKling Video V3 Pro by KwaiVGI - generate cinematic videos from text descriptions with AI. Create high-quality video content with natural motion, camera control, and optional audio generation.kwaivgi/kling-video/v3/4k/image-to-videoKling Video V3 4k by KwaiVGI - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.kwaivgi/kling-video/v3/4k/text-to-videoKling Video V3 4k is KwaiVGI's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.kwaivgi/kling-video/v3Kling Video V3 is KwaiVGI's unified video generation model. Generate or transform videos from prompts, images, reference videos, and multi-shot text while choosing standard or pro mode per request.kwaivgi/kling-video/v3/standard/image-to-videoKling Video V3 Standard by KwaiVGI - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.kwaivgi/kling-video/v3/standard/text-to-videoKling Video V3 Standard is KwaiVGI's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.kwaivgi/kling-video/3.0/omni/pro/video-to-video/editModify an existing video with natural-language instructions using Kling 3.0 Omni Pro.kwaivgi/kling-video/3.0/omni/pro/video-to-video/referenceUse a source video plus image references to generate a guided video variation with Kling 3.0 Omni Pro.