SandBase is live — $1 in free credits on signupStart free ›

KwaiVGI modelsvideo generation api

kwaivgi/kling-video/o3/standard/image-to-video

Kling Video O3 Standard by KwaiVGI - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.

Input
Text prompt for video generation. Either prompt or multi_prompt must be provided, but not both.

PNG, JPEG, WebP, or GIF · 20 MiB maximum

URL of the start frame image.

PNG, JPEG, WebP, or GIF · 20 MiB maximum

URL of the end frame image (optional).
Video duration in seconds (3-15s). Allowed values: 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15.
Whether to generate native audio for the video.
List of prompts for multi-shot video generation.
The type of multi-shot video generation. 'intelligent' lets the model automatically determine shot structure. Allowed values: customize, intelligent.
Idle

Example output — click Run to generate your own

API README

Kling Video O3 Standard

Kling Video O3 Standard is a image-led video-generation route in the Kling Video O3 family. It turns written scene direction and a starting image into a multi-second sequence with controlled subject action, camera movement, shot structure, and optional synchronized sound.

The Standard route is intended for creators who need dependable temporal continuity across a directed clip. Prompts can describe staging, motion, lens behavior, atmosphere, and audio cues; multi-prompt controls can divide the sequence into several shots while reference inputs keep important visual elements anchored.

Highlights

  • Image-anchored animation. Uses the supplied opening image to preserve composition and identity as motion develops.
  • Multi-shot direction. Supports structured prompt segments for sequences that need more than one planned shot.
  • Native audio option. Can generate sound together with the visible action when the route exposes audio generation.
  • Efficient standard output. Balances coherent O3 motion and narrative control with the Standard route pricing.

Pricing

ConfigurationBilling unitPrice
3 secondsWithout generated audio$0.252
3 secondsWith generated audio$0.336
4 secondsWithout generated audio$0.336
4 secondsWith generated audio$0.448
5 secondsWithout generated audio$0.420
5 secondsWith generated audio$0.560
6 secondsWithout generated audio$0.504
6 secondsWith generated audio$0.672
7 secondsWithout generated audio$0.588
7 secondsWith generated audio$0.784
8 secondsWithout generated audio$0.672
8 secondsWith generated audio$0.896
9 secondsWithout generated audio$0.756
9 secondsWith generated audio$1.008
10 secondsWithout generated audio$0.840
10 secondsWith generated audio$1.120
11 secondsWithout generated audio$0.924
11 secondsWith generated audio$1.232
12 secondsWithout generated audio$1.008
12 secondsWith generated audio$1.344
13 secondsWithout generated audio$1.092
13 secondsWith generated audio$1.456
14 secondsWithout generated audio$1.176
14 secondsWith generated audio$1.568
15 secondsWithout generated audio$1.260
15 secondsWith generated audio$1.680

When to Use

✅ Good fit❌ Consider alternatives
The project needs creating a directed video sequenceThe goal is a different media task or endpoint
The available inputs match the required local schemaRequired source media or permissions are unavailable
The brief can specify subject, composition, style, and deliveryThe result must be deterministic at pixel or sample level
The supported formats and controls match final placementDelivery requires unsupported dimensions, codecs, or duration
An asynchronous generated result fits the workflowA live, frame-synchronous, or real-time response is mandatory

Prompt Guide

Write the request as a production brief: identify the main subject or source material, state the intended transformation, describe composition or timing, and finish with style, atmosphere, and delivery constraints. Keep preservation requirements separate from requested changes, and use only fields exposed by this route.

{
  "prompt": "A cinematic, precisely directed scene with clear subject action, camera movement, lighting, and atmosphere",
  "image": "https://example.com/reference.jpg",
  "duration": 3,
  "end_image": "https://example.com/reference.jpg",
  "shot_type": "customize"
}

Technical Specs

SpecValue
Model IDkwaivgi/kling-video/o3/standard/image-to-video
Input fieldsimage (string)<br>prompt (string)<br>duration (integer; 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15)<br>end_image (string)<br>shot_type (string; customize, intelligent)<br>multi_prompt (array)<br>generate_audio (boolean)
Required inputprompt, image
Output fieldsurl, content_type
ExecutionAsynchronous job

Related Models

  • kwaivgi/kling-video/o3/pro/image-to-video
  • kwaivgi/kling-video/o3/standard/reference-to-video
  • kwaivgi/kling-video/o3/standard/text-to-video

Related Models

kwaivgi/kling-video/o3/standard/reference-to-videoKling Video O3 Standard by KwaiVGI - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.kwaivgi/kling-video/o3/standard/text-to-videoKling Video O3 Standard is KwaiVGI's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.kwaivgi/kling-video/o3/4k/reference-to-videoKling Video O3 4k by KwaiVGI - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.kwaivgi/kling-video/o3/4k/text-to-videoKling Video O3 4k is KwaiVGI's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.kwaivgi/kling-video/o3/pro/image-to-videoKling Video O3 Pro is KwaiVGI's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.kwaivgi/kling-video/o3/pro/reference-to-videoKling Video O3 Pro is KwaiVGI's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.kwaivgi/kling-video/o3/pro/text-to-videoKling Video O3 Pro by KwaiVGI - generate cinematic videos from text descriptions with AI. Create high-quality video content with natural motion, camera control, and optional audio generation.kwaivgi/kling-video/o3Kling Video O3 is KwaiVGI's unified omni video generation model. Generate or transform videos from prompts, images, reference videos, and multi-shot text while choosing standard or pro mode per request.