KwaiVGI modelsvideo generation api

kwaivgi/kling-video/o3

Kling Video O3 is KwaiVGI's unified omni video generation model. Generate or transform videos from prompts, images, reference videos, and multi-shot text while choosing standard or pro mode per request.

Input
Text prompt for video generation or transformation. When using reference media, refer to them as @Image1, @Image2, and @Video1.
Generation quality tier. Allowed values: standard, pro, 4k.

PNG, JPEG, WebP, or GIF · 20 MiB maximum

Optional first-frame image URL for image-to-video.
Optional reference video URL for video-to-video generation.
Allowed values: image, video.

PNG, JPEG, WebP, or GIF · 20 MiB maximum

Optional end-frame image URL for image-to-video.

PNG, JPEG, WebP, or GIF · 20 MiB maximum each

Optional reference image URLs.
Generated video duration in seconds. Allowed values: 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15.
Allowed values: 16:9, 9:16, 1:1.
Allowed values: customize, intelligent.
Enter a JSON array.
Idle

Example output — click Run to generate your own

API README

Kling Video O3

Kling Video O3 is a unified omni-video model that can create, transform, and continue video from combinations of text, images, and source footage. It is intended for workflows that need one model to reason across characters, scenes, motion references, camera direction, and sound rather than selecting a separate endpoint for every creative operation.

The model supports multi-shot storytelling and reference-led generation while maintaining characters and visual elements across a sequence. It can preserve or regenerate source audio, create synchronized sound, and use image or video orientation to guide identity and performance, making it suitable for narrative scenes, branded characters, and complex audiovisual edits.

Highlights

Unified multimodal video creation. Combines text, images, and video references within one generation and transformation model.

Reference consistency. Carries characters, objects, and visual identity across new actions, viewpoints, and scenes.

Multi-shot storytelling. Produces coherent sequences with multiple shots, camera changes, and structured scene progression.

Integrated audiovisual control. Generates synchronized sound or retains source audio while coordinating it with the new visual sequence.

Pricing

ConfigurationModePrice
3sStandard, no source video, no generated audio$0.252
3sStandard, no source video, generated audio$0.336
3sPro, no source video, no generated audio$0.336
3sPro, no source video, generated audio$0.420
3sStandard with source video$0.378
3sPro with source video$0.504
3s4K$1.260
4sStandard, no source video, no generated audio$0.336
4sStandard, no source video, generated audio$0.448
4sPro, no source video, no generated audio$0.448
4sPro, no source video, generated audio$0.560
4sStandard with source video$0.504
4sPro with source video$0.672
4s4K$1.680
5sStandard, no source video, no generated audio$0.420
5sStandard, no source video, generated audio$0.560
5sPro, no source video, no generated audio$0.560
5sPro, no source video, generated audio$0.700
5sStandard with source video$0.630
5sPro with source video$0.840
5s4K$2.100
6sStandard, no source video, no generated audio$0.504
6sStandard, no source video, generated audio$0.672
6sPro, no source video, no generated audio$0.672
6sPro, no source video, generated audio$0.840
6sStandard with source video$0.756
6sPro with source video$1.008
6s4K$2.520
7sStandard, no source video, no generated audio$0.588
7sStandard, no source video, generated audio$0.784
7sPro, no source video, no generated audio$0.784
7sPro, no source video, generated audio$0.980
7sStandard with source video$0.882
7sPro with source video$1.176
7s4K$2.940
8sStandard, no source video, no generated audio$0.672
8sStandard, no source video, generated audio$0.896
8sPro, no source video, no generated audio$0.896
8sPro, no source video, generated audio$1.120
8sStandard with source video$1.008
8sPro with source video$1.344
8s4K$3.360
9sStandard, no source video, no generated audio$0.756
9sStandard, no source video, generated audio$1.008
9sPro, no source video, no generated audio$1.008
9sPro, no source video, generated audio$1.260
9sStandard with source video$1.134
9sPro with source video$1.512
9s4K$3.780
10sStandard, no source video, no generated audio$0.840
10sStandard, no source video, generated audio$1.120
10sPro, no source video, no generated audio$1.120
10sPro, no source video, generated audio$1.400
10sStandard with source video$1.260
10sPro with source video$1.680
10s4K$4.200
11sStandard, no source video, no generated audio$0.924
11sStandard, no source video, generated audio$1.232
11sPro, no source video, no generated audio$1.232
11sPro, no source video, generated audio$1.540
11sStandard with source video$1.386
11sPro with source video$1.848
11s4K$4.620
12sStandard, no source video, no generated audio$1.008
12sStandard, no source video, generated audio$1.344
12sPro, no source video, no generated audio$1.344
12sPro, no source video, generated audio$1.680
12sStandard with source video$1.512
12sPro with source video$2.016
12s4K$5.040
13sStandard, no source video, no generated audio$1.092
13sStandard, no source video, generated audio$1.456
13sPro, no source video, no generated audio$1.456
13sPro, no source video, generated audio$1.820
13sStandard with source video$1.638
13sPro with source video$2.184
13s4K$5.460
14sStandard, no source video, no generated audio$1.176
14sStandard, no source video, generated audio$1.568
14sPro, no source video, no generated audio$1.568
14sPro, no source video, generated audio$1.960
14sStandard with source video$1.764
14sPro with source video$2.352
14s4K$5.880
15sStandard, no source video, no generated audio$1.260
15sStandard, no source video, generated audio$1.680
15sPro, no source video, no generated audio$1.680
15sPro, no source video, generated audio$2.100
15sStandard with source video$1.890
15sPro with source video$2.520
15s4K$6.300

When to Use

✅ Good fit❌ Consider alternatives
The model's named workflow matches the source material and intended outputA different input modality or model route is required
A managed asynchronous result is suitable for the production pipelineA synchronous, interactive editor is essential
The documented controls cover the required duration, framing, or formatThe project needs controls outside this endpoint's schema
Creative iteration benefits from a repeatable request structureExact deterministic pixels, frames, geometry, or samples are mandatory
A finished downloadable media asset is the desired deliverableEditable source layers or a native project file are required

Prompt Guide

For image-conditioned generation, state the intended result first, then add the subject or source treatment, progression, style, and delivery constraints. Keep one creative variable per phrase, use the documented field names for controls, and change one setting at a time when comparing results.

{
  "aspect_ratio": "16:9",
  "duration": 5,
  "image": "https://static.sandbase.ai/examples/mirrored/f77b827ca02a-TNErq9yD7ZxGRATjfAqnh_EIgJSN67.png",
  "images": [
    "https://example.com/reference.png"
  ],
  "prompt": "Based on @Video1, make the character from @Image1 dance.",
  "resolution": "standard",
  "video": "https://static.sandbase.ai/examples/mirrored/f1ecbb471a16-hklvF__w53diz6Rve7f5__JuDW2xl0mr6sJ_Kjz3Vxe_vidoeook--1-_1.mp4"
}

Technical Specs

SpecValue
Model IDkwaivgi/kling-video/o3
Inputsaspect_ratio, character_orientation, duration, end_image, generate_audio, image, images, keep_original_sound, multi_prompt, prompt, resolution, shot_type, video
Required inputsprompt
Output fieldscontent_type, url
ExecutionAsync (submit, then poll for result)
Duration3 / 4 / 5 / 6 / 7 / 8 / 9 / 10 / 11 / 12 / 13 / 14 / 15
Resolutionstandard / pro / 4k
Aspect Ratio16:9 / 9:16 / 1:1

Related Models

Related Models

kwaivgi/kling-video/o3/4k/image-to-videoKling Video O3 4k by KwaiVGI - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.kwaivgi/kling-video/o3/4k/reference-to-videoKling Video O3 4k by KwaiVGI - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.kwaivgi/kling-video/o3/4k/text-to-videoKling Video O3 4k is KwaiVGI's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.kwaivgi/kling-video/o3/pro/image-to-videoKling Video O3 Pro is KwaiVGI's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.kwaivgi/kling-video/o3/pro/reference-to-videoKling Video O3 Pro is KwaiVGI's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.kwaivgi/kling-video/o3/pro/text-to-videoKling Video O3 Pro by KwaiVGI - generate cinematic videos from text descriptions with AI. Create high-quality video content with natural motion, camera control, and optional audio generation.kwaivgi/kling-video/o3/standard/image-to-videoKling Video O3 Standard by KwaiVGI - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.kwaivgi/kling-video/o3/standard/reference-to-videoKling Video O3 Standard by KwaiVGI - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.