kwaivgi/kling-video/o3
Kling Video O3 is KwaiVGI's unified omni video generation model. Generate or transform videos from prompts, images, reference videos, and multi-shot text while choosing standard or pro mode per request.
PNG, JPEG, WebP, or GIF · 20 MiB maximum
PNG, JPEG, WebP, or GIF · 20 MiB maximum
PNG, JPEG, WebP, or GIF · 20 MiB maximum each
Example output — click Run to generate your own
Send your first request
OpenAI-compatible endpoint with unified authentication and usage tracking.
https://api.sandbase.ai/v1/runkwaivgi/kling-video/o3Input Schema
13 parameters · 1 required · 12 optional
| Parameter | Type | Required | Description |
|---|---|---|---|
prompt | string | Required | Text prompt for video generation or transformation. When using reference media, refer to them as @Image1, @Image2, and @Video1. · Max length: 2500 |
image | string | Optional | Optional first-frame image URL for image-to-video. |
video | string | Optional | Optional reference video URL for video-to-video generation. |
images | string[] | Optional | Optional reference image URLs. |
duration | integer | Optional | Generated video duration in seconds. · Options: 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 · Default: 5 3456789101112131415 |
end_image | string | Optional | Optional end-frame image URL for image-to-video. |
shot_type | string | Optional | Options: customize, intelligent · Default: "customize" customizeintelligent |
resolution | string | Optional | Generation quality tier. · Options: standard, pro, 4k · Default: "standard" standardpro4k |
aspect_ratio | string | Optional | Options: 16:9, 9:16, 1:1 · Default: "16:9" 16:99:161:1 |
multi_prompt | object[] | Optional | — |
generate_audio | boolean | Optional | Default: true |
keep_original_sound | boolean | Optional | Default: true |
character_orientation | string | Optional | Options: image, video imagevideo |
Output Schema
| Field | Type | Description |
|---|---|---|
id | string | Unique identifier for the generation task |
status | string | Task status: pending, running, completed, failed, timeout |
model | string | Model used for the generation |
outputs | array | Array of output items |
outputs[].url | string | URL of the generated artifact |
outputs[].content_type | string | MIME type (e.g. image/png, video/mp4) |
error | object | null | Error details if failed, null on success |
error.type | string | Machine-readable error type code |
error.message | string | Human-readable error description |
Async Workflow
This model uses asynchronous execution. Submit a request and poll for the result.
- Submit — POST to /v1/run, receive an
id - Poll — GET /v1/run/{id} until status is
completed,failed, ortimeout - Retrieve — Read
outputsfrom the completed response
Code Examples
Ready-to-run snippets
# Step 1: Submit
curl -X POST https://api.sandbase.ai/v1/run \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "kwaivgi/kling-video/o3",
"image": "https://static.sandbase.ai/examples/mirrored/f77b827ca02a-TNErq9yD7ZxGRATjfAqnh_EIgJSN67.png",
"video": "https://static.sandbase.ai/examples/mirrored/f1ecbb471a16-hklvF__w53diz6Rve7f5__JuDW2xl0mr6sJ_Kjz3Vxe_vidoeook--1-_1.mp4",
"prompt": "Based on @Video1, make the character from @Image1 dance.",
"duration": 5,
"shot_type": "customize",
"resolution": "standard",
"aspect_ratio": "16:9",
"generate_audio": true,
"keep_original_sound": true,
"character_orientation": "image"
}'
# Step 2: Poll result (replace <id>)
curl https://api.sandbase.ai/v1/run/<id> \
-H "Authorization: Bearer YOUR_API_KEY"API README
Kling Video O3
Kling Video O3 is a unified omni-video model that can create, transform, and continue video from combinations of text, images, and source footage. It is intended for workflows that need one model to reason across characters, scenes, motion references, camera direction, and sound rather than selecting a separate endpoint for every creative operation.
The model supports multi-shot storytelling and reference-led generation while maintaining characters and visual elements across a sequence. It can preserve or regenerate source audio, create synchronized sound, and use image or video orientation to guide identity and performance, making it suitable for narrative scenes, branded characters, and complex audiovisual edits.
Highlights
Unified multimodal video creation. Combines text, images, and video references within one generation and transformation model.
Reference consistency. Carries characters, objects, and visual identity across new actions, viewpoints, and scenes.
Multi-shot storytelling. Produces coherent sequences with multiple shots, camera changes, and structured scene progression.
Integrated audiovisual control. Generates synchronized sound or retains source audio while coordinating it with the new visual sequence.
Pricing
| Configuration | Mode | Price |
|---|---|---|
| 3s | Standard, no source video, no generated audio | $0.252 |
| 3s | Standard, no source video, generated audio | $0.336 |
| 3s | Pro, no source video, no generated audio | $0.336 |
| 3s | Pro, no source video, generated audio | $0.420 |
| 3s | Standard with source video | $0.378 |
| 3s | Pro with source video | $0.504 |
| 3s | 4K | $1.260 |
| 4s | Standard, no source video, no generated audio | $0.336 |
| 4s | Standard, no source video, generated audio | $0.448 |
| 4s | Pro, no source video, no generated audio | $0.448 |
| 4s | Pro, no source video, generated audio | $0.560 |
| 4s | Standard with source video | $0.504 |
| 4s | Pro with source video | $0.672 |
| 4s | 4K | $1.680 |
| 5s | Standard, no source video, no generated audio | $0.420 |
| 5s | Standard, no source video, generated audio | $0.560 |
| 5s | Pro, no source video, no generated audio | $0.560 |
| 5s | Pro, no source video, generated audio | $0.700 |
| 5s | Standard with source video | $0.630 |
| 5s | Pro with source video | $0.840 |
| 5s | 4K | $2.100 |
| 6s | Standard, no source video, no generated audio | $0.504 |
| 6s | Standard, no source video, generated audio | $0.672 |
| 6s | Pro, no source video, no generated audio | $0.672 |
| 6s | Pro, no source video, generated audio | $0.840 |
| 6s | Standard with source video | $0.756 |
| 6s | Pro with source video | $1.008 |
| 6s | 4K | $2.520 |
| 7s | Standard, no source video, no generated audio | $0.588 |
| 7s | Standard, no source video, generated audio | $0.784 |
| 7s | Pro, no source video, no generated audio | $0.784 |
| 7s | Pro, no source video, generated audio | $0.980 |
| 7s | Standard with source video | $0.882 |
| 7s | Pro with source video | $1.176 |
| 7s | 4K | $2.940 |
| 8s | Standard, no source video, no generated audio | $0.672 |
| 8s | Standard, no source video, generated audio | $0.896 |
| 8s | Pro, no source video, no generated audio | $0.896 |
| 8s | Pro, no source video, generated audio | $1.120 |
| 8s | Standard with source video | $1.008 |
| 8s | Pro with source video | $1.344 |
| 8s | 4K | $3.360 |
| 9s | Standard, no source video, no generated audio | $0.756 |
| 9s | Standard, no source video, generated audio | $1.008 |
| 9s | Pro, no source video, no generated audio | $1.008 |
| 9s | Pro, no source video, generated audio | $1.260 |
| 9s | Standard with source video | $1.134 |
| 9s | Pro with source video | $1.512 |
| 9s | 4K | $3.780 |
| 10s | Standard, no source video, no generated audio | $0.840 |
| 10s | Standard, no source video, generated audio | $1.120 |
| 10s | Pro, no source video, no generated audio | $1.120 |
| 10s | Pro, no source video, generated audio | $1.400 |
| 10s | Standard with source video | $1.260 |
| 10s | Pro with source video | $1.680 |
| 10s | 4K | $4.200 |
| 11s | Standard, no source video, no generated audio | $0.924 |
| 11s | Standard, no source video, generated audio | $1.232 |
| 11s | Pro, no source video, no generated audio | $1.232 |
| 11s | Pro, no source video, generated audio | $1.540 |
| 11s | Standard with source video | $1.386 |
| 11s | Pro with source video | $1.848 |
| 11s | 4K | $4.620 |
| 12s | Standard, no source video, no generated audio | $1.008 |
| 12s | Standard, no source video, generated audio | $1.344 |
| 12s | Pro, no source video, no generated audio | $1.344 |
| 12s | Pro, no source video, generated audio | $1.680 |
| 12s | Standard with source video | $1.512 |
| 12s | Pro with source video | $2.016 |
| 12s | 4K | $5.040 |
| 13s | Standard, no source video, no generated audio | $1.092 |
| 13s | Standard, no source video, generated audio | $1.456 |
| 13s | Pro, no source video, no generated audio | $1.456 |
| 13s | Pro, no source video, generated audio | $1.820 |
| 13s | Standard with source video | $1.638 |
| 13s | Pro with source video | $2.184 |
| 13s | 4K | $5.460 |
| 14s | Standard, no source video, no generated audio | $1.176 |
| 14s | Standard, no source video, generated audio | $1.568 |
| 14s | Pro, no source video, no generated audio | $1.568 |
| 14s | Pro, no source video, generated audio | $1.960 |
| 14s | Standard with source video | $1.764 |
| 14s | Pro with source video | $2.352 |
| 14s | 4K | $5.880 |
| 15s | Standard, no source video, no generated audio | $1.260 |
| 15s | Standard, no source video, generated audio | $1.680 |
| 15s | Pro, no source video, no generated audio | $1.680 |
| 15s | Pro, no source video, generated audio | $2.100 |
| 15s | Standard with source video | $1.890 |
| 15s | Pro with source video | $2.520 |
| 15s | 4K | $6.300 |
When to Use
| ✅ Good fit | ❌ Consider alternatives |
|---|---|
| The model's named workflow matches the source material and intended output | A different input modality or model route is required |
| A managed asynchronous result is suitable for the production pipeline | A synchronous, interactive editor is essential |
| The documented controls cover the required duration, framing, or format | The project needs controls outside this endpoint's schema |
| Creative iteration benefits from a repeatable request structure | Exact deterministic pixels, frames, geometry, or samples are mandatory |
| A finished downloadable media asset is the desired deliverable | Editable source layers or a native project file are required |
Prompt Guide
For image-conditioned generation, state the intended result first, then add the subject or source treatment, progression, style, and delivery constraints. Keep one creative variable per phrase, use the documented field names for controls, and change one setting at a time when comparing results.
{
"aspect_ratio": "16:9",
"duration": 5,
"image": "https://static.sandbase.ai/examples/mirrored/f77b827ca02a-TNErq9yD7ZxGRATjfAqnh_EIgJSN67.png",
"images": [
"https://example.com/reference.png"
],
"prompt": "Based on @Video1, make the character from @Image1 dance.",
"resolution": "standard",
"video": "https://static.sandbase.ai/examples/mirrored/f1ecbb471a16-hklvF__w53diz6Rve7f5__JuDW2xl0mr6sJ_Kjz3Vxe_vidoeook--1-_1.mp4"
}
Technical Specs
| Spec | Value |
|---|---|
| Model ID | kwaivgi/kling-video/o3 |
| Inputs | aspect_ratio, character_orientation, duration, end_image, generate_audio, image, images, keep_original_sound, multi_prompt, prompt, resolution, shot_type, video |
| Required inputs | prompt |
| Output fields | content_type, url |
| Execution | Async (submit, then poll for result) |
| Duration | 3 / 4 / 5 / 6 / 7 / 8 / 9 / 10 / 11 / 12 / 13 / 14 / 15 |
| Resolution | standard / pro / 4k |
| Aspect Ratio | 16:9 / 9:16 / 1:1 |
Related Models
kwaivgi/kling-video/o3/4k/image-to-video— Compare a nearby route in the same local model family.kwaivgi/kling-video/o3/4k/reference-to-video— Compare a nearby route in the same local model family.kwaivgi/kling-video/o3/4k/text-to-video— Compare a nearby route in the same local model family.
