KwaiVGI modelsimage generation api

kwaivgi/kling-video/o1/standard/edit

Kling Video O1 Standard by KwaiVGI - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.

Input
Use @Element1, @Element2 to reference elements and @Image1, @Image2 to reference images in order.

PNG, JPEG, WebP, or GIF · 20 MiB maximum each

Reference images for style/appearance. Reference in prompt as @Image1, @Image2, etc. Maximum 4 total (elements + reference images) when using video.
Reference video URL. Only .mp4/.mov formats supported, 3-10 seconds duration, 720-2160px resolution, max 200MB. Max file size: 200.0MB, Min width: 720px, Min height: 720px, Max width: 2160px, Max height: 2160px, Min duration: 3.0s, Max duration: 10.05s, Min FPS: 24.0, Max FPS: 60.0, Timeout: 30.0s
Whether to keep the original audio from the video.
Idle

Example output — click Run to generate your own

API README

Kling Video O1 Standard Edit

edit in the Kling Video lineup is built to revise an existing visual sequence from written direction. The kling video o1 standard edit configuration combines that transformation with the quality and motion profile represented by this exact model version, so teams can choose it deliberately among adjacent family variants. It is useful when the input already establishes part of the creative intent and the model must supply a polished result rather than a generic media conversion, with subject identity, scene logic, visual hierarchy, and the delivery goal kept explicit.

For production work with kwaivgi/kling-video/o1/standard/edit, begin with the non-negotiable content, then describe the desired change, framing, action, atmosphere, and finishing cues in that order. Separate what must remain recognizable from what may vary, and prefer concrete nouns and observable actions over abstract praise. That structure makes outputs easier to compare across storyboard passes, campaign variants, catalog assets, and other repeatable creative pipelines.

Highlights

Instruction-led video revision. Applies written changes to a sequence without regenerating its performance.

Temporal context preservation. Keeps useful timing, camera motion, and unaffected scene content recognizable.

Localized creative change. Targets subject appearance, environment, mood, or visual treatment.

kling video o1 standard edit consistent edited playback. Carries the requested change across frames rather than making isolated repairs. This is the defining creative strength of the kling video o1 standard edit configuration.

Pricing

DurationPrice per generated clip
Per generated second$0.1260

Billing follows params.duration * 0.126: multiply generated seconds by the $0.1260 per-second rate.

When to Use

✅ Good fit❌ Consider alternatives
The model's named workflow matches the source material and intended outputA different input modality or model route is required
A managed asynchronous result is suitable for the production pipelineA synchronous, interactive editor is essential
The documented controls cover the required duration, framing, or formatThe project needs controls outside this endpoint's schema
Creative iteration benefits from a repeatable request structureExact deterministic pixels, frames, geometry, or samples are mandatory
A finished downloadable media asset is the desired deliverableEditable source layers or a native project file are required

Prompt Guide

For generation, state the intended result first, then add the subject or source treatment, progression, style, and delivery constraints. Keep one creative variable per phrase, use the documented field names for controls, and change one setting at a time when comparing results.

{
  "images": [
    "https://static.sandbase.ai/examples/mirrored/e24d46d32342-MKvhFko5_wYnfORYacNII_AgPt8v25Wt4oyKhjnhVK5.png"
  ],
  "prompt": "Replace the character in the video with @Element1, maintaining the same movements and camera angles. Transform the landscape into @Image1",
  "video": "https://static.sandbase.ai/examples/mirrored/2038bf1f9c5a-ku8_Wdpf-oTbGRq4lB5DU_output.mp4"
}

Technical Specs

SpecValue
Model IDkwaivgi/kling-video/o1/standard/edit
Inputsimages, keep_audio, prompt, video
Required inputsprompt
Output fieldscontent_type, url
ExecutionAsync (submit, then poll for result)

Related Models

Related Models

kwaivgi/kling-video/o1/standard/video-to-videoKling Video O1 Standard is KwaiVGI's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.kwaivgi/kling-video/o1/pro/video-to-videoKling Video O1 Pro by KwaiVGI - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.kwaivgi/kling-video/o1/pro/editKling Video O1 Pro is KwaiVGI's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.kwaivgi/kling-video/create-voiceKling Video Create Voice by KwaiVGI - advanced AI model for audio-to-audio. Delivers high-quality results with fast inference, suitable for both creative and production workflows.kwaivgi/kling-video/o3/pro/editKling Video O3 Pro is KwaiVGI's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.kwaivgi/kling-video/o3/pro/video-to-videoKling Video O3 Pro by KwaiVGI - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.kwaivgi/kling-video/o3/standard/editKling Video O3 Standard by KwaiVGI - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.kwaivgi/kling-video/o3/standard/video-to-videoKling Video O3 Standard is KwaiVGI's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.