SandBase is live — $1 in free credits on signupStart free ›

KwaiVGI modelsimage generation api

kwaivgi/kolors/image-to-image

Kolors is KwaiVGI's intelligent image editing model. Transform, retouch, and reimagine existing images using text prompts - from background replacement to artistic style conversion.

Input
The prompt to generate an image from.

PNG, JPEG, WebP, or GIF · 20 MiB maximum

URL of image to use for image to image
110
The CFG (Classifier Free Guidance) scale is a measure of how close you want the model to stick to your prompt when looking for a related image to show you. Range: 1 to 10.
1150
The number of inference steps to perform. Range: 1 to 150.
0.011
The strength to use for image-to-image. 1.0 is completely remakes the image while 0.0 preserves the original. Range: 0.01 to 1.
The format of the generated image. Allowed values: jpeg, png.
The scheduler to use for the model. Allowed values: EulerDiscreteScheduler, EulerAncestralDiscreteScheduler, DPMSolverMultistepScheduler, DPMSolverMultistepScheduler_SDE_karras, UniPCMultistepScheduler, DEISMultistepScheduler.
The aspect ratio of the generated image. Allowed values: 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16.
Seed
Idle

Example output — click Run to generate your own

API README

Kolors Image to Image

Kolors Image to Image is a image-guided restyling endpoint in the Kolors family. It is built for creators who need to turn a concrete creative brief into a controlled visual sequence: the request establishes the source material, the intended subject behavior, the camera language, and the atmosphere of the finished shot. The standard route keeps that workflow explicit instead of hiding its input assumptions behind a generic video-generation label.

In practice, this route accepts image as creative context and exposes aspect ratio for delivery planning. That makes it suitable for shot-based pipelines where teams must preserve a source, direct a transformation, or control the final format without losing sight of the model's central task. Write prompts as a compact shot plan—subject, action, setting, camera, light, and timing—then use the structured fields for constraints that should remain deterministic across iterations.

Highlights

  • Bilingual visual semantics. Kolors was developed for strong Chinese and English prompt understanding across photorealistic and stylized imagery.
  • Composition control. A broad aspect-ratio set supports landscape, portrait, square, and social formats.
  • Generation-path control. Scheduler selection and reproducible seeds support intentional iteration.
  • Guided transformation. Image strength balances source preservation against reinterpretation.

Pricing

Billing unitPrice
Base request$0.000575

When to Use

ScenarioRecommendation
Choose this routeUse it when the deliverable specifically calls for image-guided restyling, rather than a neighboring generation mode.
Prepare the sourceProvide image in the format described by the request schema.
Direct the shotDescribe the subject, action, environment, camera movement, lighting, and temporal progression in that order.
Control continuityUse endpoint frames, reference media, strength, or audio controls when those fields are available instead of burying hard constraints in prose.
Plan deliverySet duration, frame count, resolution, and aspect ratio explicitly when the schema exposes them, then compare iterations with a stable seed where supported.

Prompt Guide

For image-guided restyling, describe one coherent shot rather than a list of visual keywords. Put the main subject and action first, follow with location and staging, then add camera movement, lens or framing, lighting, mood, and any timed change. Keep URLs and hard delivery choices in their dedicated fields.

{
  "prompt": "high quality image of a capybara wearing sunglasses. In the background of the image there are trees, poles, grass and other objects. At the bottom of the object there is the road., 8k, highly detailed.",
  "image": "https://static.sandbase.ai/examples/kwaivgi/kolors/image-to-image/input_image_0.png",
  "aspect_ratio": "21:9"
}

Technical Specs

SpecificationValue
Model IDkwaivgi/kolors/image-to-image
Required inputsprompt, image
ExecutionAsynchronous generation job
Request controls9 documented fields
Outputurl, content_type

Request fields

FieldType and constraints
seedinteger; Optional
imagestring; Required
promptstring; Required
strengthnumber; Optional; Range: 0.01–1; Default: 0.85
schedulerstring; Optional; Options: EulerDiscreteScheduler, EulerAncestralDiscreteScheduler, DPMSolverMultistepScheduler, DPMSolverMultistepScheduler_SDE_karras, UniPCMultistepScheduler, DEISMultistepScheduler; Default: "EulerDiscreteScheduler"
aspect_ratiostring; Optional; Options: 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16
output_formatstring; Optional; Options: jpeg, png; Default: "png"
guidance_scalenumber; Optional; Range: 1–10; Default: 5
num_inference_stepsinteger; Optional; Range: 1–150; Default: 50

Related Models

Related Models

kwaivgi/kolorsKolors by KwaiVGI - generate stunning images from text prompts with state-of-the-art AI. Supports multiple aspect ratios, styles, and high-resolution output for creative and commercial use.kwaivgi/kling-image/o3Kling Image O3 by KwaiVGI - generate stunning images from text prompts with state-of-the-art AI. Supports multiple aspect ratios, styles, and high-resolution output for creative and commercial use.kwaivgi/kling-image/o3/editKling Image O3 Edit is KwaiVGI's intelligent image editing model. Transform, retouch, and reimagine existing images using text prompts - from background replacement to artistic style conversion.kwaivgi/kling-image/v3Kling Image V3 by KwaiVGI - generate stunning images from text prompts with state-of-the-art AI. Supports multiple aspect ratios, styles, and high-resolution output for creative and commercial use.kwaivgi/kling-image/v3/editKling Image V3 Edit is KwaiVGI's intelligent image editing model. Transform, retouch, and reimagine existing images using text prompts - from background replacement to artistic style conversion.kwaivgi/kling-video/create-voiceKling Video Create Voice by KwaiVGI - advanced AI model for audio-to-audio. Delivers high-quality results with fast inference, suitable for both creative and production workflows.kwaivgi/kling-video/o1/pro/editKling Video O1 Pro is KwaiVGI's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.kwaivgi/kling-video/o1/pro/video-to-videoKling Video O1 Pro by KwaiVGI - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.