SandBase is live — $1 in free credits on signupStart free ›

KwaiVGI modelsimage generation api

kwaivgi/kolors

Kolors by KwaiVGI - generate stunning images from text prompts with state-of-the-art AI. Supports multiple aspect ratios, styles, and high-resolution output for creative and commercial use.

Input
The prompt to use for generating the image. Be as descriptive as possible for best results.
The aspect ratio of the generated image. Allowed values: 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16.
110
The CFG (Classifier Free Guidance) scale is a measure of how close you want the model to stick to your prompt when looking for a related image to show you. Range: 1 to 10.
1150
The number of inference steps to perform. Range: 1 to 150.
The format of the generated image. Allowed values: jpeg, png.
Seed
The scheduler to use for the model. Allowed values: EulerDiscreteScheduler, EulerAncestralDiscreteScheduler, DPMSolverMultistepScheduler, DPMSolverMultistepScheduler_SDE_karras, UniPCMultistepScheduler, DEISMultistepScheduler.
Idle

Example output — click Run to generate your own

API README

Kolors

Kolors is a text-to-image generation endpoint in the Kolors family. It is built for creators who need to turn a concrete creative brief into a controlled visual sequence: the request establishes the source material, the intended subject behavior, the camera language, and the atmosphere of the finished shot. The standard route keeps that workflow explicit instead of hiding its input assumptions behind a generic video-generation label.

In practice, this route accepts a written prompt as creative context and exposes aspect ratio for delivery planning. That makes it suitable for shot-based pipelines where teams must preserve a source, direct a transformation, or control the final format without losing sight of the model's central task. Write prompts as a compact shot plan—subject, action, setting, camera, light, and timing—then use the structured fields for constraints that should remain deterministic across iterations.

Highlights

  • Bilingual visual semantics. Kolors was developed for strong Chinese and English prompt understanding across photorealistic and stylized imagery.
  • Composition control. A broad aspect-ratio set supports landscape, portrait, square, and social formats.
  • Generation-path control. Scheduler selection and reproducible seeds support intentional iteration.
  • Prompt-led synthesis. Detailed text directions define subject, scene, lighting, style, and framing.

Pricing

Billing unitPrice
Base request$0.000575

When to Use

ScenarioRecommendation
Choose this routeUse it when the deliverable specifically calls for text-to-image generation, rather than a neighboring generation mode.
Prepare the sourceProvide a precise written brief in the format described by the request schema.
Direct the shotDescribe the subject, action, environment, camera movement, lighting, and temporal progression in that order.
Control continuityUse endpoint frames, reference media, strength, or audio controls when those fields are available instead of burying hard constraints in prose.
Plan deliverySet duration, frame count, resolution, and aspect ratio explicitly when the schema exposes them, then compare iterations with a stable seed where supported.

Prompt Guide

For text-to-image generation, describe one coherent shot rather than a list of visual keywords. Put the main subject and action first, follow with location and staging, then add camera movement, lens or framing, lighting, mood, and any timed change. Keep URLs and hard delivery choices in their dedicated fields.

{
  "prompt": "A young Chinese couple with fair skin, dressed in stylish sportswear, with the modern Beijing city skyline in the background. Facial details, clear pores, captured using the latest camera model, close-up shot, ultra-high quality, 8K, visual feast.",
  "aspect_ratio": "21:9"
}

Technical Specs

SpecificationValue
Model IDkwaivgi/kolors
Required inputsprompt
ExecutionAsynchronous generation job
Request controls7 documented fields
Outputurl, content_type

Request fields

FieldType and constraints
seedinteger; Optional
promptstring; Required
schedulerstring; Optional; Options: EulerDiscreteScheduler, EulerAncestralDiscreteScheduler, DPMSolverMultistepScheduler, DPMSolverMultistepScheduler_SDE_karras, UniPCMultistepScheduler, DEISMultistepScheduler; Default: "EulerDiscreteScheduler"
aspect_ratiostring; Optional; Options: 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16
output_formatstring; Optional; Options: jpeg, png; Default: "png"
guidance_scalenumber; Optional; Range: 1–10; Default: 5
num_inference_stepsinteger; Optional; Range: 1–150; Default: 50

Related Models

Related Models

kwaivgi/kolors/image-to-imageKolors is KwaiVGI's intelligent image editing model. Transform, retouch, and reimagine existing images using text prompts - from background replacement to artistic style conversion.kwaivgi/kling-image/o3Kling Image O3 by KwaiVGI - generate stunning images from text prompts with state-of-the-art AI. Supports multiple aspect ratios, styles, and high-resolution output for creative and commercial use.kwaivgi/kling-image/o3/editKling Image O3 Edit is KwaiVGI's intelligent image editing model. Transform, retouch, and reimagine existing images using text prompts - from background replacement to artistic style conversion.kwaivgi/kling-image/v3Kling Image V3 by KwaiVGI - generate stunning images from text prompts with state-of-the-art AI. Supports multiple aspect ratios, styles, and high-resolution output for creative and commercial use.kwaivgi/kling-image/v3/editKling Image V3 Edit is KwaiVGI's intelligent image editing model. Transform, retouch, and reimagine existing images using text prompts - from background replacement to artistic style conversion.kwaivgi/kling-video/create-voiceKling Video Create Voice by KwaiVGI - advanced AI model for audio-to-audio. Delivers high-quality results with fast inference, suitable for both creative and production workflows.kwaivgi/kling-video/o1/pro/editKling Video O1 Pro is KwaiVGI's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.kwaivgi/kling-video/o1/pro/video-to-videoKling Video O1 Pro by KwaiVGI - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.