SandBase is live — $1 in free credits on signupStart free ›

PixVerse modelsvideo generation api

pixverse/v6/transition

V6 Transition is PixVerse's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.

Input
The prompt for the transition
The aspect ratio of the generated image. Allowed values: 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16.

PNG, JPEG, WebP, or GIF · 20 MiB maximum

URL of the image to use as the last frame
The resolution of the generated video Allowed values: 360p, 540p, 720p, 1080p.
115
The duration of the generated video in seconds. v6 supports values from 1 to 15 seconds Range: 1 to 15.
The same seed and the same prompt given to the same version of the model will output the same video every time.

PNG, JPEG, WebP, or GIF · 20 MiB maximum

URL of the image to use as the first frame
The style of the generated video Allowed values: anime, 3d_animation, clay, comic, cyberpunk.
Enable audio generation (BGM, SFX, dialogue)
Idle

Example output — click Run to generate your own

API README

PixVerse V6 Transition

PixVerse V6 Transition is the PixVerse V6 route for first-to-last-frame transition video, turning opening and closing frames into a production-ready result while keeping the operation distinct from neighboring endpoints. V6 Transition is PixVerse's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output. The workflow is designed for creators who need the model’s specific transformation to remain visible in the request, so the source, intended change, and finished artifact can be reviewed as one coherent creative decision.

In practical use, this route exposes seed, style, prompt, duration, sync_mode, resolution, aspect_ratio, end_image_url to shape the exact deliverable. Those controls let a team preserve the important source constraints, state subject behavior or material treatment precisely, choose supported timing or output characteristics, and reproduce successful settings across alternate takes. The result fits an iterative pipeline: establish the core brief, compare controlled variations, then pass the selected asset into editorial, design, localization, visualization, or publishing work.

Highlights

  • Endpoint-controlled transition. Use opening and closing images to define both ends of the shot, then direct the motion that connects them.
  • Physics-aware motion. Render movement with stronger treatment of gravity, weight, elasticity, and material behavior.
  • Performance detail. Preserve subtle character expression and texture while the shot develops.
  • Cinematic continuity. Use improved shot language and optional synchronized audio for a more complete short-form sequence.

Pricing

ConfigurationPrice
1s · 360p · silent$0.025
1s · 360p · audio$0.035
1s · 540p · silent$0.035
1s · 540p · audio$0.045
1s · 720p · silent$0.045
1s · 720p · audio$0.060
1s · 1080p · silent$0.090
1s · 1080p · audio$0.115
2s · 360p · silent$0.050
2s · 360p · audio$0.070
2s · 540p · silent$0.070
2s · 540p · audio$0.090
2s · 720p · silent$0.090
2s · 720p · audio$0.120
2s · 1080p · silent$0.180
2s · 1080p · audio$0.230
3s · 360p · silent$0.075
3s · 360p · audio$0.105
3s · 540p · silent$0.105
3s · 540p · audio$0.135
3s · 720p · silent$0.135
3s · 720p · audio$0.180
3s · 1080p · silent$0.270
3s · 1080p · audio$0.345
4s · 360p · silent$0.100
4s · 360p · audio$0.140
4s · 540p · silent$0.140
4s · 540p · audio$0.180
4s · 720p · silent$0.180
4s · 720p · audio$0.240
4s · 1080p · silent$0.360
4s · 1080p · audio$0.460
5s · 360p · silent$0.125
5s · 360p · audio$0.175
5s · 540p · silent$0.175
5s · 540p · audio$0.225
5s · 720p · silent$0.225
5s · 720p · audio$0.300
5s · 1080p · silent$0.450
5s · 1080p · audio$0.575
6s · 360p · silent$0.150
6s · 360p · audio$0.210
6s · 540p · silent$0.210
6s · 540p · audio$0.270
6s · 720p · silent$0.270
6s · 720p · audio$0.360
6s · 1080p · silent$0.540
6s · 1080p · audio$0.690
7s · 360p · silent$0.175
7s · 360p · audio$0.245
7s · 540p · silent$0.245
7s · 540p · audio$0.315
7s · 720p · silent$0.315
7s · 720p · audio$0.420
7s · 1080p · silent$0.630
7s · 1080p · audio$0.805
8s · 360p · silent$0.200
8s · 360p · audio$0.280
8s · 540p · silent$0.280
8s · 540p · audio$0.360
8s · 720p · silent$0.360
8s · 720p · audio$0.480
8s · 1080p · silent$0.720
8s · 1080p · audio$0.920
9s · 360p · silent$0.225
9s · 360p · audio$0.315
9s · 540p · silent$0.315
9s · 540p · audio$0.405
9s · 720p · silent$0.405
9s · 720p · audio$0.540
9s · 1080p · silent$0.810
9s · 1080p · audio$1.035
10s · 360p · silent$0.250
10s · 360p · audio$0.350
10s · 540p · silent$0.350
10s · 540p · audio$0.450
10s · 720p · silent$0.450
10s · 720p · audio$0.600
10s · 1080p · silent$0.900
10s · 1080p · audio$1.150
11s · 360p · silent$0.275
11s · 360p · audio$0.385
11s · 540p · silent$0.385
11s · 540p · audio$0.495
11s · 720p · silent$0.495
11s · 720p · audio$0.660
11s · 1080p · silent$0.990
11s · 1080p · audio$1.265
12s · 360p · silent$0.300
12s · 360p · audio$0.420
12s · 540p · silent$0.420
12s · 540p · audio$0.540
12s · 720p · silent$0.540
12s · 720p · audio$0.720
12s · 1080p · silent$1.080
12s · 1080p · audio$1.380
13s · 360p · silent$0.325
13s · 360p · audio$0.455
13s · 540p · silent$0.455
13s · 540p · audio$0.585
13s · 720p · silent$0.585
13s · 720p · audio$0.780
13s · 1080p · silent$1.170
13s · 1080p · audio$1.495
14s · 360p · silent$0.350
14s · 360p · audio$0.490
14s · 540p · silent$0.490
14s · 540p · audio$0.630
14s · 720p · silent$0.630
14s · 720p · audio$0.840
14s · 1080p · silent$1.260
14s · 1080p · audio$1.610
15s · 360p · silent$0.375
15s · 360p · audio$0.525
15s · 540p · silent$0.525
15s · 540p · audio$0.675
15s · 720p · silent$0.675
15s · 720p · audio$0.900
15s · 1080p · silent$1.350
15s · 1080p · audio$1.725

When to Use

ScenarioWhy it fits
Exact workflow fitChoose this route when the required deliverable is first-to-last-frame transition video, rather than a related route with different source media.
Directed creative iterationUse it when subject, motion, material, speech, framing, or finish should be expressed explicitly and compared across controlled variants.
Existing-asset continuityUse it when supplied images, video, audio, references, or styles must remain the anchor for the generated result.
Repeatable productionUse it when successful inputs need to be saved and rerun across a campaign, asset set, localization pass, or batch.
Pipeline handoffUse it when the returned artifact will move into editorial, compositing, visualization, review, storage, or publishing.

Prompt Guide

Start with the source or subject, then describe the intended transformation, movement or behavior, camera and composition, and the desired finish. Keep media URLs reachable, use only fields documented for this exact route, and change one major control at a time when comparing results. For source-led tasks, describe what should change as well as what must remain recognizable.

{
  "prompt": "A cinematic close shot with clear subject action, camera movement, lighting, and final visual treatment.",
  "duration": 5,
  "resolution": "720p",
  "aspect_ratio": "16:9"
}

Technical Specs

PropertyValue
Model IDpixverse/v6/transition
Execution modeasync
Required inputsprompt
seedinteger
stylestring
promptstring
durationinteger; minimum: 1; maximum: 15; default: 5
sync_modeboolean; default: false
resolutionstring; options: 360p, 540p, 720p, 1080p; default: 720p
aspect_ratiostring; options: 16:9, 4:3, 1:1, 3:4, 9:16, 2:3, 3:2, 21:9; default: 16:9
end_image_urlstring
thinking_typestring
first_image_urlstring
negative_promptstring; default:
safety_tolerancestring; default: 6
generate_audio_switchboolean; default: false
generate_multi_clip_switchboolean; default: false

Related Models

  • pixverse/v6/extend
  • pixverse/v6/image-to-video
  • pixverse/v6/text-to-video
  • pixverse/c1/image-to-video

Related Models

pixverse/v6/image-to-videoV6 is PixVerse's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.pixverse/v6/text-to-videoV6 by PixVerse - generate cinematic videos from text descriptions with AI. Create high-quality video content with natural motion, camera control, and optional audio generation.pixverse/c1/text-to-videoC1 by PixVerse - generate cinematic videos from text descriptions with AI. Create high-quality video content with natural motion, camera control, and optional audio generation.pixverse/c1/transitionC1 Transition is PixVerse's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.pixverse/extendExtend by PixVerse - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.pixverse/extend/fastExtend Fast is PixVerse's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.pixverse/lipsyncLipsync is PixVerse's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.pixverse/sound-effectsSound Effects is PixVerse's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.