Bytedance modelsvideo generation api

bytedance/omnihuman/1.0

Omnihuman 1.0 is Bytedance's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.

Input

PNG, JPEG, WebP, or GIF · 20 MiB maximum

The URL of the image used to generate the video
The URL of the audio file to generate the video. Audio must be under 30s long.
Idle

Example output — click Run to generate your own

API README

Bytedance Omnihuman 1.0

Bytedance Omnihuman 1.0 is built for audio-driven human animation from a reference portrait or character image. It gives creative and production teams a focused way to move from an approved brief or source asset to a reviewable result without fragmenting the job across unrelated tools. The model is most valuable when visual intent, brand suitability, and downstream usability all matter, because its output can enter an editorial, campaign, product, or content pipeline as a purposeful asset rather than an isolated experiment.

In practice, teams can use Bytedance Omnihuman 1.0 during a structured cycle of briefing, generation, comparison, and refinement. Establish the subject, audience, visual objective, and acceptance criteria first; prepare any reference media at suitable quality; then evaluate alternatives for composition, continuity, realism, and communication value before delivery. This workflow keeps creative judgment central while making repeated production easier to review, reproduce, and scale for the specific bytedance/omnihuman/1.0 task.

Highlights

Audio-synchronized performance. Audio-synchronized performance gives bytedance › omnihuman › 1.0 a recognizable technical advantage: reviewers can assess this property directly in the generated asset instead of inferring it from request mechanics.

Identity-preserving animation. bytedance › omnihuman › 1.0 applies identity-preserving animation to the visual or temporal result itself, helping artists make a meaningful quality decision during selection and refinement.

Natural facial expression. For bytedance › omnihuman › 1.0, natural facial expression supports coherent assets across the intended creative workflow and distinguishes this capability from a simple format or delivery option.

Coordinated body motion. The practical value of coordinated body motion is visible in the finished media from bytedance › omnihuman › 1.0, where it supports repeatable art direction rather than merely exposing another request setting.

Pricing

Billing basisRate
Audio-driven output runtime$0.140000 per second

When to Use

✅ Good fit❌ Consider alternatives
Use it when audio-driven human animation from a reference portrait or character image is the central production goalChoose a different model when the required media task is fundamentally different
The team needs several reviewable creative alternativesExact deterministic reproduction is mandatory
Visual quality and practical downstream use both matterEditable source layers or native project files are required
A managed generation step fits the delivery workflowA live frame-by-frame interactive editor is essential
The documented inputs cover the available source assetsRequired source media or controls fall outside the documented fields

Prompt Guide

State the intended result first, then describe the subject, environment, action, visual treatment, and delivery constraints. Keep instructions concrete, avoid conflicting directions, and change one creative variable at a time when comparing outputs. For source-driven work, describe what should remain recognizable as clearly as what should change.

{
  "audio": "https://static.sandbase.ai/examples/bytedance/omnihuman/1.0/input_audio_1.mp3",
  "image": "https://static.sandbase.ai/examples/bytedance/omnihuman/1.0/input_image_0.png"
}

Technical Specs

PropertyDetails
audioType / options: string<br>Required: No<br>Description: The URL of the audio file to generate the video. Audio must be under 30s long.
imageType / options: string<br>Required: Yes<br>Description: The URL of the image used to generate the video

Related Models

  • bytedance/dreamactor/2.0
  • bytedance/lynx
  • bytedance/omnihuman/1.5

Related Models

bytedance/omnihuman/1.5Omnihuman 1.5 is Bytedance's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.bytedance/lynxLynx by Bytedance - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.bytedance/dreamactor/2.0DreamActor M2.0 by ByteDance generates videos by animating a reference image using motion from a driving video. It replicates motion, facial expressions, and lip movements from the template video while preserving the subject and background features of the input image.bytedance/seedance/1.0/pro/fast/image-to-videoSeedance 1.0 Pro by Bytedance - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.bytedance/seedance/1.0/pro/fast/text-to-videoSeedance 1.0 Pro is Bytedance's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.bytedance/seedance/1.0/pro/image-to-videoSeedance 1.0 Pro is Bytedance's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.bytedance/seedance/1.0/pro/text-to-videoSeedance 1.0 Pro by Bytedance - generate cinematic videos from text descriptions with AI. Create high-quality video content with natural motion, camera control, and optional audio generation.bytedance/seedance/1.5/pro/image-to-videoByteDance Seedance v1.5 Pro image-to-video model transforming still images into cinematic video with native audio generation, camera control, and professional-grade output quality.