bytedance/omnihuman/1.5
Omnihuman 1.5 is Bytedance's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.
PNG, JPEG, WebP, or GIF · 20 MiB maximum
PNG, JPEG, WebP, or GIF · 20 MiB maximum
Example output — click Run to generate your own
Send your first request
OpenAI-compatible endpoint with unified authentication and usage tracking.
https://api.sandbase.ai/v1/runbytedance/omnihuman/1.5Input Schema
6 parameters · 2 required · 4 optional
| Parameter | Type | Required | Description |
|---|---|---|---|
image | string | Required | The URL of the image used to generate the video |
prompt | string | Required | The text prompt used to guide the video generation. |
mask | string | Optional | The URL of the mask image to apply to the image. Only the person in the white area of the mask will speak. |
audio | string | Optional | The URL of the audio file to generate the video. Audio must be under 30s long for 1080p generation and under 60s long for 720p generation. |
resolution | string | Optional | The resolution of the generated video. Defaults to 1080p. 720p generation is faster and higher in quality. 1080p generation is limited to 30s audio and 720p generation is limited to 60s audio. · Options: 720p, 1080p · Default: "1080p" 720p1080p |
turbo_mode | boolean | Optional | Generate a video at a faster rate with a slight quality trade-off. · Default: false |
Output Schema
| Field | Type | Description |
|---|---|---|
id | string | Unique identifier for the generation task |
status | string | Task status: pending, running, completed, failed, timeout |
model | string | Model used for the generation |
outputs | array | Array of output items |
outputs[].url | string | URL of the generated artifact |
outputs[].content_type | string | MIME type (e.g. image/png, video/mp4) |
error | object | null | Error details if failed, null on success |
error.type | string | Machine-readable error type code |
error.message | string | Human-readable error description |
Async Workflow
This model uses asynchronous execution. Submit a request and poll for the result.
- Submit — POST to /v1/run, receive an
id - Poll — GET /v1/run/{id} until status is
completed,failed, ortimeout - Retrieve — Read
outputsfrom the completed response
Code Examples
Ready-to-run snippets
# Step 1: Submit
curl -X POST https://api.sandbase.ai/v1/run \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "bytedance/omnihuman/1.5",
"audio": "https://static.sandbase.ai/examples/bytedance/omnihuman/1.5/input_audio_0.mp3",
"image": "https://static.sandbase.ai/examples/bytedance/omnihuman/1.5/input_image_1.png",
"resolution": "1080p",
"turbo_mode": false,
"prompt": "a beautiful sunset over mountains"
}'
# Step 2: Poll result (replace <id>)
curl https://api.sandbase.ai/v1/run/<id> \
-H "Authorization: Bearer YOUR_API_KEY"API README
Bytedance Omnihuman 1.5
Bytedance Omnihuman 1.5 is built for newer audio-driven human animation with more expressive and stable whole-body performance. It gives creative and production teams a focused way to move from an approved brief or source asset to a reviewable result without fragmenting the job across unrelated tools. The model is most valuable when visual intent, brand suitability, and downstream usability all matter, because its output can enter an editorial, campaign, product, or content pipeline as a purposeful asset rather than an isolated experiment.
In practice, teams can use Bytedance Omnihuman 1.5 during a structured cycle of briefing, generation, comparison, and refinement. Establish the subject, audience, visual objective, and acceptance criteria first; prepare any reference media at suitable quality; then evaluate alternatives for composition, continuity, realism, and communication value before delivery. This workflow keeps creative judgment central while making repeated production easier to review, reproduce, and scale for the specific bytedance/omnihuman/1.5 task.
Highlights
Enhanced motion realism. Enhanced motion realism gives bytedance › omnihuman › 1.5 a recognizable technical advantage: reviewers can assess this property directly in the generated asset instead of inferring it from request mechanics.
Improved identity consistency. bytedance › omnihuman › 1.5 applies improved identity consistency to the visual or temporal result itself, helping artists make a meaningful quality decision during selection and refinement.
Expressive gesture generation. For bytedance › omnihuman › 1.5, expressive gesture generation supports coherent assets across the intended creative workflow and distinguishes this capability from a simple format or delivery option.
Stable long-form performance. The practical value of stable long-form performance is visible in the finished media from bytedance › omnihuman › 1.5, where it supports repeatable art direction rather than merely exposing another request setting.
Pricing
| Billing basis | Rate |
|---|---|
| Audio-driven output runtime | $0.160000 per second |
When to Use
| ✅ Good fit | ❌ Consider alternatives |
|---|---|
| Use it when newer audio-driven human animation with more expressive and stable whole-body performance is the central production goal | Choose a different model when the required media task is fundamentally different |
| The team needs several reviewable creative alternatives | Exact deterministic reproduction is mandatory |
| Visual quality and practical downstream use both matter | Editable source layers or native project files are required |
| A managed generation step fits the delivery workflow | A live frame-by-frame interactive editor is essential |
| The documented inputs cover the available source assets | Required source media or controls fall outside the documented fields |
Prompt Guide
State the intended result first, then describe the subject, environment, action, visual treatment, and delivery constraints. Keep instructions concrete, avoid conflicting directions, and change one creative variable at a time when comparing outputs. For source-driven work, describe what should remain recognizable as clearly as what should change.
{
"audio": "https://static.sandbase.ai/examples/bytedance/omnihuman/1.5/input_audio_0.mp3",
"image": "https://static.sandbase.ai/examples/bytedance/omnihuman/1.5/input_image_1.png",
"prompt": "Describe the intended result in clear visual terms.",
"resolution": "1080p"
}
Technical Specs
| Property | Details |
|---|---|
mask | Type / options: string<br>Required: No<br>Description: The URL of the mask image to apply to the image. Only the person in the white area of the mask will speak. |
audio | Type / options: string<br>Required: No<br>Description: The URL of the audio file to generate the video. Audio must be under 30s long for 1080p generation and under 60s long for 720p generation. |
image | Type / options: string<br>Required: Yes<br>Description: The URL of the image used to generate the video |
prompt | Type / options: string<br>Required: Yes<br>Description: The text prompt used to guide the video generation. |
resolution | Type / options: string · 720p / 1080p<br>Required: No<br>Description: The resolution of the generated video. Defaults to 1080p. 720p generation is faster and higher in quality. 1080p generation is limited to 30s audio and 720p generation is limited to 60s audio. |
turbo_mode | Type / options: boolean<br>Required: No<br>Description: Generate a video at a faster rate with a slight quality trade-off. |
Related Models
bytedance/dreamactor/2.0bytedance/lynxbytedance/omnihuman/1.0

