SandBase is live — $1 in free credits on signupStart free ›

HeyGen modelsvideo generation api

heygen/avatar5/digital-twin

Avatar5 Digital Twin by HeyGen - generate cinematic videos from text descriptions with AI. Create high-quality video content with natural motion, camera control, and optional audio generation.

Input
Text the avatar will speak. Required when ``audio_url`` is not provided.
The aspect ratio of the generated image. Allowed values: 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16.
HTTP(S) URL of an audio file for the avatar to lip-sync to. When provided, the avatar uses this audio instead of text-to-speech and ``prompt``/``voice`` are ignored.
Output resolution preset. Allowed values: 720p, 1080p, 4k.
Output container format. 'webm' produces a transparent video (automatically removes the background) and ignores the ``background`` field. Allowed values: mp4, webm.
Remove the avatar's background. Requires a matting-enabled avatar.
Name of the Avatar V-eligible avatar to use.
Name of the text-to-speech voice to use for the avatar when ``audio_url`` is not provided. Allowed values: Warm Pro Narrator, Chill Brian, Ivy, John Doe, Monika Sogam, Hope , Archer , Brittney, Patrick, David Castlemore, Michael C, Adam Stone , Juniper, Cassidy , Jessica Anne Bogart, Arabella, Andrew, Spuds Oxley , Grace Elder, Helen, Canyon Rivers, Derya - Lifelike - Excited 🤩, Mellow Marcus, Jack Sterling - Broadcaster 🎙️, Brenda - UGC - 1.mp4, Reid, Reagan, Terry, Jenny, Radio Rick, Denise, Tim in car - Excited 🤩, Iskander, Thompson, Delicate Daisy - Excited 🤩, Kingston, George UGC 1, Bold Blake, Jane, Expressive Evan, Marianne - IA, Aaron, Modern Recipe Host - Voice 1, Willow, Cute Chloe - Friendly 😊, Rafael, June - Lifelike, Crisp Chloe, Slick Simon, Nassim - Informative, Baritone Ben, Maxwell, Ellie Faye - Excited 🤩, Milani, Feisty Fiona - Excited 🤩, Professor Dean, Rose - UGC - 1.mp4, Shona, Hudson Wilder, Ann - IA, Alastair Kensington, Oxley, Christina, Andrew Rizz , Peyton, Gerardo - Outdoor, Chloe - Lifelike, Stephanie, Anthony - IA, Signal - Voice 1, Luca, Lisa - Voice 1, T.W.Tucker, Jack Sullivan - Serious 😐, Winter, Mireia - Lifelike, Georgia, Stella, Masha - Lifelike, Charming Charles - Friendly 😊, Serenity, Annie - Excited, Ralph, Bethany, Dominic, Mason Finn, Leena, Veteran Victor, Tamara, Nik Public, Calm Chloe, Sevik, Reilly, Raul, Imposing Ian, Relaxed Ray, Dexter - Professional, Relaxed Rick, Edwin, Rupert Blackwood, Ginny, Hope.
Generate a sidecar SRT caption file alongside the video.
Optional watermark image to overlay on the output video.
Optional background to composite behind the avatar. Ignored when ``output_format='webm'`` (webm output is transparent).
How the avatar fits within the output frame. 'contain' keeps the full avatar in view (may letterbox); 'cover' fills the frame (may crop). Allowed values: contain, cover.
Idle

Example output — click Run to generate your own

API README

Heygen v5 Digital Twin

heygen/avatar5/digital-twin creates a talking digital-human video whose facial movement, timing, and expression follow supplied speech or creative direction. Avatar 5 builds a highly expressive digital twin with improved facial realism, emotion, gesture, and speech synchronization for presenter performances. This combination makes the model a practical choice when the creative outcome depends on those qualities rather than on a generic media conversion.

For production work with heygen/avatar5/digital-twin, the workflow suits explainers, training, localization, personalized outreach, and presenter-led content where a reusable human presence needs to deliver changing material. The result is most reliable when the source material and creative brief clearly describe the intended subject, progression, visual or sonic character, and the qualities that must remain unchanged.

Highlights

Speech-driven facial animation aligns mouth shapes, expression, and timing with the performance.

Avatar 5 realism captures subtle facial expression, emotion, and lifelike presenter behavior.

Presenter consistency supports repeated videos without reshooting the same person.

Expressive performance turns scripted material into an engaging on-camera delivery.

Pricing

Generated or processed durationPrice
Per second$0.100

When to Use

✅ Good fit❌ Consider alternatives
The model's named workflow matches the source material and intended outputA different input modality or model route is required
A managed asynchronous result is suitable for the production pipelineA synchronous, interactive editor is essential
The documented controls cover the required duration, framing, or formatThe project needs controls outside this endpoint's schema
Creative iteration benefits from a repeatable request structureExact deterministic pixels, frames, geometry, or samples are mandatory
A finished downloadable media asset is the desired deliverableEditable source layers or a native project file are required

Prompt Guide

For audio transformation, state the intended result first, then add the subject or source treatment, progression, style, and delivery constraints. Keep one creative variable per phrase, use the documented field names for controls, and change one setting at a time when comparing results.

{
  "aspect_ratio": "21:9",
  "audio": "https://example.com/source.wav",
  "output_format": "mp4",
  "prompt": "Hello and welcome to this demonstration of HeyGen Avatar V.",
  "resolution": "720p"
}

Technical Specs

SpecValue
Model IDheygen/avatar5/digital-twin
Inputsaspect_ratio, audio, avatar, background, caption, fit, output_format, prompt, remove_background, resolution, voice, watermark
Required inputsprompt
Output fieldscontent_type, url
ExecutionAsync (submit, then poll for result)
Resolution720p / 1080p / 4k
Aspect Ratio21:9 / 16:9 / 3:2 / 4:3 / 5:4 / 1:1 / 4:5 / 3:4 / 2:3 / 9:16
Output Formatmp4 / webm

Related Models

Related Models

heygen/heygen-avatar3/digital-twinHeygen Avatar3 Digital Twin is HeyGen's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.heygen/heygen-avatar4/digital-twinHeygen Avatar4 Digital Twin is HeyGen's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.heygen/heygen-avatar4/image-to-videoHeygen Avatar4 is HeyGen's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.heygen/heygen/lipsync/precisionHeygen Lipsync Precision by HeyGen - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.heygen/heygen/lipsync/speedHeygen Lipsync Speed by HeyGen - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.heygen/heygen/translate/precisionHeygen Translate Precision by HeyGen - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.heygen/heygen/translate/speedHeygen Translate Speed by HeyGen - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.heygen/heygen/video-agent/v2Heygen Video Agent V2 is HeyGen's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.