ByteDance · Audio to video
OmniHuman v1.5 — realistic full-body talking video
A portrait plus an audio track becomes a person who moves like one.
Try OmniHuman v1.5What is OmniHuman v1.5?
OmniHuman v1.5 turns a portrait and an audio track into a realistic talking video, and the thing that sets it apart from lighter lip-sync models is body movement — gestures and posture that read as a person speaking rather than a still with a moving mouth. It supports up to 60 seconds of audio at 720p, or 30 seconds at 1080p. It is one of the most credit-expensive models on the platform, so scope the script before rendering.
Available on Floniks
Credit cost per render, read live from the model catalogue. All models draw from one balance — no per-model subscription.
- OmniHuman v1.5500 credits
What OmniHuman v1.5 is best for
- AI spokesperson and product demo videos
- Course and training narration with a presenter on screen
- Longer talking-head content where lip-sync-only output looks stiff
How to use OmniHuman v1.5
Prepare the portrait
A clear, front-facing image works best. The framing you supply sets the framing of the output.
Supply the audio
Up to 60 seconds at 720p or 30 at 1080p. Generate the voice track first with a text-to-speech model if you do not have one.
Budget the credits
This is among the most expensive models on the platform. Finalise the script before rendering rather than iterating on full takes.
Frequently asked questions
- How long can an OmniHuman video be?
- Up to 60 seconds of audio at 720p, or 30 seconds at 1080p.
- How is OmniHuman different from LTX lip sync?
- OmniHuman generates natural body movement and gesture, not just mouth motion, which is what makes longer clips watchable. It costs considerably more per render.
- Can I generate the voice on Floniks too?
- Yes. Generate the track with a text-to-speech or voice-cloning model, then feed it into OmniHuman inside the same workflow.
