Lightricks · Image to video and lip sync
LTX 2.3 — 4K video and audio-driven lip sync
The highest resolution option, plus a talking-head mode driven by audio.
Try LTX 2.3What is LTX 2.3?
LTX 2.3 covers two distinct jobs. Image-to-video renders 6 to 10 seconds at up to 4K — the highest resolution on Floniks alongside VEO 3.1 — with optional AI-generated ambient audio and first-and-last-frame support. The audio-to-video mode is different: give it a portrait and an audio clip and it produces a lip-synced talking video of 2 to 20 seconds. For maximum realism in talking-head work, OmniHuman v1.5 is the stronger option.
Available on Floniks
Credit cost per render, read live from the model catalogue. All models draw from one balance — no per-model subscription.
- LTX 2.3 Image to Video48 credits
- LTX 2.3700 credits
What LTX 2.3 is best for
- 4K delivery where resolution is the constraint
- Clips that need generated ambient audio
- Short lip-synced talking videos from a portrait plus a voice track
How to use LTX 2.3
Choose the mode
Image-to-video for general animation; audio-to-video when you have a voice track and want lip sync.
Supply your inputs
Image-to-video takes one or two frames. Audio-to-video takes a portrait plus an audio clip.
Set resolution deliberately
4K costs materially more credits per render than 1080p. Reserve it for delivery.
Frequently asked questions
- How long can an LTX 2.3 lip-sync clip be?
- 2 to 20 seconds in audio-to-video mode. For longer talking-head video, OmniHuman v1.5 supports up to 60 seconds at 720p.
- Does LTX 2.3 really reach 4K?
- Yes, in image-to-video mode — alongside VEO 3.1 it is the highest-resolution option on the platform.
- LTX 2.3 or OmniHuman for talking heads?
- OmniHuman v1.5 is more realistic and handles body movement more naturally. LTX 2.3 is the lighter, cheaper option for short clips.
