Floniks

Lightricks · Image to video and lip sync

LTX 2.3 — 4K video and audio-driven lip sync

The highest resolution option, plus a talking-head mode driven by audio.

Try LTX 2.3

What is LTX 2.3?

LTX 2.3 covers two distinct jobs. Image-to-video renders 6 to 10 seconds at up to 4K — the highest resolution on Floniks alongside VEO 3.1 — with optional AI-generated ambient audio and first-and-last-frame support. The audio-to-video mode is different: give it a portrait and an audio clip and it produces a lip-synced talking video of 2 to 20 seconds. For maximum realism in talking-head work, OmniHuman v1.5 is the stronger option.

Available on Floniks

Credit cost per render, read live from the model catalogue. All models draw from one balance — no per-model subscription.

  • LTX 2.3 Image to Video48 credits
  • LTX 2.3700 credits

What LTX 2.3 is best for

  • 4K delivery where resolution is the constraint
  • Clips that need generated ambient audio
  • Short lip-synced talking videos from a portrait plus a voice track

How to use LTX 2.3

  1. Choose the mode

    Image-to-video for general animation; audio-to-video when you have a voice track and want lip sync.

  2. Supply your inputs

    Image-to-video takes one or two frames. Audio-to-video takes a portrait plus an audio clip.

  3. Set resolution deliberately

    4K costs materially more credits per render than 1080p. Reserve it for delivery.

Frequently asked questions

How long can an LTX 2.3 lip-sync clip be?
2 to 20 seconds in audio-to-video mode. For longer talking-head video, OmniHuman v1.5 supports up to 60 seconds at 720p.
Does LTX 2.3 really reach 4K?
Yes, in image-to-video mode — alongside VEO 3.1 it is the highest-resolution option on the platform.
LTX 2.3 or OmniHuman for talking heads?
OmniHuman v1.5 is more realistic and handles body movement more naturally. LTX 2.3 is the lighter, cheaper option for short clips.

Compare with other models

Try LTX 2.3
LTX 2.3 — 4K video and audio-driven lip sync | Floniks