ByteDance · Multimodal to video
Seedance 2.0 — images, video and audio as inputs
The multimodal tier: combine up to 9 images, 3 videos and 3 audio clips in one render.
Try Seedance 2.0What is Seedance 2.0?
Seedance 2.0 is ByteDance’s video model, and its multimodal variant is unusual in accepting images, video and audio simultaneously as inputs — up to 9 images, 3 videos and 3 audio clips. That makes it the right tool when the shot is assembled from existing material rather than generated from one still. It runs 4 to 15 seconds at up to 1080p with good motion dynamics, and supports either reference images or first/last keyframe modes.
Available on Floniks
Credit cost per render, read live from the model catalogue. All models draw from one balance — no per-model subscription.
- Seedance 2.0 (Multimodal)10 credits
- Seedance 1.5 Pro6 credits
What Seedance 2.0 is best for
- Shots assembled from several existing media types
- Reference-driven video where multiple images define the subject
- Smooth, cinematically styled motion at flexible durations
- First-and-last keyframe work
How to use Seedance 2.0
Assemble your inputs
Up to 9 images, 3 videos and 3 audio clips. Keyframe mode and reference mode are mutually exclusive with video and audio inputs.
Choose a duration
4 to 15 seconds. Longer clips cost more credits and take longer to render.
Prompt the motion
Describe camera and subject movement; the references already carry look and composition.
Frequently asked questions
- What inputs does Seedance 2.0 accept?
- The multimodal variant takes images, videos and audio at once — up to 9 images, 3 videos and 3 audio clips. Keyframe mode is mutually exclusive with video and audio inputs.
- How long can Seedance 2.0 clips run?
- 4 to 15 seconds at up to 1080p.
- Seedance 2.0 or Seedance 1.5 Pro?
- Seedance 2.0 for multimodal inputs and duration range. Seedance 1.5 Pro is the premium single-path tier when 2.0 output quality is not sufficient.
