Alibaba · Image, text and reference to video
Wan 2.7 — reference-to-video for subject consistency
Up to 5 reference images to keep the same person across every shot.
Try Wan 2.7What is Wan 2.7?
Wan 2.7 is Alibaba’s video family and its reference-to-video mode is the reason to pick it: up to 5 reference images specifically for holding a subject consistent across shots. Supply several photos of the same person and the model keeps them recognisable, which is the hard part of building a series. Image-to-video takes one frame as the first, or two to pin both ends. All modes run 5–15 seconds at 720p or 1080p. Wan 2.6 Flash is the quick-iteration tier.
Available on Floniks
Credit cost per render, read live from the model catalogue. All models draw from one balance — no per-model subscription.
- Wan 2.7 (Text to Video)10 credits
- Wan 2.7 R2V (Reference to Video)10 credits
- Wan V2.650 credits
What Wan 2.7 is best for
- Keeping one character recognisable across multiple clips
- First-and-last-frame control with two input images
- Fast motion-prompt iteration on the 2.6 Flash tier
How to use Wan 2.7
Collect references of the subject
Up to 5 images in reference mode. Varied angles help more than repeated near-identical shots.
Reference them in the prompt
Wan responds to explicit references — write "image 1", "image 2" to point at specific inputs.
Iterate on Flash, finish on 2.7
Wan 2.6 Flash renders quickly for motion tests; move to 2.7 for the final.
Frequently asked questions
- How does Wan 2.7 keep a character consistent?
- Its reference-to-video mode accepts up to 5 reference images of the same subject and holds their identity across the generated clip. Reference them explicitly in the prompt as "image 1", "image 2" and so on.
- What is Wan 2.6 Flash for?
- Fast iteration. Use it to test motion prompts cheaply, then re-run the winner on Wan 2.7.
- Does Wan support first and last frame?
- Yes — supply two images in image-to-video mode to pin both ends of the motion.
