Wan 2.7
Wan 2.7 is Alibaba’s image-to-video model for clips of 2 to 15 seconds at 720p or 1080p, generated from a still and a prompt, with the option of an audio clip whose rhythm drives the motion. Use it when you already have the frame and the shot needs to move to a sound, or when you want a negative prompt to keep something out.
It runs in the editor on the left at 21 credits a second for 720p and 32 for 1080p, so a 5-second 720p clip is 105 credits and a 15-second 1080p one is 480. It does not start from text: for that, use Wan 3.0.
What Wan 2.7 adds to image to video
An audio input. Upload a WAV or MP3 of 2 to 30 seconds and the model times the motion to it, which is the feature people ask for when a clip has to land on a beat or a spoken line. No other model in the studio takes your sound as an input; the ones that generate audio invent it from the scene.
A negative prompt, up to 1,000 characters, for what must not appear. And an audio-out switch: leave it on for the generated track, turn it off for a silent file. The seed is settable, so a result you like can be re-run with one change.
What it costs
Per second of output. Audio input is free; only the generated seconds are billed.
| Size | Per second | 5 seconds | 15 seconds |
|---|---|---|---|
| 720p | 21 credits | 105 | 315 |
| 1080p | 32 credits | 160 | 480 |
Our test
The hero clip is our home-page still animated with one sentence, at 720p for 5 seconds, no audio input. It came back in about 100 seconds as a 1278×720 MP4 of 5.0 seconds with a generated audio track. At today’s rate that clip is 105 credits; it ran before a rate change and was billed 98.
The same still and sentence went through Wan 2.2, Wan 3.0 and Seedance 2.5. Of the four, this is the only one with a 720p result, which is why its file is the sharpest and its price the highest.
| Field | Value |
|---|---|
| Input | One still, 1280×720, plus the prompt in the caption |
| Settings | 720p, 5 seconds, audio out on, no audio input |
| Cost | 105 credits at the current rate ($1.05) |
| Wall time | About 100 seconds |
| Output | 1278×720, 5.0 s, MP4 with audio |
Wan 2.7 or Wan 2.2?
2.2 costs 60 credits for the same 5 seconds at 480p and comes back in about a minute, which makes it the right first pass on any still: if the motion is wrong at 60 credits it will be wrong at 105. 2.7 is the second pass, for 720p or 1080p, for a longer clip than 8 seconds, or for motion timed to your audio.
How to run Wan 2.7 here
- 1. Upload the still PNG, JPG or WebP; a sharp image between 360 and 2000 pixels a side.
- 2. Describe the motion, not the picture The model can see the image. Say what moves and how the camera moves.
- 3. Add audio if the motion has to follow it Optional. 2 to 30 seconds of WAV or MP3.
- 4. Set length and size, then create 5 seconds at 720p is 105 credits and about two minutes.
When not to use this
- Starting from a prompt alone. Wan 2.7 needs a still; the editor marks the image slot required as soon as you pick it.
- 480p drafts. The smallest size here is 720p at 21 credits a second; do the cheap check on Wan 2.2.
- Takes over 15 seconds. Wan 3.0 and Seedance 2.5 go to 30.
- An exact match to your audio. The motion follows the rhythm and energy of the track; it does not lip-sync words. That is the lip sync job, not this one.
- A photo of somebody who has not agreed to it.
Questions people ask
How much does Wan 2.7 cost?
21 credits a second at 720p and 32 at 1080p, where 1 credit is one US cent. A 5-second 720p clip is 105 credits, a 15-second 1080p clip is 480. Audio input is free, and a failed job is refunded.
Can Wan 2.7 generate video from text?
No. It is an image-to-video model and needs a first frame. Wan 3.0 is the Wan model that starts from a prompt.
What does the audio input do?
It drives the timing and energy of the motion, so the clip moves with the track. It does not make the subject speak the words; for that use lip sync.
Can I get a silent clip?
Yes. Turn audio off before generating and the file comes back without a track.
How long can the clip be?
2 to 15 seconds. Five is the safest default for a face; drift grows with length on every model.
Related tools
Checked 2026-09-23.