Wan 3.0 is Alibaba's video model for 30-second single takes with audio generated in the same pass. This page runs it in the generator above and states what its API document actually specifies — resolutions, duration range, reference inputs and price per second — rather than the spec sheets that circulated before it shipped.
Four steps on the generator above, and the second is where most clips are decided.
Starting from a still fixes what the scene looks like and leaves the model to handle only the motion, which is the more predictable path.
Say what the camera does and what moves inside the frame as two statements. Prompts that merge them tend to produce a still image with drift.
Anywhere from 2 to 15 seconds, at 720p or 1080p. Length costs linearly, so a first pass is cheaper short.
Change a single element between runs — the move, or the subject, or the lighting. Rewriting the whole prompt makes it impossible to tell what helped.
Wan 3.0 is the current generation of Alibaba's Wan video line, launched in August 2026 after a gated beta. It generates up to 30 seconds in a single continuous take at 480p, 720p or 1080p, produces an audio track together with the picture rather than after it, and accepts reference images, reference video, reference audio, documents and links as input alongside the prompt. It is an API model: the weights are not published, which is the one thing about it that has not changed.
Checked on 2 September 2026: the Wan-AI organisation on HuggingFace and the Wan-Video organisation on GitHub both still top out at the 2.2 family plus the Animate and Dancer branches. It is available through provider APIs; there is no 3.0 checkpoint to download, and any page offering one is not describing this model.
Until it shipped, this page said plainly that it had not, and ran the current release instead. That was accurate on the day it was written and is kept on the record here rather than quietly deleted.
Creative engine
Billed per second of output at the rate the model charges, by resolution. Nothing else moves it — audio is included rather than surcharged.
Resolution is 480P, 720P or 1080P. Duration is an integer from 2 to 30 seconds, or -1 to let the model choose its own length. Audio defaults to on. Aspect ratio covers 16:9, 9:16, 4:3, 3:4, 1:1 and an adaptive mode that picks from the input. Reference images, video, audio, files and links each have their own parameter, as does pinning a first or last frame.
Most video models hand you four to ten seconds and leave the rest to editing. This one's argument is that the cut is the problem, and it removes it.
A single generation covers the whole clip, so light, wardrobe and the way a character moves do not have to be matched across pieces that were never told about each other.
Sound is generated in the same pass rather than dubbed over the finished video, and it is priced into the per-second rate rather than charged as an extra.
Reference images, a reference clip, reference audio, a first or last frame — each is its own input, so look, movement and rhythm do not have to compete for room in one sentence.
Capabilities taken from the model's own API document. Each one is a parameter the model accepts, not a claim about how well it uses it.
Generate a clip from a written description alone, up to thirty seconds in one take.
Start from a still and describe the motion. Reference images are a first-class input rather than a mode.
Sound is generated together with the picture rather than added over the finished clip, and is on by default.
A reference clip supplies motion and camera language; reference audio drives pacing. Both ride alongside the prompt.
From the model's published API schema and price sheet, read on 2 September 2026.
Where a clip with its own soundtrack and a real camera move earns its keep.
Short vertical clips where the audio is part of the post rather than a track laid underneath it.
Landscape, sky and city moves used to open a sequence or bridge between two pieces of footage.
A slow move around an object or an idea, where the camera work carries the shot.
Roughing out how a sequence reads before anything expensive gets committed to it.
One credit pool covers Nano Banana images and Seedance and Veo video. The cost shows before every run, the safety check runs before the model does, and a failed run is refunded automatically. Use a subscription for ongoing work, or a one-time pack when you just need to top up.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
One-time top-ups — buy extra credits any time you run low.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
Charges appear as “SAYMAKER AI” on your card statement. You can cancel any time from Settings → Billing; cancellation takes effect at the end of the period you already paid for. Operator details, the full model list, and the refund window are on the about page.
Powered by
Use one balance across every image and video model on the shelf — the main ones are listed below — and check the credit cost before each request.
Video models
Image models
Model credit guide
A failed run is refunded automatically. The exact estimate in the generator varies by model, length, resolution, audio, and number of images.
| Type | Model | Credit cost |
|---|---|---|
| Video | Seedance 2.5 | The long-take tier, priced per second: 5s 480p ≈ 525 credits, 5s 720p ≈ 1,185. A full 30s take runs ≈ 3,150 at 480p and ≈ 7,090 at 720p. |
| Video | Seedance 2.0 | Supplying a starting frame or clip costs less than starting from words alone: 6s 720p ≈ 565 credits from an image, ≈ 925 from text. Scales with resolution and length. |
| Video | Seedance 2 Fast | Faster and lower cost. 6s 720p text-to-video ≈ 745 credits. |
| Video | Seedance 2 Mini | The cheapest tier of the Seedance 2 family. 6s 720p ≈ 465 credits, 6s 480p ≈ 215 credits. |
| Video | Seedance 1.5 Pro | Audio doubles the rate. 6s 720p ≈ 85 credits silent, ≈ 160 with audio; 1080p ≈ 175 and ≈ 340. |
| Video | Veo 3.1 | Billed per video, not per second, and the tier you pick is the whole price: Lite ≈ 115 credits at 720p and ≈ 135 at 1080p, Fast ≈ 225 at 720p, Quality ≈ 940 at 720p. |
| Video | Kling 3.0 | Audio raises the rate by about half. 6s 720p ≈ 320 credits silent, ≈ 455 with audio; 1080p ≈ 405 and ≈ 610. |
| Video | MiniMax H3 | Fixed 2K, no resolution ladder. Priced per second — a 5s clip ≈ 395 credits. |
| Image | Nano Banana 2 | Generate or edit from text and images. 20 credits per 1K image, 30 at 2K, 45 at 4K. |
| Image | Nano Banana Pro | Consistent run times across generations. 30 credits per 1K or 2K image, 55 at 4K. |
| Image | Nano Banana 2 Lite | Faster, simpler variant, and the everyday editing price: a flat 15 credits per image at every resolution. |
| Image | GPT Image 2 | The lowest-cost premium image model here. 10 credits per 1K image, 15 at 2K, 30 at 4K. |
| Image | GPT Image 2.5 Flare | OpenAI's newest image model, on the fast tier. 25 credits per 1K image, 40 at 2K, 60 at 4K — same price whether you generate or edit. |
| Image | GPT Image 2.5 Sunburst | The 2.5 tier for edits that touch only what you named. Same price as Flare: 25 credits per 1K image, 40 at 2K, 60 at 4K. |
| Image | Seedream 5.0 Lite | Flat pricing — 20 credits per image whether you generate or edit. |
What people ask about running Wan 3.0.
The model itself. The model id was probed against the live endpoint before this page was changed, and the parameters the generator sends — resolution, duration, aspect ratio, reference images — are the ones its published schema documents.
Up to thirty seconds in a single generation. The schema takes an integer from 2 to 30, and passing -1 asks the model to choose a length itself. Thirty seconds arrives as one continuous take rather than as clips stitched together, which is the difference that matters when a character has to stay the same character throughout.
480P, 720P or 1080P, in 16:9, 9:16, 4:3, 3:4, 1:1 or an adaptive ratio chosen from the input. Vertical 9:16 comes straight out of the model, so a social cut needs no reframing afterwards, and 1080p holds detail across a full thirty-second take rather than only across a short one.
No. Checked on 2 September 2026, the Wan-AI organisation on HuggingFace and the Wan-Video organisation on GitHub both still stop at the Wan 2.2 family plus the Animate and Dancer branches. It runs through provider APIs, and a download offered as 3.0 weights is not this model.
A five second clip is 150 credits at 480p, 300 at 720p and 600 at 1080p; a full thirty second take is 900, 1,800 or 3,600. Billing is per second of output, so length and resolution are the only things that change the price — audio is included in the rate rather than surcharged.
Yes, and it is on by default. Audio is produced together with the picture rather than added afterwards, which is why the sound tends to match what is happening on screen instead of running alongside it.
Wan ships synchronised audio inside the rate and runs to thirty seconds in one take, which is where it separates from everything else here. Kling is the stronger motion engine — for a dance or an action shot go there; for a long ambient scene with sound included, Wan is the one built for the length.
Describe the camera move and the subject move as two separate statements, keep the first attempt short, and see what comes back.