This is the tier built for scenes rather than shots: one call runs anywhere from 4 to 30 seconds, and the look, the motion and the pacing arrive through separate reference slots instead of all being crammed into the prompt.
Long generations reward different habits than short ones.
A five-second prompt describes a moment; a thirty-second prompt has to describe an order — what happens first, what the camera does when it changes, where it ends.
Look in the images, movement in a reference clip, pacing in reference audio. Prompts go vague when they carry all three.
Same prompt, same references, short duration. It is priced per second, so a test costs a sixth of the take.
Once the look is locked, re-run long. The full take is the deliverable; the short one was the rehearsal.
Seedance 2.5 is ByteDance's newest video model and the first tier here that generates a full scene in one pass. The Seedance 2 tiers are shot machines: describe a moment, get four to ten seconds of it, and anything longer becomes a stitching problem — matching light, matching wardrobe, matching the way a character walks, across clips that were never told about each other.
Most of the effort in AI video today goes into hiding that seam. Thirty seconds in a single generation removes it instead. The second change is what you are allowed to hand the model. One request accepts up to ten reference images, reference video clips and reference audio at the same time, and you can pin a first and a last frame and let it find the path between them.
Look, movement and rhythm each get their own slot rather than competing for room in one sentence of prompt. It runs at 480p, 720p or 1080p in seven aspect ratios, billed per second, with generated audio as a switch rather than a separate job.
Creative engine
Rate is per second and depends on the resolution you pick.
The published schema takes any integer from 4 to 30, or -1 to let the model choose.
Not just stills — motion and pacing can be supplied as footage and a track.
Pin both ends of a shot when the start and finish are non-negotiable.
Pick 2.5 when the thing you need is longer than a shot, or when you already own the look. A thirty-second take that holds one character is worth more than four eight-second takes that nearly do, because the near-misses cost you an edit and usually a re-run. The same goes for references: if you have character sheets, a product photo, a colour key, or footage whose camera move you want copied, handing them over as references is more reliable than describing them and hoping. Where the Seedance 2 tiers still win is volume — they are cheaper per second, so prompt-finding and throwaway runs belong down there. The prompt transfers up.
One generation covers a beginning, a middle and an end, so there is no seam to hide.
References carry identity, motion and rhythm far more reliably than adjectives do.
All three tiers run the same model, so find the shot on the cheap one and spend on the take you keep.
What this tier accepts, and what comes out the other side.
Any duration in range in one generation, or -1 to let the model pick the length.
Character sheets, product shots or colour keys, all in the same request.
Copy a camera move from footage; drive pacing from a track. Reference footage totals up to 30s.
Sound produced with the picture, as a switch on the request rather than a second job.
1:1, 4:3, 3:4, 16:9, 9:16, 21:9 or adaptive — vertical comes straight out of the model.
Pin either end of a shot and let the model generate the path between them.
The generation parameters for this tier, from the provider's published schema.
Jobs where the length and the references are the point.
Setup, turn and payoff inside one generation, rather than three clips and an edit.
Reference images hold identity for the full take, which is what short clips keep breaking.
Feed footage whose movement you want copied instead of describing the move in words.
Reference audio drives pacing, so moves and changes land on the beat.
One credit pool covers Nano Banana images and Seedance and Veo video. The cost shows before every run, the safety check runs before the model does, and a failed run is refunded automatically. Use a subscription for ongoing work, or a one-time pack when you just need to top up.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
One-time top-ups — buy extra credits any time you run low.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
Charges appear as “SAYMAKER AI” on your card statement. You can cancel any time from Settings → Billing; cancellation takes effect at the end of the period you already paid for. Operator details, the full model list, and the refund window are on the about page.
Powered by
Use one balance across every image and video model on the shelf — the main ones are listed below — and check the credit cost before each request.
Video models
Image models
Model credit guide
A failed run is refunded automatically. The exact estimate in the generator varies by model, length, resolution, audio, and number of images.
| Type | Model | Credit cost |
|---|---|---|
| Video | Seedance 2.5 | The long-take tier, priced per second: 5s 480p ≈ 525 credits, 5s 720p ≈ 1,185. A full 30s take runs ≈ 3,150 at 480p and ≈ 7,090 at 720p. |
| Video | Seedance 2.0 | Supplying a starting frame or clip costs less than starting from words alone: 6s 720p ≈ 565 credits from an image, ≈ 925 from text. Scales with resolution and length. |
| Video | Seedance 2 Fast | Faster and lower cost. 6s 720p text-to-video ≈ 745 credits. |
| Video | Seedance 2 Mini | The cheapest tier of the Seedance 2 family. 6s 720p ≈ 465 credits, 6s 480p ≈ 215 credits. |
| Video | Seedance 1.5 Pro | Audio doubles the rate. 6s 720p ≈ 85 credits silent, ≈ 160 with audio; 1080p ≈ 175 and ≈ 340. |
| Video | Veo 3.1 | Billed per video, not per second, and the tier you pick is the whole price: Lite ≈ 115 credits at 720p and ≈ 135 at 1080p, Fast ≈ 225 at 720p, Quality ≈ 940 at 720p. |
| Video | Kling 3.0 | Audio raises the rate by about half. 6s 720p ≈ 320 credits silent, ≈ 455 with audio; 1080p ≈ 405 and ≈ 610. |
| Video | MiniMax H3 | Fixed 2K, no resolution ladder. Priced per second — a 5s clip ≈ 395 credits. |
| Image | Nano Banana 2 | Generate or edit from text and images. 20 credits per 1K image, 30 at 2K, 45 at 4K. |
| Image | Nano Banana Pro | Consistent run times across generations. 30 credits per 1K or 2K image, 55 at 4K. |
| Image | Nano Banana 2 Lite | Faster, simpler variant, and the everyday editing price: a flat 15 credits per image at every resolution. |
| Image | GPT Image 2 | The lowest-cost premium image model here. 10 credits per 1K image, 15 at 2K, 30 at 4K. |
| Image | GPT Image 2.5 Flare | OpenAI's newest image model, on the fast tier. 25 credits per 1K image, 40 at 2K, 60 at 4K — same price whether you generate or edit. |
| Image | GPT Image 2.5 Sunburst | The 2.5 tier for edits that touch only what you named. Same price as Flare: 25 credits per 1K image, 40 at 2K, 60 at 4K. |
| Image | Seedream 5.0 Lite | Flat pricing — 20 credits per image whether you generate or edit. |
Answered from the published schema and our own runs.
Three things. Length: 2.5 generates 4 to 30 seconds in a single call, where the 2.0 tiers top out well short of that. Inputs: 2.5 takes reference video and reference audio alongside images, so motion and rhythm no longer have to be described in the prompt. Editing: it is built for changing part of an existing shot rather than only generating a new one. The 2.0 tiers stay on the site because they are cheaper per second and still suit short work.
Yes — one request, one output file, no stitching on our side. That is the whole point of the tier, and it is why the examples on this page are all exactly thirty seconds rather than a montage.
480p, 720p or 1080p, in 1:1, 4:3, 3:4, 16:9, 9:16, 21:9 or adaptive, as mp4 or mov. Vertical 9:16 comes straight out of the model, so a social cut needs no reframing afterwards. 1080p is the delivery tier: same model, same references, more detail held through a long take.
Reference images set identity and look, and the model holds them across the full take — that is what makes a thirty-second shot of the same character possible. A reference clip supplies motion and camera language, so point it at footage whose movement you want, not whose content you want. Reference audio drives pacing, so a track with a clear beat produces moves that land on it. Total reference footage can run to thirty seconds.
Per second, with the estimate shown before you run it. A 5-second clip is about 525 credits at 480p, 1,185 at 720p and 2,970 at 1080p; a full 30-second take is about 3,150, 7,090 and 17,815. The sensible workflow is the cheap one: test the idea at five seconds, then commit to the full length once the shot is right.
Long takes amplify a vague prompt. At five seconds an underspecified prompt simply picks something; at thirty it has time to drift — the camera wanders, or the scene resolves into a different one than you had in mind. The fix is sequence: say what happens in what order. Reference images tighten identity considerably, but they do not substitute for telling it where the shot is going.
The 2.5 tier is the only engine here that holds a single take up to 30 seconds, so anything story-length starts here. Veo 3.1 tops out shorter but bundles native audio into one pass — for a short clip with sound, run Veo; for length, nothing else on SayMaker competes.
Open the generator, hand it your references, and let it run the full thirty seconds.