LTX's open-weights video model, running in your browser instead of on your own GPU. Text or image in, up to 4K and 20 seconds out, with a synchronised audio track — and the resolution you pick is the resolution you pay for.
Four habits, in the order they matter.
4K costs eight times as much per second. Find the shot you want cheaply, then re-run the keeper at the size you actually need.
A slow push, a static tripod, a handheld drift. Motion is the thing a still image cannot give you, so it is the thing worth spending words on.
Rain on metal, a room tone, distant traffic. The audio is generated from the same prompt, so leaving sound undescribed leaves it to chance.
Image-to-video pins the composition, palette and subject exactly, and lets the prompt spend all its words on movement instead.
The latest open-weights video model from LTX, the generative media company spun out of Lightricks. It generates synchronised audio and video together, natively up to 4K and 50fps, from either a text prompt or a still image.
Two things separate this model from the flagships it competes with. The first is that audio is not an add-on: sound is generated with the picture in the same pass, so a clip arrives finished rather than needing a track laid under it. The second is that the weights are open — published on Hugging Face with a day-one ComfyUI integration, and free to use for organisations under $10M in annual revenue.
That is why the clips people post are so often labelled with the GPU they ran on: this is a model you can host yourself if you have the hardware.
Most people do not have that hardware, which is what this page is for. Running LTX 2.5 here costs no setup: pick 720p, 1080p, 2K or 4K, pick a length between 2 and 20 seconds, and the credit estimate updates before you commit. A 6-second 720p clip took 23 seconds end to end on the run that produced the alley example above.
The resolution ladder is the steepest of any video model on SayMaker — 4K costs eight times what 720p does per second — so the picker defaults to 720p rather than quietly starting you on the expensive rung.
Creative engine
Prompt or image in, 720p to 4K out, 2 to 20 seconds, audio included. Cost is shown before you generate.
Synchronised sound generated with the frames, included in the per-second price rather than charged as a premium.
Published on Hugging Face and integrated into ComfyUI on day one — run it on your own GPU, or run it here with none.
Pick it when the clip has to arrive finished, long, or sharp.
Ambience, footsteps, rain on metal — generated with the picture instead of sourced afterwards from a library that never quite matches.
Up to 20 seconds in one generation, which covers a whole beat rather than a fragment you have to loop.
Native 4K rather than an upscale, for anything going on a large screen or surviving a crop.
What the model accepts here, measured against the live API.
Start from a written prompt, or hand it a still and let it move — the first frame is yours to set.
720p, 1080p, 2K and 4K. The price moves with the rung, so the picker starts at 720p.
Length is a fixed ladder rather than a free number, and the longer rungs cost proportionally.
16:9 and 9:16, decided before generation so a vertical clip is composed vertically rather than cropped.
Read off the live API this model runs on here, on 2026-08-17.
Mostly work where a silent three-second loop would not have done the job.
Close work where 4K earns its cost — texture, water, small moving detail that falls apart at lower resolution.
9:16 composed as 9:16, with sound already attached, ready to post without a pass through an editor.
The 20-second ceiling covers a full atmospheric beat — a street, a landscape, a room waking up.
A photograph, a render or a generated frame, given a camera move and an ambient track.
One credit pool covers Nano Banana images and Seedance and Veo video. The cost shows before every run, the safety check runs before the model does, and a failed run is refunded automatically. Use a subscription for ongoing work, or a one-time pack when you just need to top up.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
One-time top-ups — buy extra credits any time you run low.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
Charges appear as “SAYMAKER AI” on your card statement. You can cancel any time from Settings → Billing; cancellation takes effect at the end of the period you already paid for. Operator details, the full model list, and the refund window are on the about page.
Powered by
Use one balance across every image and video model on the shelf — the main ones are listed below — and check the credit cost before each request.
Video models
Image models
Model credit guide
A failed run is refunded automatically. The exact estimate in the generator varies by model, length, resolution, audio, and number of images.
| Type | Model | Credit cost |
|---|---|---|
| Video | Seedance 2.5 | The long-take tier, priced per second: 5s 480p ≈ 525 credits, 5s 720p ≈ 1,185. A full 30s take runs ≈ 3,150 at 480p and ≈ 7,090 at 720p. |
| Video | Seedance 2.0 | Supplying a starting frame or clip costs less than starting from words alone: 6s 720p ≈ 565 credits from an image, ≈ 925 from text. Scales with resolution and length. |
| Video | Seedance 2 Fast | Faster and lower cost. 6s 720p text-to-video ≈ 745 credits. |
| Video | Seedance 2 Mini | The cheapest tier of the Seedance 2 family. 6s 720p ≈ 465 credits, 6s 480p ≈ 215 credits. |
| Video | Seedance 1.5 Pro | Audio doubles the rate. 6s 720p ≈ 85 credits silent, ≈ 160 with audio; 1080p ≈ 175 and ≈ 340. |
| Video | Veo 3.1 | Billed per video, not per second, and the tier you pick is the whole price: Lite ≈ 115 credits at 720p and ≈ 135 at 1080p, Fast ≈ 225 at 720p, Quality ≈ 940 at 720p. |
| Video | Kling 3.0 | Audio raises the rate by about half. 6s 720p ≈ 320 credits silent, ≈ 455 with audio; 1080p ≈ 405 and ≈ 610. |
| Video | MiniMax H3 | Fixed 2K, no resolution ladder. Priced per second — a 5s clip ≈ 395 credits. |
| Image | Nano Banana 2 | Generate or edit from text and images. 20 credits per 1K image, 30 at 2K, 45 at 4K. |
| Image | Nano Banana Pro | Consistent run times across generations. 30 credits per 1K or 2K image, 55 at 4K. |
| Image | Nano Banana 2 Lite | Faster, simpler variant, and the everyday editing price: a flat 15 credits per image at every resolution. |
| Image | GPT Image 2 | The lowest-cost premium image model here. 10 credits per 1K image, 15 at 2K, 30 at 4K. |
| Image | GPT Image 2.5 Flare | OpenAI's newest image model, on the fast tier. 25 credits per 1K image, 40 at 2K, 60 at 4K — same price whether you generate or edit. |
| Image | GPT Image 2.5 Sunburst | The 2.5 tier for edits that touch only what you named. Same price as Flare: 25 credits per 1K image, 40 at 2K, 60 at 4K. |
| Image | Seedream 5.0 Lite | Flat pricing — 20 credits per image whether you generate or edit. |
Common questions about generating with this model online.
You can start in the generator on this page, and every new account gets a credit balance to spend on it. Credits are shared across every model on SayMaker, so nothing is locked to one of them. Video is the most expensive thing on the site to run, so a free balance goes further at 720p than at 4K.
Not here. The weights are open, so people do run it locally — the clips posted online are often labelled with the card they were rendered on. Running it on this page needs nothing but a browser; the hardware question becomes a credit question instead.
Yes, in the same pass as the picture rather than as a second step, and it is included in the per-second price. That is one of the model's defining features. Describe the sound you want in the prompt — leave it out and you still get audio, just not audio you chose.
135 credits for 6 seconds at 720p, 270 at 1080p, and 1080 at 4K. The ladder is steep on purpose because the underlying cost is: 4K is eight times the per-second rate of 720p. The estimate is on screen before you generate.
Up to 20 seconds, chosen from a fixed ladder rather than typed as a free number. Longer clips cost proportionally more, so the usual approach is to find the shot at a short length and only extend the one that works.
Yes — image-to-video takes your still as the first frame and moves from there, which is the most reliable way to control exactly what the clip looks like. The prompt then only has to describe motion and sound rather than the entire scene.
Different shapes: this one sells a resolution ladder up to 4K and lengths to 20 seconds, while H3 runs a single flat rate with no resolution choice. Both are on SayMaker, so the honest answer is to put one prompt through each — and there is a full side-by-side of the two in the blog if you want the numbers first.
LTX 2.5 Fast at 720p costs a fraction of H3 and its audio is free, so for drafts and social clips it is usually enough. H3 justifies its rate when you need 2K fidelity or reference-image consistency. We wrote a full side-by-side — see the LTX 2.5 vs MiniMax H3 comparison on the blog.
Text or a still in, up to 4K and 20 seconds out, sound already attached. Start at 720p and only pay for size once the shot is right.