Alibaba's image model for pictures that contain writing. Type a prompt with the exact words you want on the page, pick a ratio, and it comes back with those words spelled correctly — posters, infographics, storyboards, signage and UI mockups.
Four habits that decide whether the words come out right.
Write the exact string you want rendered — reading 'MARCH 14-22', not a description of it.
Headline at the top, a second line beneath, credits along the bottom edge.
Condensed grotesk capitals, handwritten script, a thin rule under the header.
Body copy and captions hold together at 2K in a way they cannot at 1K.
This is the third generation of Alibaba's Qwen image line, and the thing it is built around is writing. Most image models treat text as decoration: ask for a shop sign and you get shapes that look like letters from a distance and fall apart up close. This model treats the words as content. Give it a headline, a set of numbered steps, a caption under every panel, and it lays the page out and spells the words.
That difference changes what you can ask for. A poster stops being a background you have to add type to in another tool, and becomes one generation. An infographic with four steps and a paragraph under each one becomes a prompt. The same capability covers newspapers, exam papers, menus, product packaging, game UI and interface mockups — anything where the information in the picture is the point of the picture.
It also renders more than one writing system, so the words do not have to be English. The tier that runs here is Pro, the top rung of the family, at 1K or 2K with a choice of seven aspect ratios.
Creative engine
Two resolutions, seven ratios. The ratio never changes the price.
Headlines, paragraphs and captions come back spelled, not suggested.
Grids, numbered sequences, panels and headers hold their structure.
Start from a prompt, or hand it a picture and describe the change.
Pick it the moment your prompt contains a word in quotation marks. If you are describing a scene — a fox in autumn woodland, a portrait in mixed light — a photoreal model like [Seedream 5.0 Pro](/image/seedream-5-pro) will serve you better, and [Nano Banana Pro](/image/nano-banana-pro) is the one to reach for when a design has to look art-directed. But the second the picture has to say something, the ranking inverts. A model that renders a beautiful café and puts MOFFEE CAFE over the door has not made you a usable image, and no amount of re-rolling fixes it reliably. This is the model that gets the sign right, and then gets the opening hours under it right too. It is also the model to use when the layout carries meaning: step one above step two, a header separated by a rule, a caption tied to the panel it belongs to. Those are structural decisions, and it makes them.
Posters, covers, flyers, packaging — where the type is the design.
Infographics, menus, exam papers, documents held in frame.
In-scene lettering that has to survive being looked at closely.
What the model accepts here, and what each control does.
A prompt in, a finished page out, with the wording carried through.
Hand it a picture and describe the change you want made to it.
1:1, 16:9, 9:16, 4:3, 3:4, 3:2 and 2:3 — portrait, landscape and square.
2K is the one to use when the smallest text on the page still has to read.
Measured against the live API this model runs on here, on 2026-08-05.
The pictures that have something to say.
A headline, a date line and a credit block, finished in one generation.
Numbered steps with real body copy under each one, laid out on a grid.
Panelled sheets where every frame carries its own caption.
Shopfronts, menus, packaging and interface screens with legible labels.
One credit pool covers Nano Banana images and Seedance and Veo video. The cost shows before every run, the safety check runs before the model does, and a failed run is refunded automatically. Use a subscription for ongoing work, or a one-time pack when you just need to top up.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
One-time top-ups — buy extra credits any time you run low.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
Charges appear as “SAYMAKER AI” on your card statement. You can cancel any time from Settings → Billing; cancellation takes effect at the end of the period you already paid for. Operator details, the full model list, and the refund window are on the about page.
Powered by
Use one balance across every image and video model on the shelf — the main ones are listed below — and check the credit cost before each request.
Video models
Image models
Model credit guide
A failed run is refunded automatically. The exact estimate in the generator varies by model, length, resolution, audio, and number of images.
| Type | Model | Credit cost |
|---|---|---|
| Video | Seedance 2.5 | The long-take tier, priced per second: 5s 480p ≈ 525 credits, 5s 720p ≈ 1,185. A full 30s take runs ≈ 3,150 at 480p and ≈ 7,090 at 720p. |
| Video | Seedance 2.0 | Supplying a starting frame or clip costs less than starting from words alone: 6s 720p ≈ 565 credits from an image, ≈ 925 from text. Scales with resolution and length. |
| Video | Seedance 2 Fast | Faster and lower cost. 6s 720p text-to-video ≈ 745 credits. |
| Video | Seedance 2 Mini | The cheapest tier of the Seedance 2 family. 6s 720p ≈ 465 credits, 6s 480p ≈ 215 credits. |
| Video | Seedance 1.5 Pro | Audio doubles the rate. 6s 720p ≈ 85 credits silent, ≈ 160 with audio; 1080p ≈ 175 and ≈ 340. |
| Video | Veo 3.1 | Billed per video, not per second, and the tier you pick is the whole price: Lite ≈ 115 credits at 720p and ≈ 135 at 1080p, Fast ≈ 225 at 720p, Quality ≈ 940 at 720p. |
| Video | Kling 3.0 | Audio raises the rate by about half. 6s 720p ≈ 320 credits silent, ≈ 455 with audio; 1080p ≈ 405 and ≈ 610. |
| Video | MiniMax H3 | Fixed 2K, no resolution ladder. Priced per second — a 5s clip ≈ 395 credits. |
| Image | Nano Banana 2 | Generate or edit from text and images. 20 credits per 1K image, 30 at 2K, 45 at 4K. |
| Image | Nano Banana Pro | Consistent run times across generations. 30 credits per 1K or 2K image, 55 at 4K. |
| Image | Nano Banana 2 Lite | Faster, simpler variant, and the everyday editing price: a flat 15 credits per image at every resolution. |
| Image | GPT Image 2 | The lowest-cost premium image model here. 10 credits per 1K image, 15 at 2K, 30 at 4K. |
| Image | GPT Image 2.5 Flare | OpenAI's newest image model, on the fast tier. 25 credits per 1K image, 40 at 2K, 60 at 4K — same price whether you generate or edit. |
| Image | GPT Image 2.5 Sunburst | The 2.5 tier for edits that touch only what you named. Same price as Flare: 25 credits per 1K image, 40 at 2K, 60 at 4K. |
| Image | Seedream 5.0 Lite | Flat pricing — 20 credits per image whether you generate or edit. |
Common questions about generating with this model online.
You can start in the generator on this page without an account, and every new account gets a credit balance to spend on it. Credits are shared across every model on SayMaker, so nothing is locked to one of them.
Yes. The tier wired here is the Pro one, the top rung of the family, which is the one that holds small text together. It runs at 1K or 2K.
The three examples on this page are the honest answer: a poster headline, four paragraphs of infographic body copy, and six handwritten captions, all generated here on 2026-08-05 and none of them typeset afterwards. Write the exact wording into your prompt and it renders that wording.
Yes — multilingual rendering is one of the things the 3.0 generation was built for. Put the exact characters you want in the prompt.
25 credits at 1K and 45 at 2K, and the aspect ratio does not change either number.
Yes. Upload a source image in the generator and describe the change — the same text-rendering strength applies when you are adding or replacing writing in a picture you already have.
25 credits at 1K and 45 at 2K. What you are paying for is the writing: headlines, captions and numbered steps come back spelled and laid out, which is the one thing most image models cannot do. Banana 2 recovers better when an edit goes wrong. Anything with words in it → Qwen; tricky single edits → Banana.
Write the words you want on the page, pick a ratio, and let it typeset the thing for you.