Every New Image Model on Genfire (September 2026): GPT Image 2.5, Seedream 5, Qwen Image 3, Recraft V4.1, Grok Imagine 2.0 and Muse Image
Six new image families landed this season. Reference limits, edit modes, text rendering and aspect ratios compared, with a job-by-job pick list.
Six New Families in One Season
The image side of Genfire moved faster than the video side this summer. OpenAI shipped GPT Image 2.5 in two flavours, ByteDance shipped Seedream 5.0, Alibaba shipped Qwen Image 3, Recraft shipped V4.1 plus a custom-style system, xAI shipped Grok Imagine 2.0, and Meta shipped Muse Image. All are live in the AI image generator and on the /v1 API now.
The honest answer to "which one should I use" comes down to four things: reference images, editing, text inside the picture, and the shape of the frame. What follows is model by model, with facts from Genfire's model registry — then one table and a job-by-job pick list.
GPT Image 2.5 — Sunburst and Flare
OpenAI's newest image model arrived as two endpoints, and the split is speed against precision.
- Variants: Sunburst (precision) and Flare (speed)
- Reference images: up to 16
- Editing: yes on both, including masked inpainting
- Quality control: low, medium, high
- Aspect ratios: 1:1, 16:9, 9:16, 4:3, 3:4
- Images per request: up to 4
- Pricing: identical to GPT Image 2 — same controls, same price table
Flare is OpenAI's own recommended default: better output than GPT Image 2 at roughly half the latency. Point it at everyday work, social assets and anything you generate in volume. Sunburst spends longer per image for extra fidelity on intricate detail and tighter control across an edit chain — it shows when you are six passes into the same frame and need pass six to still look like pass one.
Best for: reference-heavy composition. Sixteen reference images is the deepest pool of any image model here, which makes GPT Image 2.5 the fit for brand systems where a shot has to respect a logo, a product, a colour card and a face at once. Full spec sheet at /gpt-image-2-5.
Seedream 5.0 — Pro and Lite
ByteDance's line jumped from 4.5 to 5.0 with a flagship and a lightweight sibling.
- Seedream 5.0 Pro: flagship photoreal generation and multi-reference editing, up to 2K, up to 10 reference images
- Seedream 5.0 Lite: multilingual text rendering, web-search-aware — currently in preview status on Genfire
- Editing: yes on both
- Aspect ratios: 1:1, 16:9, 9:16
- Images per request: up to 6 on Pro (the highest on the platform), up to 4 on Lite
Pro is for when the brief is "make this look like a photograph". It carries the ad-and-product strength that made Seedream 4.5 popular and adds a real multi-reference edit path. Lite is the oddity: the only model here described as web-search-aware, rendering text in several languages. Because it is marked preview, experiment with it rather than build a pipeline on it.
Best for: photoreal product and ad frames (Pro), multilingual text-in-image experiments (Lite). The narrow ratio set is the trade-off.
Qwen Image 3
Alibaba's typography specialist got a version bump. The headline is bilingual text.
- Text rendering: bilingual Chinese and English, up to 2K
- Reference images: 1–3 in edit mode, addressed in the prompt as "image 1", "image 2", "image 3"
- Editing: yes — edits cap at 1440×1440, while text-to-image goes to 2048×2048
- Prompt length: up to 800 characters
- Aspect ratios: 1:1, 16:9, 9:16, 4:3, 3:4
- Images per request: up to 4
The addressing convention is the part worth internalising. Qwen Image 3's edit mode does not silently blend your references — you refer to them by position in the prompt ("replace the bottle in image 1 with the one from image 2"), which makes multi-image composites predictable. The 800-character cap is real: this model rewards a tight instruction over a sprawling paragraph.
Best for: posters, packaging, signage and anything where a Chinese or English string has to be spelled correctly at 2K. Also the cleanest small-reference composite tool here — three slots with names beats ten slots without. Spec sheet at /qwen-image-3.
Recraft V4.1 — and Recraft V4 Styles
Recraft is the design-system model: where you go when output has to sit inside a brand rather than look impressive on its own. V4.1 sharpens prompt control and produces cleaner composition than V4, and landed as a six-endpoint family.
- Raster: Recraft V4.1 and V4.1 Pro
- Vector: V4.1 Vector and V4.1 Pro Vector — true SVG output, not a traced raster
- Utility runtime: V4.1 Utility and V4.1 Utility Pro — same price and parameters as the standard tiers; the difference is throughput, for high-volume ideation and A/B pipelines
- Aspect ratios: 1:1, 16:9, 9:16, 4:3, 3:4
- Editing: no — the V4.1 family is text-to-image and text-to-vector only
Alongside V4.1 sits Recraft V4 Styles, a different idea. It ships in four variants (standard, Pro, Vector, Pro Vector) and takes a custom style two ways: a style_id you trained earlier, or 1–10 style reference images that train one inline and hand the style_id back. That tray is not a source frame to edit — it is the style, which is why the studio relabels it "Style references".
Best for: logos and flat illustration that has to scale (Vector), large-format production work (Pro), batch exploration (Utility), and locking a house look across hundreds of assets (Styles). More at /recraft-v4.
Grok Imagine 2.0
xAI's latest image generation, and the widest frame options on Genfire.
- Resolution: 1K or 2K
- Aspect ratios: 13 — 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 2:1, 1:2, 20:9, 9:20, 19.5:9, 9:19.5
- Quality tiers: low or medium (medium is the default)
- Editing: yes, up to 3 reference images
- Images per request: up to 4
The 13-ratio list is the reason to care. The 20:9 and 19.5:9 pairs are phone-screen and ultra-wide shapes almost nothing else offers natively, so you get the crop you wanted instead of the crop you cut down to. The quality enum really is low and medium only — Genfire rejects a "high" rather than silently billing you for medium. Its sibling Grok Imagine Pro is the previous generation's Quality-mode endpoint: sharper detail, stronger text rendering, up to 2K, the same 13 ratios.
Best for: stylized work that needs an unusual frame. More at /grok-imagine.
Muse Image
Meta's entry is the instruction-follower, and it has the most specific strengths in the batch.
- Fine detail: built for accurate text, plots and QR codes
- Editing: yes — precise, localized changes, composing from up to 10 reference images
- Aspect ratios: 21:9, 16:9, 4:3, 3:2, 1:1, 2:3, 3:4, 9:16, 9:21 — and omitting the ratio lets Muse size the output from the prompt
- Resolution and quality controls: none — there are no dials to set
- Images per request: up to 4
"Plots and QR codes" is an unusually literal claim for a model description, and it tells you what Muse is tuned for: images where a detail is either correct or worthless. A QR code that is 98% right does not scan. The absent resolution control is a feature in practice — nothing to tune, and omitting the ratio is how you ask Muse to decide the shape from what you described.
Best for: infographics, diagrams, mock UI, labelled product shots, and edits where you want one thing changed and everything else left alone. Spec sheet at /muse-image.
Nano Banana 2 — Still the Default
Worth restating in this company: Nano Banana 2 is still Genfire's default image model. Google's general-purpose model is fast, covers a broad range of subjects, edits with masked inpainting, and is the only model here with a full 1K / 2K / 4K resolution control and the widest ratio set in the studio — 21:9, 9:21, 5:4, 4:5, 3:2 and auto among them.
Nano Banana 2 Lite sits beneath it: ultra-low-latency text-to-image on Gemini 3.1 Flash Lite Image, with no edit endpoint. It generates; it does not revise.
Best for: the first thing you try. 4K plus inpainting covers most of what the other models specialise in. More at /nano-banana.
The Comparison Table
Everything below is what the shipped registry states. Where a cell says "—", the model does not declare that capability rather than being known to lack it.
| Model | Resolution | Reference images | Edit mode | Text rendering | Aspect ratios |
|---|---|---|---|---|---|
| GPT Image 2.5 Flare / Sunburst | Quality low/med/high, no resolution dial | 16 | Yes + masked inpaint | — | 5 |
| Seedream 5.0 Pro | Up to 2K | 10 | Yes, multi-reference | — | 3 |
| Seedream 5.0 Lite (preview) | — | Yes | Yes | Multilingual | 3 |
| Qwen Image 3 | 2K text-to-image, 1440px edits | 3, addressed by name | Yes | Bilingual CN/EN | 5 |
| Recraft V4.1 / Pro | — | — | No | Design-grade | 5 |
| Recraft V4.1 Vector / Pro Vector | True SVG | — | No | Design-grade | 5 |
| Recraft V4.1 Utility / Utility Pro | Same as V4.1 / Pro | — | No | Design-grade | 5 |
| Recraft V4 Styles | Raster or true SVG | 10, as style references | No | Design-grade | 5 |
| Grok Imagine 2.0 | 1K or 2K | 3 | Yes | — | 13 |
| Grok Imagine Pro | 1K or 2K | Yes | Yes | Yes | 13 |
| Muse Image | No control | 10 | Yes, localized | Text, plots, QR codes | 9, or omit |
| Nano Banana 2 | 1K, 2K or 4K | Yes | Yes + masked inpaint | — | 11 |
| Nano Banana 2 Lite | — | — | No | — | 11 |
Which Model for Which Job
A photoreal product shot for an ad. Seedream 5.0 Pro first, GPT Image 2.5 Flare second — Pro gives you six images per request, Flare gives you sixteen reference slots to constrain with. For a paid campaign, the AI ad generator wires the same models into an ad-shaped workflow.
A poster with a headline that must be spelled right. Qwen Image 3 for Chinese or English, Muse Image if the layout also carries a chart or fine print, Recraft V4.1 if the piece has to match an existing brand.
A logo or an icon set. Recraft V4.1 Vector or Pro Vector, because you get a real SVG rather than a raster you have to trace. Add Recraft V4 Styles when the set has to stay consistent across dozens of pieces.
A composite from several photos. Qwen Image 3 for two or three inputs you want to address individually; Muse Image or GPT Image 2.5 for more.
One small change to an existing image. Muse Image for localized edits, Nano Banana 2 when you want to mask the region explicitly. Skip Nano Banana 2 Lite and the Recraft V4.1 family here — neither has an edit path.
An unusual frame. Grok Imagine 2.0 for 20:9, 19.5:9 and their portrait mirrors. Muse Image for 21:9 and 9:21, or omit the ratio. Nano Banana 2 for 5:4 and 4:5.
Hundreds of images on a deadline. Recraft V4.1 Utility or Utility Pro — same price, same parameters, more throughput. Nano Banana 2 Lite for the fastest pass when nothing needs revising afterwards.
No idea yet. Nano Banana 2. It is the default for a reason.
Running Them, and What They Cost
In the browser, open the AI image generator and pick a model — the panel reshapes itself around what that model accepts, and attaching a reference image drops models without an edit path out of the picker rather than failing at generation time. On the API, call POST /v1/images/generations with the model id; the MCP server and the CLI expose the identical set.
Genfire is pay-as-you-go. Credit packs start at $19 for 1,000 credits, purchased credits last 12 months, and every purchase includes a commercial license; optional monthly plans start at $29 a month. GPT Image 2.5 is priced exactly the same as GPT Image 2, and Recraft's Utility runtimes cost the same as the standard tiers they mirror — in both cases the newer option is not the more expensive one. The studio quotes the exact credit cost before you generate; current rates are on the pricing page.
The Short Version
Nano Banana 2 remains the default and the right first try. Past that the batch sorts cleanly: GPT Image 2.5 for reference depth, Seedream 5.0 Pro for photoreal, Qwen Image 3 for bilingual text and named composites, Recraft V4.1 for brand systems and true vector, Grok Imagine 2.0 for unusual frames, Muse Image for detail that has to be literally correct. Open the image studio to try any of them.