GPT Image 2 leads the arena, Nano Banana Pro owns 4K, and Ideogram spells nine words out of ten right. A model-by-model guide, plus the endpoints being shut down this year.
Somewhere out there, a marketing team's image pipeline is about to stop working. Google has retired the entire Imagen line, with shutdown starting as early as 17 August 2026. OpenAI pulled DALL-E 2 and DALL-E 3 from its API on 12 May 2026, and GPT Image 1 goes on 23 October 2026. If any tool you use still calls one of those endpoints, it will quietly stop returning images. That is the least glamorous thing in this post and probably the most useful.
The rest of it is about which model to actually pick.
Start with GPT Image 2. It sits at the top of the Artificial Analysis blind-vote arena at 1,369 Elo, well clear of everything else, and it is the model most likely to just do what your prompt said. If you need 4K output or you plan to spend more time editing images than generating them, use Google's Nano Banana Pro instead. If you care more about how an image feels than whether it is literally correct, that is Midjourney, and no benchmark will tell you that.
Everything below is the detail behind those three sentences.
| Model | Best for | Main weakness | Price | Text in image | |---|---|---|---|---| | GPT Image 2 | Prompt accuracy, all-round default | Yellow colour cast, slow, pricey at top quality | ~$0.005 (low) to ~$0.165–0.211 (high) per image via OpenAI API | Excellent | | Nano Banana Pro | Native 4K, conversational editing | Closed, opaque token pricing | $0.134/image at 1K–2K, $0.24 at 4K (Google API, standard) | Excellent | | Nano Banana 2 | Cheap, fast, high quality | Lower ceiling than Pro | $67 per 1,000 images | Excellent | | Midjourney V8.1 | Art direction, concept work | No public API, weak on precise edits | $10 / $30 / $60 / $120 per month | Fair on short strings | | Ideogram 4.0 | Typography, logos, posters | Weaker photoreal faces | Free tier 10/day; $7–$48/month; API $0.03–$0.10 per image | Best in class | | Seedream 5.0 Pro | Photoreal product and people shots | Closed API, older versions oversaturated | $90 per 1,000 images | Very good | | Reve 2.1 | Editable layouts, design work | Faces and hands trail the leaders | $5/month Pro plan | Strong | | FLUX.2 | Best open-weight quality, self-hosting | [dev] licence blocks selling output | Free self-hosted; $12/1,000 [dev], $30 [pro], $70 [max] | Very good | | MAI-Image-2.5 | Image editing | Mostly locked to Microsoft's ecosystem | $48.10 per 1,000; Pro $108.50 | Strong | | Firefly Image 5 | Commercially safe output, Photoshop | Raw quality below the leaders | Included in Creative Cloud | Good |
Prices marked per 1,000 images come from Artificial Analysis, normalised to 1024x1024 at default settings. Real bills move around with resolution, quality tier and batch mode, so check before you commit budget.
Released 21 April 2026. It reasons about the prompt before it renders, which is why it handles multi-subject scenes and awkward layout instructions better than anything else. Give it a paragraph of fussy detail and it will get most of it.
The flaw you will hit in week one: a warm yellow cast on almost everything. Users call it the piss filter, and it seems to come from OpenAI's prompt rewriting adding warm-tone language you cannot switch off. Writing "neutral 6000K, no sepia" into your prompt helps. It should not be necessary.
Cost scales hard with quality. Low quality runs about half a cent an image, high quality up to roughly 21 cents. Editing is supported, with up to 16 reference images.
Nano Banana Pro arrived 20 November 2025 and remains the cheapest route to genuine 4K, at $0.24 per image on standard processing and $0.12 on batch. It holds identity across up to five people in one frame and blends up to 14 objects, which matters if you are building anything with recurring characters. Because it is grounded in Gemini, it can also produce infographics where the labels are factually right rather than plausible-looking gibberish.
Nano Banana 2 landed around February 2026 as the fast, cheap sibling at $67 per 1,000 images. Curiously it scores slightly higher than Pro on the editing arena, 1,254 against 1,248, so for edit-heavy work at volume it is arguably the better buy. Opaxes gives you both alongside the others in this list, which is the easy way to settle an argument like that on your own images rather than someone else's benchmark.
Default since 10 June 2026. Nothing else produces images with as much personality, and for concept art, editorial illustration or album covers that is the whole job.
V8 changed something, though. It follows prompts more literally now, and the community reaction was blunt. One widely quoted Reddit line: "the camera got more precise, but the photographer left." Output can look over-polished in a way people describe as AI sheen. Midjourney's own suggested fix is the --raw flag, which works but hands the problem back to you.
The practical blocker for businesses is that there is no general public API. Discord and web only. You cannot put it in a pipeline.
Released 3 June 2026, and the one to use when the words in the image are the point. Independent 2026 reviews put its short-text accuracy around 90 to 95 percent. Midjourney and Stable Diffusion land closer to 30 to 40 percent. Nine correct spellings out of ten against three out of ten is not a marginal difference, it is the difference between shipping a poster and regenerating it eleven times.
Faces are its weak spot, so it is a bad choice for anything where a person is the hero.
Worth checking the maths on pricing: at volume the subscription badly undercuts the API. Roughly 3,500 Turbo images costs about $42 on the Pro plan versus around $210 through the API.
ByteDance's 5.0 Pro shipped July 2026, after 4.0 in September 2025. This is the one reviewers reach for on product photography and photoreal people: natural skin texture, film-like colour, hands that mostly behave. In hands-on testing, 4.0 handled a fiddly edit (a burger with a gap where the ingredients had been removed) that Nano Banana fumbled.
ByteDance claims 95 percent anatomical accuracy. Treat that as marketing, not a measurement. Earlier versions had oversaturated skin, which reviewers say 5.0 has largely fixed. At $90 per 1,000 images it is not the cheap option.
Reve builds an editable, code-like layout before it renders, so you can re-render one element without redrawing the whole image. Misspelled subhead, logo in the wrong corner: fix that piece, keep everything else. For posters, packaging and UI mockups this saves real time.
It ranks third on the text-to-image arena at 1,322. Faces and hands trail GPT Image 2 noticeably, and the same literalism that makes layouts precise makes loose creative prompts come out mechanical. No mobile app. At $5 a month for the Pro plan it is the cheapest serious tool here.
FLUX.2 launched 25 November 2025 and is the best open-weight option: 4 megapixel output, multi-reference conditioning for consistent characters and products, and the [klein] variant runs on about 13GB of VRAM. Self-hosting is free.
Read the licence before you sell anything. [klein]-4B is Apache 2.0 and fine for commercial use. [dev] is a non-commercial community licence, and people get caught by this.
Qwen-Image is the local favourite for dense text and Chinese characters. But Qwen-Image-3.0, released 21 July 2026, shipped with no open weights, no model card, no benchmark scores and no technical report, which is a step backwards from earlier versions and drew criticism from independent outlets. Every claim about it rests on hand-picked demos. Day-one testers also said the humans look lifeless. Treat it with more caution than its reputation suggests.
Product photos. Seedream 5.0 for the shot itself. Nano Banana Pro if you need it at 4K or need the same product across twenty images.
Logos and posters. Ideogram, on the subscription rather than the API. Nano Banana Pro if the piece also needs a photorealistic element.
Realistic people. Seedream 5.0 or Nano Banana Pro. FLUX.2 if you need to own the pipeline. Avoid Ideogram, Reve and Qwen here.
Art and illustration. Midjourney V8.1, with --raw when the polish gets in the way.
Editing a photo you already have. Microsoft's MAI-Image-2.5 and 2.6 top the editing arena, but the top five (MAI, Reve 2.1, GPT Image 2, Seedream 5.0 Pro) are close enough that the confidence intervals overlap. Nano Banana Pro is the most pleasant to actually use, because you describe the change in plain language. Firefly if legal needs to sign off, since Adobe trains on licensed content and offers enterprise indemnification.
Running it yourself. FLUX.2 [dev] for quality, Qwen-Image for text, SDXL if you need the deep library of LoRAs and ControlNets. Stable Diffusion 3.5 sits awkwardly between them: more VRAM than SDXL, less quality than FLUX.
- Google Imagen, all versions. Shutdown starting 17 August 2026, across both the Gemini Developer API and Agent Platform. Move to gemini-3.1-flash-image. - Gemini 2.5 image models. Shut down October 2026. - DALL-E 2 and DALL-E 3. Removed from the OpenAI API on 12 May 2026. - GPT Image 1. Deprecating 23 October 2026. GPT Image 1.5 survives and is still useful if you need transparency support. - Google ImageFX. Shut down 30 April 2026, folded into Flow. - FLUX.1. Superseded by FLUX.2 as flagship. - Midjourney v6 and v7. Superseded, though v7 is still selectable.
One caveat: some sources reference a "Stable Diffusion 4". That could not be confirmed, so do not plan around it.
Elo scores measure average correctness across thousands of generic prompts. Your brief is not average. A model that ranks fourth overall may be the best one for photographing a matte black bottle on a grey background, and no leaderboard will ever tell you that.
So run the real thing. Take one actual brief, the fussy one with the brand colours and the text overlay, and put it through two models side by side. Score them on the things you care about: did it follow the instruction, is the text readable, how many rounds to an approved asset, what did those rounds cost. You will know within an hour, and it will be a better answer than this post can give you. Opaxes is built for exactly that comparison, but the method matters more than where you run it.