Blog · September 18, 2026 · 17 min read

AI Image Generation Compared

Graphic title card for the article “AI Image Generation Compared” with a stylised video frame with editing timeline.
Grafik: HumanITy

In a working business it is not the single most beautiful picture that decides, it is whether the tenth picture still matches the first. The vendors say strikingly different amounts about exactly that, and it makes for a better comparison than sample galleries do. Google documents up to 14 reference images per request for the Gemini 3 image models, of which the Pro model takes up to five character images for character consistency and up to three as pure style references. Midjourney separates style from subject: a Style Reference transfers colors, medium, texture and lighting according to the documentation, and expressly no objects or people, and a style code once drawn stays the same across rerolls and variations. Adobe takes the detour of a model of your own, for which you upload 10 to 30 of your own images and train an illustration, photographic or character style. Anyone running open weights locally trains a style LoRA, for which Black Forest Labs names 20 to 40 images as the optimal dataset size. And OpenAI writes the weak spot into its own limitations list: the model may occasionally struggle to keep recurring characters or brand elements visually consistent across multiple generations. This is a documentation comparison, not a controlled quality test of text accuracy, style fidelity or German prompts. All figures are as of September 2026.

I do not generate AI images as an end in themselves, I generate them as supply: as B-roll in a reel, as the background for a graphic. So price is not my first question. Which tools have a free tier, and what you are legally allowed to do with the result, is covered separately in Creating AI Images: Free and Free to Use. This post is about the question after that: what do you work with when you need not one picture but twenty that look like they came from one hand. Everything I collect about AI in daily work is on the AI hub; how I decide what gets handed over at all is in Which Tasks You Can Hand Off to AI.

Seven criteria, and why one of them beats the rest

  1. Image quality for your purpose. Not in general, but for your subject: a product shot, a comic panel and a photorealistic portrait are three different jobs. Test your most frequent subject, not the sample on the landing page.
  2. Handling of non-English prompts. Does the tool understand your prompt, or does it translate it first?
  3. Text inside the image. The classic weak spot, and one of the few the vendors themselves concede in the small print.
  4. Style consistency across several images. Does the look hold across a series, or does every picture start from scratch?
  5. Editing existing images. Can you change an existing picture on purpose instead of rolling the dice again?
  6. Resolution and formats. Which pixel dimensions, which aspect ratios, which file formats come out?
  7. Your own style, reusable. Can your look be captured and called up again on the next job?

Points 4 and 7 belong together and beat all the others in my business. A single impressive image is available on all five routes. The difference is made by the tenth tile of a series that is supposed to sit next to the first without anyone noticing.

What the vendors say about recognizability across a series

Tool What is documented for a series What it hangs on
Gemini 3 (Nano Banana) up to 14 reference images per request; on the Pro model up to 5 character images for character consistency and up to 3 pure style references you have to supply the references with every request, they are not a stored profile
GPT Image (OpenAI) several input images as references, editing over multiple rounds in the same conversation OpenAI names consistency across multiple generations as an express limitation
Midjourney Style Reference --sref with a strength dial --sw (0 to 1000, default 100), style codes, moodboards as a style profile of your own Style Reference carries only the look, not objects or people; characters use a separate reference type
Adobe Firefly custom model on 10 to 30 of your own images, either an illustration style, a photographic style or a character described as a beta for certain Creative Cloud plans at the time of writing
FLUX.2 [klein] locally style LoRA, optimal dataset size 20 to 40 images per the vendor open weights, but different licences by variant: 4B under Apache 2.0, 9B under the FLUX Non-Commercial License; training, compute time and upkeep are on you

Three things stand out in that table. First, the difference between a reference and a profile: with Gemini you hand over your exemplars again with every request, while Midjourney and Adobe produce something you call up later. Midjourney even documents that for moodboards a code once generated keeps working even if you delete the moodboard afterwards. That sounds like a detail and is in fact the question of whether your style is a file or a habit.

Second, the separation of style and subject that Midjourney draws in its documentation: Style Reference takes the look, not the content. To hold a character across several pictures you need a second reference type there.

Third, the most honest line in the table. OpenAI lists among its limitations that the model may occasionally struggle to keep recurring characters or brand elements consistent across multiple generations, and adds that it may not place elements precisely in layout-sensitive compositions. Read that before planning a 20-tile series and you save yourself an afternoon.

Text inside the image: what the vendors themselves concede

This is where the marketing image and everyday use sit furthest apart. Google describes an enhanced text rendering for the Gemini 3 image models that it says can generate legible, stylized text for infographics, menus, diagrams and marketing assets. That is a vendor statement, not a measurement with a disclosed method. OpenAI, by contrast, files text rendering expressly under its limitations: significantly improved, but the model can still struggle with precise text placement and clarity. Adobe lists a separate entry for distorted text in generated images among the known limitations of Firefly; that page is dated December 2025, so it is older than the rest of this post.

Black Forest Labs publishes concrete prompt rules for FLUX, and they work as a checklist for any tool: put the exact text in quotation marks, move the text description to the front, name the color and the effect, give brand colors as hex codes, and above all keep it short, because long strings are harder to render accurately.

The check you can run yourself takes five minutes: have your tool put the same short text into the same subject three times, once as a single word, once as a line with an umlaut or accent, once as two lines with digits. If umlauts, hyphens and digits survive all three runs, you can trust the tool with captions.

Prompt language: translated or understood

The clearest vendor statement on this point comes from Adobe. Firefly supports prompts in over 100 languages, by way of machine translation into English through Microsoft Translator. Adobe itself writes that generations based on translated prompts may be inaccurate or unexpected because of the nuances of each language. So if you write a linguistically precise German or French prompt in Firefly, you are working against an intermediate step.

Black Forest Labs claims the opposite for FLUX.2 and lists multilingualism as a use case of its own: the model understands several languages, and a prompt in the language of the content often produces more culturally authentic results than a translated English one. The documentation shows a German example prompt for it.

Midjourney gives a note on input style that matters just as much in practice: short, simple prompts typically generate the best images, while long lists and detailed instructions tend to confuse the process. A prompt with three subordinate clauses is therefore, by Midjourney's own recommendation, the weaker input style there.

For Gemini and OpenAI, the documentation reviewed here gives no directly comparable statement about German prompt language. That gap is not a reason to guess; use the five-minute test from the previous section.

In practice I run two tracks: the subject description in whichever language the tool understands according to its documentation, and the text that is supposed to appear in the picture always quoted word for word, with its accents and umlauts, exactly as it has to end up.

Editing, resolution, formats

Tool Changing an existing image Resolution and format
Gemini 3 (Nano Banana) editing in conversation, a mask can be defined for part of the image 1K, 2K and 4K, plus 512 pixels on the Flash model; 9:16 at 2K is 1,536 x 2,752 pixels
GPT Image (OpenAI) dedicated edits endpoint, a mask marks the area to replace, editing across several rounds recommended 1024x1024, 1536x1024, 1024x1536; custom sizes as multiples of 16, ratio between 1:3 and 3:1, no edge above 3840 pixels; PNG, JPEG, WebP, transparent background possible
Midjourney editor on the website, reference types for editing depending on the model version in version 8.2 at 1:1, per the vendor, 2048 x 2048 pixels as HD and 1024 x 1024 as SD, aspect ratio via --ar
Adobe Firefly editing functions in Firefly and in Adobe's own applications training images for your own model at least 1000 pixels, aspect ratio up to 16:9
FLUX locally and via API dedicated tools for masked removal, extending beyond the edge and deblurring; editing with up to 10 reference images output up to 4 megapixels

Two things about this matter more in daily work than they look. One is the difference between aspect ratio and pixel dimensions, which Midjourney puts into its own documentation: an aspect ratio is not an image size, the actual dimensions depend on the version and the upscaler. For a reel in 9:16 what the tools deliver is usually plenty; for a printed piece, Midjourney works out that 2048 pixels at 300 dpi come to roughly 6.8 inches of edge length, and points to third-party upscalers for more.

The other is the mask. A tool that can only generate anew sends you back to the start on every correction, and with every new roll you lose the match to the rest of the series. A tool that is designed to confine the change to a marked area holds a series together better. The vendors do not guarantee that every pixel outside the mask stays identical, so lay the before and after over each other once before you rely on the mask. If you have to choose between two otherwise equal candidates, still take the one with the mask.

Where these images end up in my business

A generated image is not a result here, it is a component. The editing runs through Eddi, my AI employee for video: he gets raw material and a script and returns a finished, technically checked video, with subtitles, overlays and in 9:16. He works on the command line with ffmpeg, a locally running Whisper model and Remotion, he does not film himself, and he does not overwrite raw material. For generations at an AI image service he has a rule of his own: they cost money, and larger series only run after I approve them. How he works and where his limit sits is on his page, Eddi.

That approval is exactly where the criteria in this post come together. Approving a series means knowing beforehand that the images will match, because a series you have to generate twice costs twice. Which program does the editing around it is something I went through in AI Video Editing Software Compared; how the roles are distributed here overall is in My AI Workforce.

Frequently asked questions

Which AI image generator is the best?

That can only be answered per criterion, not in general. For text inside the image Google makes the furthest-reaching claim, while OpenAI names an express limitation. For a reusable style of your own, Midjourney with style codes and moodboards, Adobe with custom models and the local route with a LoRA each offer a different mechanism. Decide by your most frequent job, not by a ranking.

Which tool holds a style across a whole series?

Most likely the ones that store the style instead of describing it anew every time: Midjourney via style codes and moodboards, Adobe via a custom model from 10 to 30 of your own images, a local setup via a LoRA from 20 to 40 images. With Gemini you supply up to three style references per request on the Pro model. Test it on ten images before you commit.

Can image generation be run locally?

Yes, using open model weights on your own hardware. FLUX is a model family with different licences: Black Forest Labs provides the 4B variant of FLUX.2 [klein] under Apache 2.0 and the 9B variant under the FLUX Non-Commercial License. Running locally and using commercially are therefore separate questions. Its style-LoRA guide names 20 to 40 images as the optimal dataset size, saying that below 20 the variation is missing and above 40 the style gets diluted. You pay no fee to a provider, but you pay in compute time, setup and upkeep. What that means for licensing is in my post on free AI images.

Does AI reliably write text into an image?

Not reliably enough to go unchecked. OpenAI still lists text rendering among the model's limitations despite improvements, and Adobe lists distorted text as a known limitation. The prompt rules from Black Forest Labs help: exact text in quotation marks, keep it short, name the color and the character of the type. For captions that matter, set the text into the image yourself afterwards.

Do I have to write my prompts in English?

That depends on the tool. Adobe machine-translates prompts from over 100 languages into English and points out itself that translated prompts can produce inaccurate or unexpected results. Black Forest Labs expressly recommends the language of the content for FLUX.2. Midjourney advises short, simple prompts in general. Text that is meant to appear in the image you always supply word for word.

Can I edit existing images with AI instead of generating new ones?

Yes, and for series work that is the more important route. OpenAI offers a dedicated endpoint for edits where a mask marks the area to be replaced. Google describes multi-round editing in the same conversation. FLUX brings its own tools for removing, extending and deblurring. Whatever the tool, compare before and after because generative editing does not guarantee identical pixels outside the intended area.

How to continue

Run the series test before you commit, and use a subject you need anyway. Generate ten images of the same visual world, not ten variations of the same image: different subjects, the same look, the same brief and the same number of references per run. Also record whether the tool fixes a seed or other random behaviour. Then put image one and image ten side by side and see whether anyone spots the break. That check tells you more about a tool's fit for your business than any sample gallery, and it costs less than a failed set of tiles.

And if you then want to know how these images make their way into a finished video instead of sitting in a downloads folder: that exact workflow, from raw material through the edit to approval, is what I show in my community, together with the ready-made templates for the employees who handle it here.

Kevin Welter

Kevin Welter

Developer, IT architect, author of technical books (Kubernetes, cloud infrastructures) and speaker. Runs his business with an AI workforce of fourteen AI employees and shows solo business owners in his community how to hire their first AI employee.

More about AI employees

Your first AI employee up and running within an hour

Join the community