On September 20, Alibaba's Qwen team released the trained weights for Qwen-Image-2.1, a model that unifies image generation and editing. Although its image-generation component is held to 7B parameters, it scores higher than Google's Nano Banana 2.0 and OpenAI's GPT Image 1.5 on the team's own benchmark. It still trails the leader, GPT Image 2.5 Sunburst, and open weights do not mean free commercial use. To judge whether it can stand in for a closed model, you need to look separately at the score gap to each competitor and at what you can control yourself across the image-production workflow.

AD

Close to Nano Banana 2.0, 6.73 points behind GPT Image 2.5

In the comparison chart Qwen published, Qwen-Image-2.1 ranks 7th among 29 models, with an overall score of 60.28. Above it are models from OpenAI and xAI, as well as Qwen's own Qwen Image 3 Pro. The claim that it "beat closed models" means different things depending on which competitor you compare it with.

The table below pulls the main models from the September 20 chart and shows how many points each is above or below Qwen-Image-2.1. Model names follow the chart's notation.

Model Qwen-Image-Bench overall score Compared with Qwen-Image-2.1
GPT Image 2.5 Sunburst 67.01 6.73 points higher
GPT Image 2 64.69 4.41 points higher
Qwen Image 3 Pro 62.36 2.08 points higher
MAI Image 2.5 Pro 61.02 0.74 points higher
Qwen Image 2.1 60.28 Baseline
Nano Banana 2.0 59.82 0.46 points lower
GPT Image 1.5 59.65 0.63 points lower
Seedream 5 Pro 59.53 0.75 points lower
Nano Banana Pro 59.45 0.83 points lower

Source: the same Qwen chart. The comparison column takes each row's model as the subject and shows its difference from Qwen-Image-2.1's 60.28 points. These are not independent measurements or win rates. The chart gives no error ranges or number of repeated runs, so the statistical significance of narrow gaps cannot be judged.

Qwen-Image-2.1's 60.28 is 0.46 points above Nano Banana 2.0 on the same chart, but 4.41 points short of GPT Image 2 and 6.73 short of GPT Image 2.5 Sunburst.

The fair reading, then, is that a downloadable option has entered the cluster of models scoring close to some of Google's models and to GPT Image 1.5. It is not a result showing it took the top spot, and a 0.46-point gap does not mean it always produces better pictures than Nano Banana.

Moreover, this comparison is a text-to-image evaluation. The ranking does not prove the features Qwen-Image-2.1 newly promotes, such as preserving people, localized editing and easier handling of transparent images. Being competitive at generation and being trustworthy for editing work are separate things that each need to be verified.

What the 60.28 measures

Qwen-Image-Bench is not a scheme that scores only how good a picture looks. According to the public dataset description, it evaluates quality and aesthetics and alignment with instructions, as well as real-world fidelity and creative expression. The whole is divided into 5 major categories, 23 sub-capabilities and 56 fine-grained items. The design separates, for example, whether text is correct, whether specified objects' relationships are depicted, and whether the result works as a design.

The scoring model, Q-Judger, rates each item as fail, pass or excellent, converted to 0, 60 and 100 points respectively. Scores are averaged up through sub-capabilities and major categories, and the category averages are then averaged per prompt. Finally, 1,000 prompts are averaged, so 60.28 cannot be read as "60.28% of images succeeded" or "it beats competitors 60.28% of the time."

This aggregation method has a practical aim. The design paper by Niantong Li and colleagues explains that it weights an improvement from fail to pass more heavily than one from pass to excellent: 60 points versus 40. It is a scoring scheme that prizes lifting broken images to a usable level, and so differs in character from a popularity vote on "beautiful pictures." The paper is a preprint published in May 2026 and is not an independent verification of this 2.1 release.

The evaluator also has a provenance. Q-Judger is built on Qwen3.6-27B, and its developers say they trained it on more than 130,000 image-instruction pairs rated by 80 art-specialist evaluators. A rank correlation of 0.92 with experts is reported, but that does not mean the 80 people directly reviewed the 29 models here. That the evaluator comes from the same development team does not invalidate the results, but the robustness of the ranking can only be confirmed through follow-up tests with other evaluators or humans.

In addition, the detailed table in the public README is an older version covering 18 models, scored on images generated from Chinese-language prompts, and it says English results will be published later. The 29-model chart announced this time carries no item-by-item scores for Qwen-Image-2.1 and no run settings. We cannot assert that the old table's Chinese-language conditions apply unchanged to the new chart, nor that it beat Google or OpenAI at rendering Japanese text.

So someone making posters for Japan needs to know, beyond the overall rank, how much it can reduce Japanese typos and text-layout errors. For editing portrait photos, what matters is whether it can change only the specified areas while keeping faces and clothing intact. The overall score helps narrow down candidates to try, but it does not measure success or failure for each use case.

AD

The computation outside the 7B

The 7B that Qwen cites is the parameter count of the DiT, the part that generates images. The official implementation description specifies, outside it, Qwen3-VL 8B, which understands text and reference images. A VAE, which reconstructs images from their compressed representation, is also required. Judging the size of the full model set or the GPU memory needed from the 7B alone would misread the actual configuration.

Care is also needed in reading the announcement chart. In the lower panel, the closed models' bars extend to around the 80B mark, but they are labeled with a lock symbol and "parameter count undisclosed." That height is not an actual figure. You cannot divide by that bar to conclude "equivalent quality at less than one-tenth the size."

On the other hand, the computation savings described alongside the miniaturization are concrete. Image generation repeats a denoising computation. Because the reference images and editing instructions do not change during that process, Qwen-Image-2.1 saves the input-side computation result at the first step and reuses it in subsequent steps. The more images are involved, the more it matters to cut the waste of reprocessing the same input many times.

The default generation uses 40 steps, but the cache saves only the input-side recomputation; the process of updating the output image continues. It is not a mechanism that makes things 40 times faster. To judge whether it is faster or cheaper than closed models, you need to measure processing time and cost at the same resolution and the same number of reference images.

The official team also recommends an auxiliary model that expands short instructions into detailed descriptions. It is a fine-tuned Qwen3.5-VL 9B, released separately for generation and for editing. This is an optional preprocessing step, not a required component that runs at every step, but the sentence a user enters and the text actually passed to the image model may differ. Even in a test that enters the same instruction as a competitor, whether rewriting was enabled should be recorded as a condition. The announcement chart does not let us confirm that setting.

The significance of having the weights is that users can inspect this configuration and vary the conditions themselves. For adoption decisions, being able to choose where to quantize, what to place on the GPU, and whether to rewrite instructions is more useful than stopping at "it's light because it's 7B."

Do transparent images and 10-image references set it apart from competitors?

Qwen-Image-2.1 handles both ordinary images and RGBA images with transparency, in generation and editing, within the same model. It consolidates the transparent-image capability that Qwen introduced in December 2025 with the dedicated Qwen-Image-Layered model into a single model that handles both generation and editing. The official team shows examples such as changing an expression while keeping a transparent background, fixing text inside a transparent layer, and extracting a subject from a photo.

However, transparent backgrounds and multi-image references themselves exist in closed models too. Lining up the official materials of each company as of September 21, the presence of a feature alone does not explain the differences.

Feature compared Qwen-Image-2.1 Official specs of closed models What the comparison shows
Transparent background Generates RGBA and edits transparent images directly OpenAI's GPT Image 2.5 Sunburst can output transparent backgrounds; GPT Image 2 also supports them in preview Transparent-background output is not unique to Qwen
Multiple reference images Up to 10 Google's Nano Banana 2 accepts up to 14, with a breakdown of 10 objects handled at high fidelity and 4 people for consistency The number of images alone cannot support a claim of Qwen's advantage
Output resolution Native 2K; default for square is 2,048×2,048 Google advertises up to 4K for Nano Banana 2 Supporting 2K and image-quality superiority must be evaluated separately

Sources: Qwen official implementation, OpenAI API specification, Google image generation guide. This compares supported features and is not an image-quality test under equal conditions. Google's API name is Nano Banana 2 (Gemini 3.1 Flash Image), while Qwen's comparison chart labels it Nano Banana 2.0.

The reason to choose Qwen would arise when you want to combine these tasks in your own environment. For example, you could create a transparent asset, place it on a different background, and correct it while preserving the features of a person or product. If the whole pipeline can be verified locally, it is also conceivable to assign the model only the steps it does well. Still, the official examples are success cases, and how often it can preserve fine details, and how much quality degrades over repeated revisions, must be measured separately.

Commercial deployment carries further clear conditions. The Qwen Research License Agreement limits the non-commercial scope to research and evaluation purposes and requires obtaining a separate license for commercial use. Even if the weights can be obtained for free, that does not mean you can run a business service with them as is. Even if the published score is higher than competitors', whether it comes out cheaper once contract terms and operating costs are included cannot yet be judged.

What Qwen-Image-2.1 opens up is a chance to try, with your own inputs and evaluation criteria, a model whose published score is close to leading cloud image models. If it meets the required quality in tests of Japanese text and product-detail preservation, and if the commercial license and operating costs work out, it becomes a concrete candidate for moving part of the production process from the cloud to your own hands.