xAI's new image model "Imagine Image 2.0" needs to be understood by separating where it's available from what Arena has measured. On August 7, 2026, xAI made the model widely available in Quality Mode on Grok.com/imagine, as well as on iOS and Android. On Arena, as displayed that same day, it ranked second in both text-to-image generation and single-image editing.

However, that ranking doesn't represent a definitive verdict on the product overall. Arena is a venue that aggregates votes from anonymous head-to-head matchups, and the display for Image 2.0 carries a "Preliminary" label. There's also a gap between the editing features xAI presents and the current state of the API available to developers. Looking separately at the rankings, the creative workflow, and the API reveals what has been opened up to users with this general release, and what remains undetermined for developers.

AD

The "Preliminary" Condition Attached to "No. 2"

On Arena's Text-to-Image leaderboard, "grok-imagine-image-2.0 (low)" ranked second with a score of 1,320±12. The top model, "gpt-image-2 (medium)," scored 1,380±5, and Image 2.0 has garnered 2,722 votes. On the Image Edit leaderboard for single-image editing, Image 2.0 also placed second at 1,439±8, with gpt-image-2 (medium) again in first at 1,463±4. The Image Edit leaderboard shows 53 models and 28,831,297 total votes.

These two "No. 2" rankings are not the result of combining generation and editing scores into a single figure. They come from separate leaderboards with different aggregation conditions. The top model on Text-to-Image shows 69,194 votes, while Image 2.0 shows only 2,722, so the score gap alone cannot be taken as a definitive measure of the performance difference between models.

Furthermore, Image 2.0's entries on both leaderboards carry the "Preliminary" label. Under Arena's policy, scores based on votes collected before public release are treated as provisional until a sufficient number of new votes accumulate after launch. What we can say is that the model achieved a high ranking immediately after release. What this display does not guarantee is long-term ranking stability or overall suitability across the full range of image-creation use cases.

A Feature Set Shifting from Generation to Iterative Editing

For Image 2.0, xAI touts fine-grained instruction following, preservation of text and layout within complex images, and consistency of input content across both generation and editing. These are claims made by xAI itself, not results from an independent text-accuracy test.

The lineup of features suggests a workflow oriented less toward generating a single image and calling it done, and more toward continuously refining an existing image. Magic Wand alters only a specified region, while Segmentation lets users select the area to be edited. The model also supports background removal, and Multi-Ref Editing allows referencing up to five input images. It's a design that shrinks the portion being redone while allowing multiple references to be brought in.

Smart Resize reconstructs a single image into nine different aspect ratios ranging from 1:2 to 2:1. The supported ratios are 1:2, 9:16, 2:3, 3:4, 1:1, 4:3, 3:2, 16:9, and 2:1. The templates target photo editing, product imagery, and marketing use cases, as well as design work, game assets, and streaming emotes.

xAI has shown an example called "Built image by image," in which characters, locations, and props are generated separately while maintaining consistent appearance—positioning this as a stepping stone toward video production. However, this remains within the scope of the company's own demo and explanation. It cannot be read as a feature that guarantees video continuity or reproducibility in commercial production work.

AD

What Exactly Does Arena's Ranking Aggregate?

Arena's ranking is not a static benchmark solving a fixed set of questions just once. Users vote after comparing the outputs of two models without knowing their names, and the results are aggregated using the Bradley-Terry method. The resulting scores therefore represent community preference in anonymous head-to-head matchups.

In February 2026, Image Arena introduced use-case-specific categories and quality filtering, based on an analysis of over 4 million prompts. This is a mechanism designed to address the fact that results in image generation look different depending on the use case. The overall ranking cannot be used as a complete proxy for every possible creative condition.

Arena's publication policy covers not only public APIs but also widely accessible public services. The fact that the model was made generally available on Grok.com and the apps meets Arena's criteria for inclusion, but this does not mean a developer-facing API was released on the same day. Image 2.0 entered evaluation with its consumer-facing entry point and its developer-facing entry point still separate.

The "(low)" label also requires caution in interpretation. While it appears in the model name on Arena, xAI's announcement does not explain any tiering of speed or quality. There is no basis for reading it as a "faster version" or a "lower-quality version."

Between General Availability and API Access

Users can already access Image 2.0 through Grok.com/imagine, iOS, and Android. However, xAI's official announcement lists API access as "coming soon." As of August 7, no model ID, pricing, or regional availability has been disclosed for a dedicated Image 2.0 API.

xAI's API documentation lists an image model with the aliases "grok-imagine-image-quality" and "grok-imagine-image-quality-latest." Its image output pricing is $0.05 per image at 1K resolution and $0.07 per image at 2K. However, that page does not explicitly identify this model as Imagine Image 2.0. The pricing currently shown in the documentation cannot be treated as the API pricing for the 2.0 model announced this time.

In short, the stability of the ranking will become clearer through post-launch Arena voting, and the conditions for developers will become clear once a dedicated API for Image 2.0 is released. Separately from that, how well the partial editing, multi-reference support, and aspect-ratio changes hold up in actual production work is something that needs to be tested independently.