xAI

Imagine Image 2.0

the one that gets the text right

Released 7 August 20263 min readImage GenerationLast updated:

Editorial illustration: xAI's Imagine Image 2.0, a poster being edited region by region with its small type staying sharp

Key facts

xAIGrok Imagine
Maker
Quality Modegrok.com, iOS, Android
Where
grok-imagine-image-qualityflat per-image pricing
API model
Region-levelwand, segmentation, cut-out
Editing
Up to 3images per edit
References

The pitch is images you can use in real work rather than images you look at. Legible small text and layout that holds together are the claims, and the editing tools are what make it a working surface instead of a slot machine: point at a region and change only that.

Imagine Image 2.0 is xAI’s second-generation image model, generally available as the new Quality Mode on grok.com/imagine and in the Grok iOS and Android apps. xAI states the goal in one line: make images you can use in real work.

That phrasing is doing more than marketing. It separates two things image models are usually judged as one: how good a picture looks, and whether the picture can carry information. A model that renders a beautiful poster with unreadable body text has failed at the second while passing the first.

What it claims to do differently

Typography and layout are planned, not decorated. xAI’s description is that the model “plans typography and layout the way a designer would, so dense, multi-part visuals hold together and small text comes out sharp.” Small legible text is the single hardest thing in image generation and the reason most generated diagrams cannot be published: the headline renders, the labels turn to soup.

Instruction following is tight. The company frames it as following instructions closely, down to the details, and preserving what you put in across generations and edits. Preservation is the harder half. A model that regenerates the whole frame every time you ask for one change is not an editing tool.

The editing tools are the actual product

Three named capabilities turn it from a generator into a surface you can iterate on:

Tool What it does
Magic wand Edits the region you point at and leaves the rest untouched
Segmentation Selects precise areas of the image to change
Background removal Exports any subject on a transparent background

Background removal is the most practical of the three. Lifting a clean subject out of a generated scene is normally a separate job with a separate tool, and having it in the same surface removes a step from every workflow that composites.

Infographic: what Image 2.0 adds, showing the magic wand, segmentation and background removal tools, why typography is the test, and the API model ids

Through the API

The Imagine API exposes it as grok-imagine-image-quality, alongside grok-imagine-video-1.5 for video. xAI documents the surface as image generation, image editing with up to three reference images, video generation from text or stills, video editing, reference-to-video, video extension from a final frame, and a Files API for storing and referencing generated assets.

Pricing is flat per image regardless of prompt length, which is worth knowing when you are writing the long, specific prompts this kind of model rewards. There is no penalty for detail.

Where it sits

This is the successor to Grok Imagine, which was the fast and permissive option rather than the precise one. Image 2.0 is aimed at a different job, and the competition it is aimed at is the models people already use for work with text in it: Nano Banana Pro and GPT Image 2.

What has not been published

xAI has released no benchmark table with this model, so the typography and instruction-following claims stand on demonstrations rather than a measured comparison against Nano Banana Pro or GPT Image 2. The per-image price is described as flat but the figure is not stated in the capability documentation. And Quality Mode is a mode inside Grok rather than a separate product, so its behaviour can change without a version number moving.

The claims are testable, which is the useful part: put dense small type in a prompt, generate it in all three, and read the labels.