xAI

Grok 4.6

xAI's current flagship, built for long-running agents

Released 12 August 20268 min readLarge Language ModelsLast updated:

Editorial collage: the SpaceXAI mark on torn paper beside an API card reading model equals grok-4.6, a spec sheet listing a 500,000 token context, a 1 February 2026 cutoff and $2 / $6 pricing, and a scoreboard slip showing Grok 4.6 on three, Fable 5 Max on five and GPT-5.6 Sol on two, over a halftone server hall

Key facts

12 Aug 2026current flagship
Released
500ktokens
Context
$2 / $6per M in / out
Price
$0.50up from $0.30
Cached in
3on xAI's own table
Best of 10
1 Feb 2026knowledge
Cutoff

xAI's fastest-moving flagship yet, and the honest read is in its own launch table: a clear jump over Grok 4.5 on all ten benchmarks, and the best score on only three of them.

What it is

Grok 4.6 is xAI’s current flagship model, released on 12 August 2026, thirty-five days after Grok 4.5. The API model name is grok-4.6, and the company’s own documentation describes it as “SpaceXAI’s frontier model built for coding, agentic tasks, and knowledge work”.

The launch announcement puts the emphasis on stamina rather than raw intelligence: Grok 4.6 “builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work”, and “stays with complex tasks across many steps, whether researching a topic, analyzing information, working across a codebase, or turning an idea into a polished application or work artifact.”

Screenshot of the Grok 4.6 model card on docs.x.ai, showing a Latest badge, the line Grok 4.6 is SpaceXAI's frontier model built for coding, agentic tasks, and knowledge work, and a Python example setting model to grok-4.6
The model card, showing the API name and the company's own one-line description. Note the masthead: xAI's own documentation now says SpaceXAI. Screenshot of docs.x.ai/developers/grok-4-6, 12 August 2026.

What does xAI’s own benchmark table actually show?

It shows Grok 4.6 taking the best score in three rows out of ten. xAI published a ten-benchmark comparison against Grok 4.5, GPT-5.6 Sol Max and Claude Fable 5 Max, with the winner of each row in bold, and its own model is not the one in bold most often.

Benchmark Grok 4.6 Grok 4.5 Leader
AA Intelligence Index 61 56 62 Fable 5 Max
GDPVal-AA v2 1753 1526 Grok 4.6
CursorBench v3.2 69.9% 66.7% 70.5% Fable 5 Max
DeepSWE v1.1 65.9% 54% 73% GPT-5.6 Sol Max
FrontierCode v1.1 61.3% 56.6% 63.6% Fable 5 Max
APEX-Agents 57.5% 47.1% 59.2% Fable 5 Max
Terminal-Bench v3.0 26% 15.7% 34.6% GPT-5.6 Sol Max
APEX-SWE 56.4% 53.6% 58.8% Fable 5 Max
AA-Briefcase 1577 1313 Grok 4.6
Harvey LAB (Vals) 15.8% 12.9% Grok 4.6

Counted up, Claude Fable 5 Max takes five rows, Grok 4.6 takes three, and GPT-5.6 Sol Max takes two. Grok’s three wins are the two economic and professional-work measures, GDPVal-AA and AA-Briefcase, plus a legal benchmark where the field is weak: Harvey LAB, where 15.8 per cent is the best score on offer and GPT-5.6 Sol manages 2.5 per cent.

The losses are not close everywhere either. On Terminal-Bench v3.0 Grok 4.6 scores 26 per cent against 34.6 for GPT-5.6 Sol, and on DeepSWE v1.1 it takes 65.9 per cent against 73. Both are agentic coding tests, which is the ground the launch is pitched on.

Screenshot of the evaluation table on the xAI announcement page, showing ten benchmarks with the best score in each row highlighted, and the get started section listing Cursor, Grok Build, the API, OpenRouter, Vercel and Cloudflare
xAI's own table, with its own highlighting. The footnote reads: "Best score per evaluation in bold. Third-party model scores are the best of self-reported or publicly available results." Screenshot of x.ai/news/grok-4-6, 12 August 2026.

Publishing a table you lose is unusual enough to be worth saying out loud. The company’s own framing is careful and accurate: Grok 4.6 “achieves frontier intelligence across several agentic coding and knowledge work benchmarks” and “matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index”. Both claims survive the table.

Infographic: Grok 4.6 takes the best score in three of ten benchmarks on xAI’s own table, against five for Claude Fable 5 Max and two for GPT-5.6 Sol Max. Against Grok 4.5 it improves on every one: the AA Intelligence Index from 56 to 61, GDPVal-AA v2 from 1526 to 1753, DeepSWE v1.1 from 54 to 65.9 per cent, and Terminal-Bench v3.0 from 15.7 to 26 per cent. Pricing is unchanged at 2 and 6 dollars per million tokens with a 500,000 token context, but cached input costs 67 per cent more.

Where it does win, decisively

Against its own predecessor. Grok 4.6 improves on Grok 4.5 in all ten rows, and several are step changes rather than increments: Terminal-Bench v3.0 goes from 15.7 to 26 per cent, DeepSWE from 54 to 65.9, GDPVal-AA from 1526 to 1753. Artificial Analysis, which runs the index independently, recorded the same jump: “Grok 4.6 gains 5 points over Grok 4.5 on the Intelligence Index just over one month after its release.”

Elon Musk marked the launch with one figure, on X, on 12 August: “Grok 4.6 reaches 1753 ELO.”

The specification

Model name grok-4.6
Context window 500,000 tokens
Knowledge cutoff 1 February 2026
Modalities Text and image input, text output
Output limit None stated
Reasoning effort Low, medium, high (default), xhigh
APIs Responses API, Chat Completions
Tools Function calling, web search, X search, code execution

xAI recommends setting a prompt_cache_key so a conversation’s requests route to the same server: without it, the docs warn, “you often pay full input price on a cache-cold server.”

How much does it cost to run?

The headline price has not moved from Grok 4.5: $2.00 per million input tokens and $6.00 per million output. What has changed is the cached rate, from $0.30 to $0.50 per million, an increase of about 67 per cent on the tokens an agent loop re-reads most.

The number that catches people is the long-context tier. Once a prompt reaches 200,000 tokens, every rate doubles, and xAI is explicit that this applies to the whole request: “Models with long context pricing bill the long context rates for all tokens in a request once its prompt reaches the model’s long context threshold.”

The cliff, in four steps. A long agent loop does not pay double on the tokens past 200,000; it pays double on all of them.

Tools are billed separately, at $5 per thousand calls each for web search, X search and code execution.

Where you can use it

Grok 4.6 shipped into Cursor and Grok Build on day one, and is also on the xAI API, OpenRouter, Vercel and Cloudflare. It is the default model of the Grok Build coding agent, on both the API and the CLI. xAI offered double the included usage inside Grok Build and Cursor for the launch week.

The announcement also mentions “a fast variant which is twice the price”. As of 12 August xAI has published no separate model ID or price line for it, so there is nothing to point an API call at yet.

How it was trained

xAI is more forthcoming than usual here. Grok 4.6 had “a longer supplemental training run than Grok 4.5”, using curated model-generated data for reasoning and technical concepts, high-quality engineering data, and what the company calls an improved optimiser and training recipe.

The step worth noticing is the next one: xAI used Grok 4.5 to regenerate the supervised fine-tuning trajectories for the new model, across reasoning efforts, agent harnesses and domains, then filtered out bad traces with model-based checks. The previous flagship became the teacher for its successor. Reinforcement learning followed on agentic tasks including kernel optimisation, web development and computer-aided design.

xAI has published no parameter count. The figure in circulation, 1.5 trillion on the same foundation as Grok 4.5, comes from a Musk post on 28 July, before release, and is his personal statement rather than anything in a model card. The same post says Grok 4.7 will be a 2.1-trillion-parameter model “released a few weeks later”.

The company changed its name and did not announce it

xAI’s own pages, including the Grok 4.6 announcement and the API documentation, now carry SpaceXAI in the masthead, while the copyright line still reads X.AI LLC. Our Grok 4.5 page recorded SpaceXAI in July as a name other coverage had started using. It is now the name on the company’s own products. We could not find a dated announcement of the rebrand itself on any xAI channel.

What is worth checking before you commit to it

The safety section of the announcement is thin by the standards of a frontier release: safeguards “improved and calibrated in line with the model’s capabilities”, the “widest-ever suite of pre-deployment testing”, and third-party testing referred to but not named or linked. No separate system card accompanied the launch, which is a change from xAI’s earlier practice and leaves the strongest claims on this page resting on the company’s own summary.

The gap this leaves is measurable elsewhere. Grok 4.6 does not yet appear on the standard SWE-bench Verified leaderboard, on ARC-AGI, or in a dated Arena placement, so the coding case rests on xAI’s own choice of benchmarks, three of which are its own or its partner’s. Artificial Analysis has confirmed the headline index score and the GDPVal figure independently, and its cost-per-task measurement of $0.84 is the most useful outside number available: it puts Grok 4.6 at roughly half the turns and a quarter of the input tokens of Claude Opus 5 on the same work.

For where this sits against the rest of the field, see our large language models hub and the wider AI section.