YFarmX logoYFarmX

xAI

Grok 4.7

xAI's current flagship, and the cheapest model on its own comparison table

Released 21 September 20267 min readLarge Language Models

Editorial collage headed GROK 4.7 with the subtitle 500K context, $2 / $6, wins 2 of 7: the angular xAI X mark on a torn card, a code card reading model equals grok-4.7, a spec card listing context 500,000, cutoff May 2026, in / out $2 / $6, reasoning low to xhigh, and a price card setting $2 against $10

Key facts

21 Sep 2026current flagship
Released
500ktokens
Context
$2 / $6per M in / out
Price
$0.50per M tokens
Cached in
2on xAI's own table
Best of 7
May 2026knowledge
Cutoff

Grok 4.7 is the model xAI now sells as its best for coding and knowledge work. It costs $2 per million input tokens against Fable 5.1's $10, it beats the Grok 4.6 it replaces on every benchmark xAI published, and on two of those seven rows it takes the top score outright.

What it is

Grok 4.7 is a large language model from xAI, released on 21 September 2026, forty days after Grok 4.6. You reach it by setting the API model name to grok-4.7, and the company’s own documentation describes it as “SpaceXAI’s frontier model built for coding, agentic tasks, and knowledge work”.

The launch post is blunter about the pitch: Grok 4.7 “is our most capable model for coding and knowledge work”, it “works longer on difficult tasks, checks its own work more carefully, and comes with our best-calibrated safeguards to date”, and it is “served at the same price and speed as Grok 4.6”.

That last clause is the whole strategy. xAI has not raised the price in three releases, so the interesting question about Grok 4.7 is what $2 per million input tokens now buys you against models that charge five times as much.

The Grok 4.7 page on docs.x.ai, showing a Latest badge beside Grok 4.7 in the sidebar, a Meet grok-4.7 button, the line Grok 4.7 is SpaceXAI's frontier model built for coding, agentic tasks, and knowledge work, and a Python example setting model to grok-4.7
The model card, carrying the API name and xAI's own one-line description. The description under the buttons opens "Grok 4.7 is SpaceXAI's frontier model", which is the company's name now. Screenshot of docs.x.ai/developers/grok-4-7, 21 September 2026.

What does xAI’s own benchmark table show?

xAI published seven benchmarks on launch day, and Grok 4.7 takes the best score on two of them. It beats the Grok 4.6 it replaces on all seven.

Benchmark Grok 4.7 Grok 4.6 Leader
CursorBench 4.0 46.3% 40.4% 51.8% Fable 5.1 Max
DeepSWE v1.1 71.0% 65.2% 72.7% GPT-5.6 Sol Max
EEBench 64.0% 53.0% Grok 4.7
AA Briefcase v1.1 1,657 1,546 1,678 Fable 5.1 Max
Terminal-Bench 4.0 38.0% 20.3% 57.9% Fable 5.1 Max
Harvey Legal Agent 19.6% 15.8% Grok 4.7
HealthBench Professional 56.7% 48.5% 62.1% Fable 5.1 Max

Read the column headings before the numbers. xAI is comparing Grok 4.7 at xHigh reasoning against Grok 4.6 at High, so part of the generational gain is a longer thinking budget rather than a better model. The DeepSWE figure carries an asterisk on xAI’s own page marking it a high-effort score, which makes it the one row not run at the setting the column advertises.

Terminal-Bench is the row that moved most: 20.3% to 38.0%, very nearly double, on a benchmark built from long multi-step shell work. That is consistent with the training change xAI describes, and it is still twenty points behind the leader.

xAI's GDPval bar chart of Elo scores on professional knowledge work: Fable 5.1 at max effort 1735, Grok 4.7 at xhigh 1695, Grok 4.6 at high 1605 and GPT-6 Astra at max 1542
GDPval measures work of the kind lawyers, nurses and financial analysts do. Grok 4.7 clears the Grok 4.6 it replaces by 90 Elo and sits 40 behind Fable 5.1. Chart from xAI's launch post, 21 September 2026.

It costs $2 where the table’s winner costs $10

The argument for Grok 4.7 is the cost column, and xAI puts it in the table itself. Grok 4.7 charges $2 and $6 per million input and output tokens. Fable 5.1 Max, which takes four of the seven rows, charges $10 and $50. GPT-5.6 Sol Max charges $4 and $20.

Set those side by side and Grok 4.7 costs a fifth of Fable 5.1 Max on input tokens and a little over an eighth on output. The two rows Grok 4.7 takes, it takes outright at those prices.

The Harvey Legal Agent Benchmark is the widest gap on the page, and it runs the other way from the rest: Grok 4.7 scores 19.6% where GPT-5.6 Sol Max manages 2.5% and Fable 5.1 Max manages 6.7%. A benchmark where the two most expensive models score under 7% is worth treating with care, because a gap that wide usually says more about the task format than about legal ability.

xAI's CursorBench 4.0 chart plotting score against average cost per task for Fable 5.1, Opus 5, Grok 4.7, Sonnet 5 and GPT-5.6 Sol, with cost falling from 18 dollars on the left to zero on the right
The chart xAI leads with: CursorBench 4.0 score against average cost per task, cost falling left to right. Fable 5.1 holds the top of the score axis and Grok 4.7 holds a lower curve at a fraction of the spend. From xAI's launch post, 21 September 2026.

The fast variant lives in Cursor and Grok Build

Grok 4.7 Fast is a separate product, reachable inside Cursor and inside xAI’s own coding agent, and it bills at double the standard rate. The launch page’s subtitle promises a model “twice as fast, at half the price of comparable models”, and the speed half of that sentence belongs to this variant.

xAI’s documentation sets out the restrictions: “Grok 4.7 Fast is the same model served on faster infrastructure, billed at twice the standard token rates. It is available only in Cursor and Grok Build, and it is not included in Grok Build’s free tier. It is not available on the public xAI API.

So a team building on the API gets the $2 and $6 model at standard speed. Double speed costs $4 and $12, and it is reachable only inside xAI’s own coding agent or inside Cursor.

The specification

Model name grok-4.7
Context window 500,000 tokens
Knowledge cutoff May 2026
Modalities Text and image input, text output
Output limit No text output limit
Reasoning effort Low, medium, high (default), xhigh
APIs Responses API, Chat Completions
Tools Function calling, web search, X search, code execution

Two behaviours are worth knowing before you build on it. Encrypted reasoning is always returned on the Responses API, so a grok-4.7 response carries reasoning.encrypted_content whether or not you asked for it, which is how a multi-turn conversation keeps the model’s reasoning. And the documentation says “We highly recommend setting a prompt_cache_key”, because without one a conversation’s requests scatter across servers and you pay full input price on a cache-cold machine.

What does it cost to run?

Pricing is tiered on prompt length, and the tier changes at 200,000 tokens.

Prompt size Input Cached input Output
Under 200k tokens $2.00 $0.50 $6.00
200k tokens and over $4.00 $1.00 $12.00

All figures are per million tokens. The threshold works in a way that catches people out, and xAI’s pricing page spells it out: a model on long-context pricing “bill[s] the long context rates for all tokens in a request once its prompt reaches the model’s long context threshold”. A 201,000-token prompt therefore costs $4 per million across the whole thing, not $2 on the first 200,000 and $4 on the remainder.

The cached input rate is the one that decides the bill on a long agent run, which is why the prompt_cache_key advice above is a cost instruction rather than a performance tip. Routing inference through the US regional endpoint at us.api.x.ai carries a 10% premium on token usage.

Where you can use it

Grok 4.7 is the default model in Grok Build, xAI’s coding agent, and it is available in Cursor on all plans. On the API you reach it from the xAI console, or through the OpenRouter, Vercel and Cloudflare gateways. The US regional endpoint keeps inference inside the United States for the 10% premium.

How it was trained

xAI describes three changes from Grok 4.6. Grok 4.7 “uses a new, larger base model”. It “was trained with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete”. And it was trained “to natively understand the Grok Bot harness”, which xAI says makes it better at conversational tasks and general knowledge work.

The company reports two safety figures alongside that. Grok 4.7 tops LatchBio’s biosafety benchmark at 62.4%, and on HackerBench v0.3, xAI’s own benchmark for risky and malicious cyber tasks, it allows 3.3% of risky dual-use prompts through. Both come from xAI rather than an outside evaluator, so they describe what the lab measured on its own tests.

What to check before you commit to it

The knowledge cutoff is May 2026, four months before release, so anything later reaches the model through web search or X search rather than its weights.

The benchmark table compares reasoning settings that differ by column, and the effort setting changes both the score and the bill. A number produced at xHigh is not the number you get at the default High.

The seven benchmarks are the ones xAI chose to publish. Four of the seven go to Fable 5.1 Max, one to GPT-5.6 Sol Max, and xAI printed all of it, which is a point in the table’s favour rather than against it.

And the headline speed claim belongs to Grok 4.7 Fast, a separate product that costs twice as much and lives inside Cursor and Grok Build.