ARC Prize scored GPT-6 Astra at 62.7% and 99.9% on the same benchmark on the same day
OpenAI released GPT-6 Astra on 3 September. ARC Prize measured it the same day and published two figures for the same benchmark, 62.7% and 99.9%. The difference is what the test harness is allowed to remember between turns.
Listen to this articleListen

OpenAI released GPT-6 Astra on 3 September 2026. The same day, ARC Prize published its own measurement of the model on ARC-AGI-3, the third and hardest of its reasoning benchmarks.
It published two.
Under what the foundation calls its Standard harness, Astra scored 62.7% on the ARC-AGI-3 Semi-Private set, at a cost of about $26,000. Under a second arrangement it calls the Provider Adapter harness, the same model on the same benchmark scored 99.9%, for about $19,000. ARC Prize describes both as state of the art, and it is right on both counts. The previous best on that benchmark was a long way below either.
The second number is the one that travelled.
Memory between turns is what separates the two runs
ARC Prize defines both harnesses on its own page, and the definitions are the story.
The Standard harness “enables a model to carry forward notes it chooses to keep with it throughout the environment”. The model plays, decides what is worth writing down, and carries that note into the next turn. Everything else is discarded. Every model on the leaderboard is measured this way, which is what makes the column comparable.
The Provider Adapter harness “preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work”.
“Opaque” is the load-bearing word. What carries across is the provider’s own internal reasoning state, held and restored by OpenAI’s infrastructure. A rival can be measured the same way only if it exposes an equivalent, and the harness is named after the provider for that reason.
So the gap between 62.7% and 99.9% is the difference between a model that must write down what it wants to remember and one that is handed back everything it was thinking.
The human comparison is about moves taken
ARC Prize’s own account on X posted that Astra “surpasses human performance on 96% of ARC-AGI-3 levels”.
Its blog says something narrower. “GPT-6 Astra surpasses the human baseline in action efficiency on ARC-AGI-3. It used fewer actions than the median tested human on 96% of levels.”
Action efficiency counts the moves a player used. Completion is a separate measurement, and the 96% figure is about the first of the two. Both sentences describe that same measurement, and one of them names it.
The baseline itself is worth knowing. ARC Prize built it before launching the benchmark by testing “approximately 500 members of the general public”, and notes that participants “were not selected for puzzle-solving experience or ability”. It measures how quickly ordinary people got through the environments, which is the bar the 96% figure is set against.
Three numbers, one foundation, one day
The Standard harness figure appears three times in ARC Prize’s own output on 3 September, carrying a different number each time.
| Where | Standard harness score |
|---|---|
| ARC Prize blog | 62.7% |
| ARC Prize on X | 63% |
| François Chollet on X | 66% |
The first two agree: 62.7% rounds to 63%. The third sits three points above both. Chollet, who created the benchmark series, wrote that Astra “scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game”, three minutes after the foundation’s account posted 63%.
All three figures still stand as published, which leaves the choice of which to quote to the reader.
Where the AGI framing came from
ARC Prize addressed this directly, and against its own interest.
“When we launched ARC-AGI-3, we made it clear that saturating the benchmark would not represent ‘proof of achieving AGI.’ Therefore, while we believe Astra represents meaningful progress towards generalization, we are not claiming that it is AGI.”
That is the body whose benchmark is being cited, setting the limit on what its own numbers mean. The “AGI era” language reported around the launch traces to remarks by OpenAI president Greg Brockman at a press briefing, reported by Fortune as a hedged personal reflection. It belongs to that briefing, and the benchmark analysis stands separately from it.
The model, on its own terms
Astra’s published specification, from OpenAI’s developer documentation:
| GPT-6 Astra | |
|---|---|
| Context window | 1,050,000 tokens |
| Max output | 128,000 tokens |
| Knowledge cutoff | 30 April 2026 |
| Price | $10 in, $50 out per million |
| Cached input | $1.00 per million |
Cache writes are $12.50 per million. Input is text and images, output text.
How far the evidence goes
Astra is genuinely at the top of this benchmark. On the neutral harness it roughly doubles the previous best, and it does so while using fewer moves than a median member of the public on most levels. That is a real result and ARC Prize says so without hedging.
What the evidence supports is the precise version. The 99.9% is a score on a provider-specific harness, the 62.7% is the score on the harness every model is measured on, the 96% figure measures action efficiency, and the foundation that owns the benchmark has in writing stopped at calling Astra meaningful progress towards generalisation.
Sources
- ARC Prize: OpenAI's GPT-6 Astra on ARC-AGI-3arcprize.org
- ARC Prize leaderboardarcprize.org
- GPT-6 Astra model reference, OpenAI developer docsdevelopers.openai.com
- ARC Prize on X, 3 September 2026x.com
- François Chollet on X, 3 September 2026x.com
- Fortune: OpenAI debuts GPT-6 Astra, and Greg Brockman on the AGI erafortune.com


