AI State of Play

State of Play

AI State of Play

The house verdict on where this frontier stands: a considered read, hard numbers, every shift dated.

Time-sensitiveVerified

LLMs

GPT-6 Astra began rolling out on 3 September at $10/$50 with a 1,050,000-token context, two days after Claude Fable 5.1 shipped and took the top of the Artificial Analysis index at 57; Gemini 3.8 Flash and Muse Spark 1.3 both landed on 2 September.

$10/$50
GPT-6 Astra, per million
57
Fable 5.1 leads the AA index
$0.75/$3.75
Gemini 3.8 Flash, to 31 Dec
  1. 3 SepGPT-6 Astra begins rolling out: a new OpenAI generation at $10 in and $50 out per million, $1 cached input, a 1,050,000-token context, and the first Critical cybersecurity classification under OpenAI's Preparedness Framework (developers.openai.com; deploymentsafety.openai.com).
  2. 2 SepGoogle ships Gemini 3.8 Flash at 3.7 Flash's introductory $0.75 and $3.75 per million to 31 December, plus Gemini 3.8 Flash Cyber, gated to vetted defenders through the new Fairwind Program (blog.google).
  3. 2 SepMeta publishes Muse Spark 1.3: about 20% fewer tool calls and 25% fewer tokens than 1.2 on coding tasks (research.meta.ai).
  4. 1 SepClaude Fable 5.1 and Mythos 5.1 ship at Fable 5's $10/$50 with cache reads cut 75% to $0.25 per million; Anthropic estimates typical workloads about 25% cheaper than Fable 5, and Terminal-Bench-Science rises from 24.7% to 52.6% (anthropic.com).
  5. 28 AugZ.ai publishes the full GLM-5.3 weights: 753 billion parameters, 756GB across 141 shards, under a custom glm-5.3 licence rather than the MIT of the 5.2 line. The model card's CyberGym score reads 84.5 against GLM-5.2's 77.2. Read our report →
  6. 27 AugTencent publishes Hy4 preview under Apache 2.0, weights out on the day: 770 billion parameters with 49 billion active per token on Tencent's own figures, 78 layers, 256 routed experts plus one shared, and a one-million-token context. Read our report →
  7. 26 AugGLM-5.3-Flash launches under MIT: 320B parameters with 18B active, multimodal, 1M context. Z.ai confirms it is the model it ran anonymously as ox-alpha from 20 August, the stealth model YFarmX fingerprinted to Z.ai's tokenizer. Read our report →
  8. 26 AugAlibaba opens the weights of Qwen3.8-Flash-Next, 125 billion parameters with about 6 billion active, on Qwen Community 1.0 terms at $0.15 in and $0.47 out. Qwen calls it an experimental preview of the architecture Qwen4 will be built on. Read our report →
  9. 21 AugOpenAI cuts GPT-5.6 Sol to $4 in and $20 out per million on a promotion running to at least 21 November, with long context at $8 and $30. Terra and Luna hold. Read our report →
  10. 21 AugA Hugging Face search for Qwen3.8-27B-abliterated returns 150 refusal-stripped rebuilds of Alibaba's Apache-2.0 model, carrying more than 1.2 million downloads between them a week after release. Read our report →
  11. 16 AugDeepSeek's peak and off-peak pricing takes effect as announced: V4 Pro output runs $3.96 per million at peak against $1.98 off-peak, from a flat $0.87, with peak defined as 01:00-04:00 and 06:00-10:00 UTC on weekdays only. Read our report →
  12. 14 AugZ.ai ships GLM-5.3: the GLM-5.2 base post-trained harder. Coding improves, exploit-finding more than doubles, and the weights wait a fortnight for safety work. Read our report →
  13. 13 AugGemini 3.7 Flash arrives as Google's workhorse for coding and agents: the fastest model Artificial Analysis tracks, at $0.75 and $3.75 per million until 31 December. The rate doubles on 1 January 2027. Read our report →
  14. 13 AugDeepSeek V4 Pro goes general: MIT licence, 1M-token context. The flat price ends 16 August, with output up to 4.5 times the old rate at peak. Read our report →
  15. 13 AugQwen3.8-Max's promised open weights land on Hugging Face, standard and FP8 with a 27B sibling, ten days after the 2.4-trillion-parameter flagship launched at $2 and $6. Read our report →
  16. 12 AugGrok 4.6 beats Grok 4.5 on all ten of xAI's own benchmarks and wins three rows against the field. Cached input rises to $0.50 per million; $2 and $6 hold. Read our report →
  17. 11 AugMeta opens Muse Glimmer under Apache 2.0, dropping the Llama licence rules. Read our report →
Explore llms

Image

Image Generation

GPT Image 2 still leads the arena and the lead is back out to 51 points, with MAI-Image-2.6 second and Nano Banana 2 seventh. Meta's Muse Image reached OpenRouter at a cent an image and sits fifth on text-to-image.

#1
GPT Image 2, arena
51pts
lead over MAI-Image-2.6
$0.01
Muse Image, OpenRouter
  1. 30 AugGPT Image 2 still leads the arena at 1381.7 Elo, 51 points clear of Microsoft's MAI-Image-2.6, with Grok Imagine Image 2.0 third, Reve 2.1 fourth, Meta's Muse Image fifth and Nano Banana 2 seventh (arena.ai, read this day).
  2. 26 AugMeta's Muse Image lands on OpenRouter at $0.01 an image with a 66K context: the reasoning image model that breaks a brief down, refines its own output and invokes web search on knowledge-heavy prompts. Read our report →
  3. 10 AugMicrosoft ships MAI-Image-2.6 and it takes second place on the arena text-to-image board at launch, ahead of Google, Meta and xAI, with text rendering up 91 Elo on the company's own figures.
  4. 7 AugxAI ships Grok Imagine Image 2.0: localised Magic Wand edits, segmentation, background removal, up to five reference images combined, and aspect-ratio conversion, live in Quality Mode. Read our report →
  5. 30 JunNano Banana 2 Lite becomes the fast tier at roughly $0.034 an image; Nano Banana Pro keeps precise editing and native 4K.
  6. 24 JulMidjourney V8.2 ships and takes the default seat from V8.1; Personalization profiles are the headline upgrade.
  7. StandingFLUX.2 is the strongest open-weight family; Firefly Image 5 wins wherever licensing decides the tool.
Explore image

Video

Video Generation

Gemini Omni 1.1 Flash arrived on 27 August and leads text-to-video, Wan 3.0 took the editing board from Seedance 2.5 by four points, and Sora 2's API switches off on 24 September.

27 Aug
Omni 1.1 Flash ships
24 Sep
Sora 2 API off
4pts
Wan 3.0 over Seedance
  1. 24 SepSora 2's API switches off.
  2. 29 AugThe boards, checked this day at arena.ai: Gemini Omni 1.1 Flash leads text-to-video with FLUX 3 Video third; MiniMax-H3 holds image-to-video with Omni 1.1 Flash second and Seedance 2.5 third; Wan 3.0 tops video editing four points clear of Seedance 2.5.
  3. 27 AugGoogle ships Gemini Omni 1.1 Flash: end frames, scene extension in 10-second steps to a cumulative 40 seconds, a 360p draft mode about 60% faster at a third of the cost, and 1080p or 4K export (Google's own launch post).
  4. 24 AugAlibaba's Wan 3.0 goes wide: 30-second single-pass clips, documents, spreadsheets, slides and web pages as input, and an API from about $0.05 a second at 480p. Within two days it tops the video-editing arena. No open weights yet.
  5. 16 JulByteDance opens Seedance 2.5's API with native 30-second single-pass clips, generating picture and sound in one model. Read our report →
  6. StandingKling 3.0 is the value pick at native 4K and 60fps; Veo 3.1 the safest Western choice, with native dialogue and SynthID provenance.
Explore video

Agents

Agentic Frameworks / Routing

Hermes Agent now browses as you: version 0.20.6's real-profile browsing drives a managed copy of your Chrome profile, off by default and flagged in its own docs as a convenience, not an isolation boundary.

0.20.6
Hermes, 27 Aug
50
Grok Bot roster cap
$200
Cursor Ultra, per month
  1. 27 AugHermes Agent 0.20.6 ships consent-gated real-profile browsing: the agent works from a managed copy of your Chrome profile, cookies and saved logins included, off by default. The docs call it a convenience, not an isolation boundary. Fifty-plus vendor-hosted remote MCP servers and keychain secret encryption land with it. Read our report →
  2. 11 AugGrok Bot opens in early beta: xAI's persistent agents with their own shared cloud computer, a roster cap of fifty, access through SuperGrok Heavy and Cursor's Ultra and Teams Premium plans. xAI does not name the model powering it. Read our report →
  3. 3 AugHermes Agent's Herald release lands real-time conversational voice with barge-in, on-device wake words, signed outbound webhooks and version 1.0 of an agent-to-agent protocol. Roughly 1,400 merged pull requests from 650-plus contributors sit behind it. Read our report →
  4. 4 AugOpenRouter launches Ori, a CLI that configures Claude Code, Codex, OpenCode and Hermes to run on its gateway, roughly halving system-prompt tokens out of the box.
  5. 20 JulHermes Agent's Quicksilver release cuts first-turn time to first token by roughly 80%. Finishing a task still writes the agent a reusable skill, so it improves with use.
  6. 13 JulNous Research, which builds Hermes, is in talks at a $1.5B valuation.
  7. 9 JulOpenAI's ChatGPT Work joins, shipping finished documents, spreadsheets and web apps. Read our report →
  8. JulyClaw Chain chains four OpenClaw flaws across roughly 245,000 reachable servers. All four are patched.
  9. StandingOpenClaw is the fastest-growing open-source project in GitHub history, Cowork leads desktop agents, and MCP is the universal tool standard.
Explore agents

Coding

AI Coding & Dev Tools

OpenAI gave Cursor three months' notice: model access ends 12 November unless the talks resolve it, a fortnight after SpaceX's $60bn purchase of Anysphere closed.

12 Nov
proposed OpenAI shutoff
5%
Cursor traffic on OpenAI
$2.5B
Claude Code run-rate
  1. 28 AugOpenAI notifies SpaceX it intends to wind down Cursor's access to OpenAI models, proposing a 12 November shutoff and citing its experience of Musk companies breaking contracts. Michael Truell says OpenAI models carry about 5% of Cursor traffic and the two sides are talking (OpenAI's own post; Truell on X).
  2. 14 AugSpaceX's $60bn all-stock acquisition of Anysphere closes: 'Cursor is now a part of SpaceX', the company's own blog says.
  3. 3 AugCursor ships Google Workspace plugins: its agents get direct Gmail, Drive and Calendar access, searching files, drafting mail and scheduling on their own.
  4. 24 JulOpus 5 becomes Claude Code's default at the same price as the model it replaced. Fable 5 is still what you reach for when a task needs the highest capability Anthropic sells. Read our report →
  5. StandingClaude Code is the fastest-growing developer product on record, at a reported $2.5B revenue run-rate inside nine months.
  6. StandingCursor and Claude Code are tied near 18% adoption each; Copilot's larger installed base has stopped growing.
  7. JuneWindsurf re-emerges as Devin Desktop under Cognition, and Codex runs on the GPT-5.6 Sol family.
Explore coding

Companies

AI Companies

Nvidia's Hugging Face acquisition is signed: an 8-K records the definitive agreement of 2 September at $12.93bn all told, with completion expected in the first half of 2027 subject to regulatory approvals; Anthropic's S-1 had still not landed by 5 September.

$12.93bn
Nvidia-Hugging Face, signed
12 Nov
proposed Cursor shutoff
3.5x
Unitree vs issue, 4 Sep
  1. 4 SepUnitree closes at 530.30 yuan, down 3.66% on the day and about 3.5 times the 150.80 issue price, from near six times on debut day (exchange data via stockanalysis.com).
  2. 3 SepNvidia files an 8-K recording a definitive agreement, dated 2 September, to acquire Hugging Face: about $11.9bn to stockholders plus up to $1.0bn in retention awards, $12,930,300,000 all told in Jensen Huang's newsroom post, completion expected in the first half of 2027 subject to regulatory approvals (SEC EDGAR; blogs.nvidia.com). Read our report →
  3. 2 SepHiddenLayer closes a $100m Series B led by Delta-v Capital, with M12 and Booz Allen Ventures aboard and revenue up more than tenfold in a year; Wonderful raises a $550m Series C the same week, and Bloomberg reports a $13bn Crusoe cloud deal (company newsrooms; Bloomberg). Read our report →
  4. 1 SepThe Department of War launches OpenAI's ChatGPT Mil and Starshield AI's Grok for Government on GENAI.mil at Impact Level 5 for controlled unclassified information, extending access towards roughly three million personnel (war.gov releases). Read our report →
  5. 31 AugThe European Commission designates ChatGPT a very large online search engine under the Digital Services Act, alongside Reddit and Roblox as very large online platforms, with obligations applying from January 2027 (digital-strategy.ec.europa.eu). Read our report →
  6. 31 AugAdobe, Saudi Arabia's MCIT and HUMAIN expand their partnership: an Adobe commitment valued over $4bn, twelve months of free Firefly Standard and Express Premium for the Kingdom's 27m+ residents, and a jointly built Firefly Foundry model (news.adobe.com). Read our report →
  7. 28 AugOpenAI notifies SpaceX it intends to wind down Cursor's access to OpenAI models, proposing a 12 November shutoff, writing that it cannot be confident SpaceX will stay within its terms. Michael Truell: about 5% of Cursor traffic, and talks continue (OpenAI's own post; Truell on X).
  8. 28 AugUnitree closes at 585.00 yuan, down 4.9% on the day and about 3.9 times the 150.80 issue price, from near six times on debut day (exchange data via stockanalysis.com).
  9. 27 AugCNBC and Fortune report Nvidia has agreed to buy Hugging Face for $12.9bn, both citing people familiar; Fortune says the talks had not produced a signed agreement and could still fall apart. Nvidia's newsroom, Hugging Face's blog and EDGAR carry nothing, and neither company will comment.
  10. 26 AugNvidia reports $96.2bn of revenue for the quarter ended 26 July, up 18% on the quarter and 106% in a year; AWS announces two million more NVIDIA GPUs the same day (both companies' own releases).
  11. 26 AugOpenAI publishes 'The Hugging Face incident and the road ahead', its follow-up report on the July compromise of Hugging Face's infrastructure by its own evaluation agents. Read our report →
  12. 26 AugSam Altman tells TIME that OpenAI expects to have, by the end of the year, an internal system he would call AGI; research chief Mark Chen puts the company '80% of the way' there.
  13. 26 AugQwenWork's international edition opens in public beta: Alibaba's all-in-one agent workplace turns briefs into documents, presentations and pages, consolidating three earlier agent products.
  14. 25 AugReuters, citing the Wall Street Journal, reports Anthropic's IPO pitch describes a market above $30tn and projects 2028 revenue of $190bn to $200bn. No public S-1 had reached EDGAR by 29 August. Read our report →
  15. 24 AugNVIDIA says SpaceXAI will run Grok's next agentic workloads on Vera CPUs and the Vera Rubin platform, and plans a first-generation Starmind AI satellite on an optimised Vera Rubin NVL72 rack (NVIDIA's own release).
  16. 21 AugClaude Security starts scanning on Claude Mythos 5, in public beta for Claude Enterprise: findings come back with a CWE category, severity, confidence and a suggested patch, billed as ordinary token usage, with no direct access to the model itself. Read our report →
  17. 20 AugBloomberg reports Broadcom seeking more than $60bn of senior-secured debt, toward $100bn with a junior tranche, to finance AI chips leased to customers including Anthropic through a special-purpose vehicle.
  18. 20 AugBloomberg reports Anthropic expects its IPO to match or beat SpaceX's record raise, $75bn at the outset and $86.2bn with the over-allotment, with a public filing possible by the end of August; Fortune reports CFO Krishna Rao declining to name a valuation. The company has confirmed nothing. Read our report →
  19. 19 AugUnitree lists in Shanghai and finishes the morning session at 883.87 yuan against a 150.80 yuan issue price, valuing the company near $50bn on a roughly $904m raise (Reuters). The first humanoid-robotics A-share.
  20. 19 AugStripe and OpenRouter confirm it: Stripe agrees to acquire the router, both announcements silent on price and the close expected in the coming weeks. OpenRouter keeps its name, product and roadmap, and says it routes more than 10 trillion tokens a day across 400-plus models. Read our report →
  21. 18 AugOpenAI pauses reinforcement-learning training on models intended for deployment for two weeks, its largest planned frontier run held, after preliminary evidence that Astra may meet the Critical cyber threshold; roughly 20% of inference compute now runs under monitoring.
  22. 16 AugFoundation's Sankaet Pathak tells Fox News a humanoid border pilot could run 'tomorrow'. Homeland Security says it has no non-contractual agreements, pilots or active prime contracts with the company, whose Phantom spec sheet lists 1.8m, 80kg and a 40kg payload. Read our report →
  23. 15 AugUnitree prices its Shanghai float at 150.80 yuan, raising about 6.1 billion yuan; the retail tranche was oversubscribed more than 8,000 times, a STAR Market record. Read our report →
  24. 14 AugThe FT reports Anthropic investors discussing a listing at valuations near $2 trillion. The company has named no number and announced no IPO; the last confirmed mark is May's $965B round. Read our report →
  25. 11 AugMistral takes regional EU-or-US inference general, announces a European compute coalition, and starts hosting third-party open models, beginning with Z.ai's GLM-5.2.
  26. 10 AugOpenAI expands Daybreak into Blue and Red tiers and introduces GPT-5.6-Cyber, a model built for authorised offensive security work: zero-day discovery and exploit-chain development. On AWS Bedrock from 11 August.
  27. 10 AugRiot Platforms discloses a $9.1bn, 20-year lease of 191MW at its Rockdale, Texas campus to an unnamed frontier AI lab. Widely reported as Anthropic; neither company has confirmed it.
  28. 7 AugOpenAI says preliminary evaluations of its unreleased Astra model cannot rule out critical cyber capability under its Preparedness Framework, the first such statement it has made about a model in development.
  29. 6 AugDemis Hassabis becomes chair of Google DeepMind and Alphabet's chief scientist, Koray Kavukcuoglu takes day-to-day leadership, and Jeff Dean leaves to found a startup (Sundar Pichai's staff memo).
  30. 2 AugThe EU AI Act grows teeth: the Commission's AI Office can now demand documentation from general-purpose model providers, run technical evaluations and order a model withdrawn, with fines up to 3% of global turnover or 15m euros. Deepfake-labelling and AI-disclosure duties begin the same day.
Explore companies

Resources

Learning & Resources

Published benchmark scores saturate within months of release, so the live arenas and usage boards are what practitioners now read.

  1. StandingBenchmarks saturate within months of release, so a published score dates fast. The arenas and usage boards track what people actually reach for.
  2. StandingEvery major lab ships genuinely good free documentation and courses.
  3. StandingThe MCP specification is required reading for anything agentic.
Explore resources