AI State of Play
State of Play
AI State of Play
Where things stand today, area by area: the read, the numbers, and the dated log of what changed.
Time-sensitiveVerified · 7 areas · 43 dated developments
LLMs
Four flagships in three days: Grok 4.6, Gemini 3.7 Flash, DeepSeek V4 Pro and GLM-5.3. The cheap flat prices are ending with them: DeepSeek goes peak and off-peak on 16 August, and Google's launch rate doubles in January.
- 4
- flagships, 12-14 Aug
- $0.75/$3.75
- Gemini 3.7 Flash intro
- 4.5x
- DeepSeek peak rise, 16 Aug
- 14 AugZ.ai ships GLM-5.3: the GLM-5.2 base post-trained harder. Coding improves, exploit-finding more than doubles, and the weights wait a fortnight for safety work.
- 13 AugGemini 3.7 Flash arrives as Google's workhorse for coding and agents: the fastest model Artificial Analysis tracks, at $0.75 and $3.75 per million until 31 December. The rate doubles on 1 January 2027.
- 13 AugDeepSeek V4 Pro goes general: MIT licence, 1M-token context. The flat price ends 16 August, with output up to 4.5 times the old rate at peak.
- 13 AugQwen3.8-Max's promised open weights land on Hugging Face, standard and FP8 with a 27B sibling, ten days after the 2.4-trillion-parameter flagship launched at $2 and $6.
- 13 AugOpenAI previews Ultrafast: GPT-5.6 Sol served at up to 750 output tokens a second on Cerebras hardware, API only, by the company's own figures.
- 12 AugGrok 4.6 beats Grok 4.5 on all ten of xAI's own benchmarks and wins three rows against the field. Cached input rises to $0.50 per million; $2 and $6 hold.
- 11 AugMeta opens Muse Glimmer under Apache 2.0, dropping the Llama licence rules.
- 24 JulClaude Opus 5 lands at $5 and $25 per million and leads the Artificial Analysis index, with Fable 5 level behind it and first on the Arena text board.
- DealSonnet 5, the everyday default with a 1M-token window, runs on introductory pricing of $2 and $10 per million until 31 August; the standard $3 and $15 returns from 1 September.
Image
Image Generation
GPT Image 2 still leads the arena by a record margin, and it now plans or searches the web before it draws.
- #1
- GPT Image 2, arena
- $0.034
- per image, NB2 Lite
- 4K
- native, NB Pro
- 7 AugxAI ships Grok Imagine Image 2.0: localised Magic Wand edits, segmentation, background removal, up to five reference images combined, and aspect-ratio conversion, live in Quality Mode.
- 6 AugGPT Image 2 still leads the arena by a record margin, and plans or searches the web before it draws.
- 30 JunNano Banana 2 Lite becomes the fast tier at roughly $0.034 an image; Nano Banana Pro keeps precise editing and native 4K.
- 24 JulMidjourney V8.2 ships and takes the default seat from V8.1; Personalization profiles are the headline upgrade.
- StandingFLUX.2 is the strongest open-weight family; Firefly Image 5 wins wherever licensing decides the tool.
Video
Video Generation
Seedance 2.5 will generate 30 seconds in a single pass, and Sora 2's API switches off on 24 September.
- 30s
- single pass, Seedance 2.5
- 24 Sep
- Sora 2 API off
- 4K/60
- Kling 3.0
- 16 JulByteDance opens Seedance 2.5's API with native 30-second single-pass clips, the longest one-shot generation anyone offers.
- 24 SepSora 2's API switches off.
- StandingGemini Omni Flash tops text-to-video while Seedance 2.0 keeps the image-to-video and editing crowns.
- StandingKling 3.0 is the value pick at native 4K and 60fps; Veo 3.1 the safest Western choice, with native dialogue and SynthID provenance.
Agents
Agentic Frameworks / Routing
xAI joined the agent race: Grok Bot's teammates live on their own cloud computer, in early beta, sold through Cursor's top plans. Hermes, meanwhile, can talk: the Herald release added live voice, wake words and an agent-to-agent protocol.
- 50
- Grok Bot roster cap
- $200
- Cursor Ultra, per month
- 650+
- Herald contributors
- 11 AugGrok Bot opens in early beta: xAI's persistent agents with their own shared cloud computer, a roster cap of fifty, access through SuperGrok Heavy and Cursor's Ultra and Teams Premium plans. xAI does not name the model powering it.
- 3 AugHermes Agent's Herald release lands real-time conversational voice with barge-in, on-device wake words, signed outbound webhooks and version 1.0 of an agent-to-agent protocol. Roughly 1,400 merged pull requests from 650-plus contributors sit behind it.
- 4 AugOpenRouter launches Ori, a CLI that configures Claude Code, Codex, OpenCode and Hermes to run on its gateway, roughly halving system-prompt tokens out of the box.
- 20 JulHermes Agent's Quicksilver release cuts first-turn time to first token by roughly 80%. Finishing a task still writes the agent a reusable skill, so it improves with use.
- 13 JulNous Research, which builds Hermes, is in talks at a $1.5B valuation.
- 9 JulOpenAI's ChatGPT Work joins, shipping finished documents, spreadsheets and web apps.
- JulyClaw Chain chains four OpenClaw flaws across roughly 245,000 reachable servers. All four are patched.
- StandingOpenClaw is the fastest-growing open-source project in GitHub history, Cowork leads desktop agents, and MCP is the universal tool standard.
Coding
AI Coding & Dev Tools
Cursor's agents can now work your Gmail, Drive and Calendar directly, while Claude Code runs Opus 5 as default at a reported $2.5B run-rate.
- $2.5B
- Claude Code run-rate
- 9 mths
- to get there
- 18%
- adoption, tied
- 3 AugCursor ships Google Workspace plugins: its agents get direct Gmail, Drive and Calendar access, searching files, drafting mail and scheduling on their own.
- 24 JulOpus 5 becomes Claude Code's default at the same price as the model it replaced. Fable 5 is still what you reach for when a task needs the highest capability Anthropic sells.
- StandingClaude Code is the fastest-growing developer product on record, at a reported $2.5B revenue run-rate inside nine months.
- StandingCursor and Claude Code are tied near 18% adoption each; Copilot's larger installed base has stopped growing.
- JuneWindsurf re-emerges as Devin Desktop under Cognition, and Codex runs on the GPT-5.6 Sol family.
Companies
AI Companies
Investors are floating an Anthropic IPO near $2 trillion, the EU can now fine a frontier lab 3% of global turnover, and the first humanoid-robot IPO opened for subscriptions on 10 August.
- ~$2tn
- reported Anthropic IPO talk
- 3%
- max EU AI Act fine
- $623m
- Unitree IPO raise
- 14 AugThe FT reports Anthropic investors discussing a listing at valuations near $2 trillion. The company has named no number and announced no IPO; the last confirmed mark is May's $965B round.
- 13 AugOpenAI appoints Dali Rajic, formerly of Wiz, as chief revenue officer, citing more than one billion weekly users and two million business customers, twice a year ago, on its own figures.
- 11 AugChatGPT's advertising pilot leaves the US: sponsored results reach free tiers in the UK, Mexico, Brazil, Japan and South Korea.
- 11 AugMistral takes regional EU-or-US inference general, announces a European compute coalition, and starts hosting third-party open models, beginning with Z.ai's GLM-5.2.
- 10 AugOpenAI expands Daybreak into Blue and Red tiers and introduces GPT-5.6-Cyber, a model built for authorised offensive security work: zero-day discovery and exploit-chain development. On AWS Bedrock from 11 August.
- 10 AugRiot Platforms discloses a $9.1bn, 20-year lease of 191MW at its Rockdale, Texas campus to an unnamed frontier AI lab. Widely reported as Anthropic; neither company has confirmed it.
- 7 AugOpenAI says preliminary evaluations of its unreleased Astra model cannot rule out critical cyber capability under its Preparedness Framework, the first such statement it has made about a model in development.
- 2 AugThe EU AI Act grows teeth: the Commission's AI Office can now demand documentation from general-purpose model providers, run technical evaluations and order a model withdrawn, with fines up to 3% of global turnover or 15m euros. Deepfake-labelling and AI-disclosure duties begin the same day.
- 31 JulThe Shanghai Stock Exchange confirms Unitree's STAR Market IPO: a raise of about $623m, 85% of it earmarked for R&D, subscriptions from 10 August. It would be the first humanoid-robotics A-share.
Resources
Learning & Resources
Published benchmark scores saturate within months of release, so the live arenas and usage boards are what practitioners now read.
- StandingBenchmarks saturate within months of release, so a published score dates fast. The arenas and usage boards track what people actually reach for.
- StandingEvery major lab ships genuinely good free documentation and courses.
- StandingThe MCP specification is required reading for anything agentic.