AI State of Play

State of Play

AI State of Play

Where things stand today, area by area: the read, the numbers, and the dated log of what changed.

Time-sensitiveVerified · 7 areas · 43 dated developments

LLMs

Four flagships in three days: Grok 4.6, Gemini 3.7 Flash, DeepSeek V4 Pro and GLM-5.3. The cheap flat prices are ending with them: DeepSeek goes peak and off-peak on 16 August, and Google's launch rate doubles in January.

4
flagships, 12-14 Aug
$0.75/$3.75
Gemini 3.7 Flash intro
4.5x
DeepSeek peak rise, 16 Aug
  1. 14 AugZ.ai ships GLM-5.3: the GLM-5.2 base post-trained harder. Coding improves, exploit-finding more than doubles, and the weights wait a fortnight for safety work.
  2. 13 AugGemini 3.7 Flash arrives as Google's workhorse for coding and agents: the fastest model Artificial Analysis tracks, at $0.75 and $3.75 per million until 31 December. The rate doubles on 1 January 2027.
  3. 13 AugDeepSeek V4 Pro goes general: MIT licence, 1M-token context. The flat price ends 16 August, with output up to 4.5 times the old rate at peak.
  4. 13 AugQwen3.8-Max's promised open weights land on Hugging Face, standard and FP8 with a 27B sibling, ten days after the 2.4-trillion-parameter flagship launched at $2 and $6.
  5. 13 AugOpenAI previews Ultrafast: GPT-5.6 Sol served at up to 750 output tokens a second on Cerebras hardware, API only, by the company's own figures.
  6. 12 AugGrok 4.6 beats Grok 4.5 on all ten of xAI's own benchmarks and wins three rows against the field. Cached input rises to $0.50 per million; $2 and $6 hold.
  7. 11 AugMeta opens Muse Glimmer under Apache 2.0, dropping the Llama licence rules.
  8. 24 JulClaude Opus 5 lands at $5 and $25 per million and leads the Artificial Analysis index, with Fable 5 level behind it and first on the Arena text board.
  9. DealSonnet 5, the everyday default with a 1M-token window, runs on introductory pricing of $2 and $10 per million until 31 August; the standard $3 and $15 returns from 1 September.
Explore llms

Image

Image Generation

GPT Image 2 still leads the arena by a record margin, and it now plans or searches the web before it draws.

#1
GPT Image 2, arena
$0.034
per image, NB2 Lite
4K
native, NB Pro
  1. 7 AugxAI ships Grok Imagine Image 2.0: localised Magic Wand edits, segmentation, background removal, up to five reference images combined, and aspect-ratio conversion, live in Quality Mode.
  2. 6 AugGPT Image 2 still leads the arena by a record margin, and plans or searches the web before it draws.
  3. 30 JunNano Banana 2 Lite becomes the fast tier at roughly $0.034 an image; Nano Banana Pro keeps precise editing and native 4K.
  4. 24 JulMidjourney V8.2 ships and takes the default seat from V8.1; Personalization profiles are the headline upgrade.
  5. StandingFLUX.2 is the strongest open-weight family; Firefly Image 5 wins wherever licensing decides the tool.
Explore image

Video

Video Generation

Seedance 2.5 will generate 30 seconds in a single pass, and Sora 2's API switches off on 24 September.

30s
single pass, Seedance 2.5
24 Sep
Sora 2 API off
4K/60
Kling 3.0
  1. 16 JulByteDance opens Seedance 2.5's API with native 30-second single-pass clips, the longest one-shot generation anyone offers.
  2. 24 SepSora 2's API switches off.
  3. StandingGemini Omni Flash tops text-to-video while Seedance 2.0 keeps the image-to-video and editing crowns.
  4. StandingKling 3.0 is the value pick at native 4K and 60fps; Veo 3.1 the safest Western choice, with native dialogue and SynthID provenance.
Explore video

Agents

Agentic Frameworks / Routing

xAI joined the agent race: Grok Bot's teammates live on their own cloud computer, in early beta, sold through Cursor's top plans. Hermes, meanwhile, can talk: the Herald release added live voice, wake words and an agent-to-agent protocol.

50
Grok Bot roster cap
$200
Cursor Ultra, per month
650+
Herald contributors
  1. 11 AugGrok Bot opens in early beta: xAI's persistent agents with their own shared cloud computer, a roster cap of fifty, access through SuperGrok Heavy and Cursor's Ultra and Teams Premium plans. xAI does not name the model powering it.
  2. 3 AugHermes Agent's Herald release lands real-time conversational voice with barge-in, on-device wake words, signed outbound webhooks and version 1.0 of an agent-to-agent protocol. Roughly 1,400 merged pull requests from 650-plus contributors sit behind it.
  3. 4 AugOpenRouter launches Ori, a CLI that configures Claude Code, Codex, OpenCode and Hermes to run on its gateway, roughly halving system-prompt tokens out of the box.
  4. 20 JulHermes Agent's Quicksilver release cuts first-turn time to first token by roughly 80%. Finishing a task still writes the agent a reusable skill, so it improves with use.
  5. 13 JulNous Research, which builds Hermes, is in talks at a $1.5B valuation.
  6. 9 JulOpenAI's ChatGPT Work joins, shipping finished documents, spreadsheets and web apps.
  7. JulyClaw Chain chains four OpenClaw flaws across roughly 245,000 reachable servers. All four are patched.
  8. StandingOpenClaw is the fastest-growing open-source project in GitHub history, Cowork leads desktop agents, and MCP is the universal tool standard.
Explore agents

Coding

AI Coding & Dev Tools

Cursor's agents can now work your Gmail, Drive and Calendar directly, while Claude Code runs Opus 5 as default at a reported $2.5B run-rate.

$2.5B
Claude Code run-rate
9 mths
to get there
18%
adoption, tied
  1. 3 AugCursor ships Google Workspace plugins: its agents get direct Gmail, Drive and Calendar access, searching files, drafting mail and scheduling on their own.
  2. 24 JulOpus 5 becomes Claude Code's default at the same price as the model it replaced. Fable 5 is still what you reach for when a task needs the highest capability Anthropic sells.
  3. StandingClaude Code is the fastest-growing developer product on record, at a reported $2.5B revenue run-rate inside nine months.
  4. StandingCursor and Claude Code are tied near 18% adoption each; Copilot's larger installed base has stopped growing.
  5. JuneWindsurf re-emerges as Devin Desktop under Cognition, and Codex runs on the GPT-5.6 Sol family.
Explore coding

Companies

AI Companies

Investors are floating an Anthropic IPO near $2 trillion, the EU can now fine a frontier lab 3% of global turnover, and the first humanoid-robot IPO opened for subscriptions on 10 August.

~$2tn
reported Anthropic IPO talk
3%
max EU AI Act fine
$623m
Unitree IPO raise
  1. 14 AugThe FT reports Anthropic investors discussing a listing at valuations near $2 trillion. The company has named no number and announced no IPO; the last confirmed mark is May's $965B round.
  2. 13 AugOpenAI appoints Dali Rajic, formerly of Wiz, as chief revenue officer, citing more than one billion weekly users and two million business customers, twice a year ago, on its own figures.
  3. 11 AugChatGPT's advertising pilot leaves the US: sponsored results reach free tiers in the UK, Mexico, Brazil, Japan and South Korea.
  4. 11 AugMistral takes regional EU-or-US inference general, announces a European compute coalition, and starts hosting third-party open models, beginning with Z.ai's GLM-5.2.
  5. 10 AugOpenAI expands Daybreak into Blue and Red tiers and introduces GPT-5.6-Cyber, a model built for authorised offensive security work: zero-day discovery and exploit-chain development. On AWS Bedrock from 11 August.
  6. 10 AugRiot Platforms discloses a $9.1bn, 20-year lease of 191MW at its Rockdale, Texas campus to an unnamed frontier AI lab. Widely reported as Anthropic; neither company has confirmed it.
  7. 7 AugOpenAI says preliminary evaluations of its unreleased Astra model cannot rule out critical cyber capability under its Preparedness Framework, the first such statement it has made about a model in development.
  8. 2 AugThe EU AI Act grows teeth: the Commission's AI Office can now demand documentation from general-purpose model providers, run technical evaluations and order a model withdrawn, with fines up to 3% of global turnover or 15m euros. Deepfake-labelling and AI-disclosure duties begin the same day.
  9. 31 JulThe Shanghai Stock Exchange confirms Unitree's STAR Market IPO: a raise of about $623m, 85% of it earmarked for R&D, subscriptions from 10 August. It would be the first humanoid-robotics A-share.
Explore companies

Resources

Learning & Resources

Published benchmark scores saturate within months of release, so the live arenas and usage boards are what practitioners now read.

  1. StandingBenchmarks saturate within months of release, so a published score dates fast. The arenas and usage boards track what people actually reach for.
  2. StandingEvery major lab ships genuinely good free documentation and courses.
  3. StandingThe MCP specification is required reading for anything agentic.
Explore resources