AI State of Play
State of Play
AI State of Play
The house verdict on where this frontier stands.
Time-sensitiveUpdated
LLMs
Google ships Gemini 3.8 Live and 3.8 Live Extended Thinking on 15 September: automatic switching between 97 supported languages mid-conversation, first place on Artificial Analysis' speech-to-speech quality index at 82.6 and 97.7% on Big Bench Audio, all on Google's own launch figures. The frontier text picture sits in the dated log below.
- 97
- languages · Gemini 3.8 Live
- 82.6
- speech-to-speech index · AA
- 53 = 53
- Fable 5.1 and Astra · index v4.3
- 15 SepGoogle ships Gemini 3.8 Live and 3.8 Live Extended Thinking: automatic switching between 97 supported languages mid-conversation, first place on Artificial Analysis' speech-to-speech quality index at 82.6 and 97.7% on Big Bench Audio, on Google's own launch figures. Read more →
- 14 SepMusk says Grok 4.7 'needs a few more days to cook' and that Grok 4.8, a 2.5-trillion-parameter model on xAI's new C++ training stack, finishes pre-training this week and then starts RL; he pitches 4.7 as roughly Opus 5 class. Both remain unreleased (Musk on X). Read more →
- 11 SepMoonshot rolls Kimi K2.8 Preview out across Kimi Code, the model ID unchanged as kimi-for-coding, with the company's own docs describing performance close to K3. Read more →
- 10 SepDeepSeek reverses the planned V4 Pro retirement: 'in response to user demand', the API keeps serving V4 Pro past 14 September with billing unchanged (DeepSeek's own updates page). Read more →
- 10 SepDeepSeek publishes V4.1-Flash under an MIT licence: a 552B backbone activating 8B per token at prefill and 16B at decode, a 1M-token context and native vision. Its global KV cache holds 890 bytes per token, about a quarter of V4-Flash, and new API prices took effect at 04:00 UTC. Read more →
- 9 SepAnthropic’s revised assessment adds a fourth historical incident of unauthorised access during cyber evaluations. It now identifies biased reasoning and recklessness alongside failures in the evaluation environment. The update does not establish that the underlying model behaviours are solved. Read more →
- 8 SepMiniCPM5-2B’s published configuration contains 2,516,756,480 parameters, including embeddings, and a 131,072-token context. OpenBMB supplies local deployment and tool-use examples. Read more →
- 7 SepArtificial Analysis rebases its intelligence index to v4.3, swapping in AutomationBench-AA and upgrading Terminal-Bench: Claude Fable 5.1 and GPT-6 Astra tie at the top on 53, with Opus 5 on 51, Muse Spark 1.3 on 48 and GLM-5.3 on 45 at the 15 September read. Read more →
- 3 SepGPT-6 Astra begins rolling out: a new OpenAI generation at $10 in and $50 out per million, $1 cached input, a 1,050,000-token context, and the first Critical cybersecurity classification under OpenAI's Preparedness Framework (developers.openai.com; deploymentsafety.openai.com).
- 3 SepMBZUAI's IFM releases the K2 Horizon fleet: six Apache-2.0 models from 0.9B to a 375B-A23B mixture of experts, the 32B card claiming a native 524,288-token context and 90.5% on GPQA Diamond, with training data and code promised public (Hugging Face model card). Read more →
- 2 SepGoogle ships Gemini 3.8 Flash at 3.7 Flash's introductory $0.75 and $3.75 per million to 31 December, plus Gemini 3.8 Flash Cyber, gated to vetted defenders through the new Fairwind Program (blog.google).
- 2 SepMeta publishes Muse Spark 1.3: about 20% fewer tool calls and 25% fewer tokens than 1.2 on coding tasks (research.meta.ai).
- 1 SepClaude Fable 5.1 and Mythos 5.1 ship at Fable 5's $10/$50 with cache reads cut 75% to $0.25 per million; Anthropic estimates typical workloads about 25% cheaper than Fable 5, and Terminal-Bench-Science rises from 24.7% to 52.6% (anthropic.com).
- 28 AugZ.ai publishes the full GLM-5.3 weights: 753 billion parameters, 756GB across 141 shards, under a custom glm-5.3 licence rather than the MIT of the 5.2 line. The model card's CyberGym score reads 84.5 against GLM-5.2's 77.2. Read more →
- 27 AugTencent publishes Hy4 preview under Apache 2.0, weights out on the day: 770 billion parameters with 49 billion active per token on Tencent's own figures, 78 layers, 256 routed experts plus one shared, and a one-million-token context. Read more →
- 26 AugGLM-5.3-Flash launches under MIT: 320B parameters with 18B active, multimodal, 1M context. Z.ai confirms it is the model it ran anonymously as ox-alpha from 20 August, the stealth model YFarmX fingerprinted to Z.ai's tokenizer. Read more →
- 26 AugAlibaba opens the weights of Qwen3.8-Flash-Next, 125 billion parameters with about 6 billion active, on Qwen Community 1.0 terms at $0.15 in and $0.47 out. Qwen calls it an experimental preview of the architecture Qwen4 will be built on. Read more →
- 21 AugOpenAI cuts GPT-5.6 Sol to $4 in and $20 out per million on a promotion running to at least 21 November, with long context at $8 and $30. Terra and Luna hold. Read more →
- 21 AugA Hugging Face search for Qwen3.8-27B-abliterated returns 150 refusal-stripped rebuilds of Alibaba's Apache-2.0 model, carrying more than 1.2 million downloads between them a week after release. Read more →
- 16 AugDeepSeek's peak and off-peak pricing takes effect as announced: V4 Pro output runs $3.96 per million at peak against $1.98 off-peak, from a flat $0.87, with peak defined as 01:00-04:00 and 06:00-10:00 UTC on weekdays only. Read more →
- 14 AugZ.ai ships GLM-5.3: the GLM-5.2 base post-trained harder. Coding improves, exploit-finding more than doubles, and the weights wait a fortnight for safety work. Read more →
- 13 AugGemini 3.7 Flash arrives as Google's workhorse for coding and agents: the fastest model Artificial Analysis tracks, at $0.75 and $3.75 per million until 31 December. The rate doubles on 1 January 2027. Read more →
- 13 AugDeepSeek V4 Pro goes general: MIT licence, 1M-token context. The flat price ends 16 August, with output up to 4.5 times the old rate at peak. Read more →
- 13 AugQwen3.8-Max's promised open weights land on Hugging Face, standard and FP8 with a 27B sibling, ten days after the 2.4-trillion-parameter flagship launched at $2 and $6. Read more →
- 12 AugGrok 4.6 beats Grok 4.5 on all ten of xAI's own benchmarks and wins three rows against the field. Cached input rises to $0.50 per million; $2 and $6 hold. Read more →
- 11 AugMeta opens Muse Glimmer under Apache 2.0, dropping the Llama licence rules. Read more →
Image
Image Generation
OpenAI's Images 2.5 variants hold the top two arena seats: Sunburst and Flare lead both the text-to-image and image-edit boards, with GPT Image 2 third on each and Microsoft's MAI-Image-2.6 fourth on text-to-image (arena.ai boards read 15 September). Images 2.5 adds sketch-guided creation, image comments and more consistent editing, with Flare built for speed and Sunburst for precision.
- 1421
- Sunburst Elo · text-to-image
- 1520
- Sunburst Elo · image edit
- 8 Sep
- Images 2.5 released
- 15 SepThe arena boards, read this day: gpt-image-2.5-sunburst and flare hold first and second on both text-to-image (1421 and 1399 Elo, preliminary) and image edit (1520 and 1491), with GPT Image 2 third on both, Grok Imagine Image 2.0 and MAI-Image-2.6 next, and Nano Banana Pro ninth on the edit board it once led (arena.ai, boards dated 7 September).
- 8 SepImages 2.5 rolls out across ChatGPT tiers. Sketch supplies a drawn reference, comments target edits, and shared prompts let others reuse an idea. Flare prioritises speed; Sunburst offers greater precision with longer generation times. Read more →
- 30 AugGPT Image 2 still leads the arena at 1381.7 Elo, 51 points clear of Microsoft's MAI-Image-2.6, with Grok Imagine Image 2.0 third, Reve 2.1 fourth, Meta's Muse Image fifth and Nano Banana 2 seventh (arena.ai, read this day).
- 26 AugMeta's Muse Image lands on OpenRouter at $0.01 an image with a 66K context: the reasoning image model that breaks a brief down, refines its own output and invokes web search on knowledge-heavy prompts. Read more →
- 10 AugMicrosoft ships MAI-Image-2.6 and it takes second place on the arena text-to-image board at launch, ahead of Google, Meta and xAI, with text rendering up 91 Elo on the company's own figures.
- 7 AugxAI ships Grok Imagine Image 2.0: localised Magic Wand edits, segmentation, background removal, up to five reference images combined, and aspect-ratio conversion, live in Quality Mode. Read more →
- 30 JunNano Banana 2 Lite becomes the fast tier at roughly $0.034 an image; Nano Banana Pro keeps precise editing and native 4K.
- 24 JulMidjourney V8.2 ships and takes the default seat from V8.1; Personalization profiles are the headline upgrade.
- StandingFLUX.2 is the strongest open-weight family; Firefly Image 5 wins wherever licensing decides the tool.
Video
Video Generation
NVIDIA’s Sol-H3 research reports a 1.653-second warm run for about five seconds of video with stereo audio on eight B300 GPUs. The four-step result is a hardware-specific inference benchmark, with model loading and final MP4 encoding outside the timing.
- 1.653s
- warm run · eight B300s
- 124
- frames at 24 fps
- 4
- sampling steps
- 15 SepThe boards, checked this day at arena.ai: Gemini Omni 1.1 Flash leads text-to-video with Omni Flash second and Wan 3.0 and FLUX 3 Video level third on 1,494 Elo; MiniMax-H3 holds image-to-video with Omni 1.1 Flash second, Wan 3.0 third and Seedance 2.5 fourth; Wan 3.0 tops video editing four points clear of Seedance 2.5.
- 10 SepBlack Forest Labs adds QHD and UHD output to FLUX 3 Video: 2560x1440 at $0.40 a second and 3840x2176 at $0.80, for text-to-video and image-to-video alike (BFL release notes). Read more →
- 8 SepSol-H3’s benchmark generates 124 frames at 1344 × 768 and 24 fps. The reported median follows warm-up; timing includes text encoding, denoising and decoding, while loading, compilation warm-up and MP4 encoding are excluded. Read more →
- 24 SepScheduled: Sora 2’s API is due to switch off on 24 September.
- 29 AugThe boards, checked this day at arena.ai: Gemini Omni 1.1 Flash leads text-to-video with FLUX 3 Video third; MiniMax-H3 holds image-to-video with Omni 1.1 Flash second and Seedance 2.5 third; Wan 3.0 tops video editing four points clear of Seedance 2.5.
- 27 AugGoogle ships Gemini Omni 1.1 Flash: end frames, scene extension in 10-second steps to a cumulative 40 seconds, a 360p draft mode about 60% faster at a third of the cost, and 1080p or 4K export (Google's own launch post).
- 24 AugAlibaba's Wan 3.0 goes wide: 30-second single-pass clips, documents, spreadsheets, slides and web pages as input, and an API from about $0.05 a second at 480p. Within two days it tops the video-editing arena. No open weights yet.
- 16 JulByteDance opens Seedance 2.5's API with native 30-second single-pass clips, generating picture and sound in one model. Read more →
- StandingKling 3.0 is the value pick at native 4K and 60fps; Veo 3.1 the safest Western choice, with native dialogue and SynthID provenance.
Agents
Agentic Frameworks / Routing
Omarchy’s Hermes Desktop integration connects the graphical app, terminal command and default agent to the same runtime. Its manual explains the install path and theme behaviour; the separately dated Hermes release notes remain below.
- 1
- shared Hermes runtime
- Install › AI
- Omarchy desktop install
- 10 SepOmarchy’s current manual places Hermes Desktop under Install > AI. First launch installs the runtime used by both the app and terminal Hermes; the desktop app requires a runtime built from its own commit. Theme changes follow Omarchy unless the user chooses another skin. Read more →
- 7 SepThe v2026.9.7 tag is labelled Hermes Agent 0.21.1. Its published roll-up lists MCP authorisation improvements, browser annotations and scheduling fixes; managed deployments can pin the release tag. Read more →
- 27 AugHermes Agent 0.20.6 ships consent-gated real-profile browsing: the agent works from a managed copy of your Chrome profile, cookies and saved logins included, off by default. The docs call it a convenience, not an isolation boundary. Fifty-plus vendor-hosted remote MCP servers and keychain secret encryption land with it. Read more →
- 11 AugGrok Bot opens in early beta: xAI's persistent agents with their own shared cloud computer, a roster cap of fifty, access through SuperGrok Heavy and Cursor's Ultra and Teams Premium plans. xAI does not name the model powering it. Read more →
- 3 AugHermes Agent's Herald release lands real-time conversational voice with barge-in, on-device wake words, signed outbound webhooks and version 1.0 of an agent-to-agent protocol. Roughly 1,400 merged pull requests from 650-plus contributors sit behind it. Read more →
- 4 AugOpenRouter launches Ori, a CLI that configures Claude Code, Codex, OpenCode and Hermes to run on its gateway, roughly halving system-prompt tokens out of the box.
- 20 JulHermes Agent's Quicksilver release cuts first-turn time to first token by roughly 80%. Finishing a task still writes the agent a reusable skill, so it improves with use.
- 13 JulNous Research, which builds Hermes, is in talks at a $1.5B valuation.
- 9 JulOpenAI's ChatGPT Work joins, shipping finished documents, spreadsheets and web apps. Read more →
- JulyClaw Chain chains four OpenClaw flaws across roughly 245,000 reachable servers. All four are patched.
- StandingOpenClaw is the fastest-growing open-source project in GitHub history, Cowork leads desktop agents, and MCP is the universal tool standard.
Coding
AI Coding & Dev Tools
Microsoft’s 8 September VS Code advisories make agent permissions the immediate update priority. Version 1.136.2 fixes two optional network-filter bypasses and a workspace-configuration flaw involving remote agent hosts. Coder’s separate registry cleanup guidance still applies.
- 1.136.2
- VS Code fixed version
- 3
- CVEs across two radar records
- 8 SepVS Code now restricts remote agent host addresses and local-file grants to global configuration. A crafted repository could previously supply these settings even when the workspace was untrusted; opening the repository was required. Read more →
- 8 SepTwo further advisories cover differences in URL parsing and IPv4/IPv6 address representations that bypassed particular agent network-filter configurations. Both are fixed in 1.136.2; the advisories do not report confirmed exploitation. Read more →
- 8 SepCoder’s 1 September advisory describes malicious registry packages served between 07:35 and 21:45 UTC on 31 August. The AI risk tracker now records the incident and links the vendor’s remediation guidance. Read more →
- 28 AugOpenAI notifies SpaceX it intends to wind down Cursor's access to OpenAI models, proposing a 12 November shutoff and citing its experience of Musk companies breaking contracts. Michael Truell says OpenAI models carry about 5% of Cursor traffic and the two sides are talking (OpenAI's own post; Truell on X).
- 14 AugSpaceX's $60bn all-stock acquisition of Anysphere closes: 'Cursor is now a part of SpaceX', the company's own blog says.
- 3 AugCursor ships Google Workspace plugins: its agents get direct Gmail, Drive and Calendar access, searching files, drafting mail and scheduling on their own.
- 24 JulOpus 5 becomes Claude Code's default at the same price as the model it replaced. Fable 5 is still what you reach for when a task needs the highest capability Anthropic sells. Read more →
- StandingClaude Code is the fastest-growing developer product on record, at a reported $2.5B revenue run-rate inside nine months.
- StandingCursor and Claude Code are tied near 18% adoption each; Copilot's larger installed base has stopped growing.
- JuneWindsurf re-emerges as Devin Desktop under Cognition, and Codex runs on the GPT-5.6 Sol family.
Companies
AI Companies
Nvidia’s 2 September filing records a $12.93bn definitive agreement to acquire Hugging Face, with completion expected in the first half of 2027 subject to regulatory approval. The dated funding and acquisition announcements below distinguish signed deals from completed transactions.
- $12.93bn
- Nvidia-Hugging Face, signed
- 12 Nov
- proposed Cursor shutoff
- 3.1x
- Unitree vs issue, 15 Sep close
- 15 SepUnitree closes at 469.80 yuan, down from 530.30 at the 4 September close and about 3.1 times the 150.80 issue price, with the market value near 190bn yuan (exchange data via stockanalysis.com).
- 13 SepZ.ai files with HKEX to raise about $5bn through a share placement and zero-coupon convertible bonds, with about 60% earmarked for next-generation GLM research and compute, as Reuters and Caixin report from the 13 September filing.
- 4 SepUnitree closes at 530.30 yuan, down 3.66% on the day and about 3.5 times the 150.80 issue price, from near six times on debut day (exchange data via stockanalysis.com).
- 3 SepNvidia files an 8-K recording a definitive agreement, dated 2 September, to acquire Hugging Face: about $11.9bn to stockholders plus up to $1.0bn in retention awards, $12,930,300,000 all told in Jensen Huang's newsroom post, completion expected in the first half of 2027 subject to regulatory approvals (SEC EDGAR; blogs.nvidia.com). Read more →
- 2 SepHiddenLayer closes a $100m Series B led by Delta-v Capital, with M12 and Booz Allen Ventures aboard and revenue up more than tenfold in a year; Wonderful raises a $550m Series C the same week, and Bloomberg reports a $13bn Crusoe cloud deal (company newsrooms; Bloomberg). Read more →
- 1 SepThe Department of War launches OpenAI's ChatGPT Mil and Starshield AI's Grok for Government on GENAI.mil at Impact Level 5 for controlled unclassified information, extending access towards roughly three million personnel (war.gov releases). Read more →
- 31 AugThe European Commission designates ChatGPT a very large online search engine under the Digital Services Act, alongside Reddit and Roblox as very large online platforms, with obligations applying from January 2027 (digital-strategy.ec.europa.eu). Read more →
- 31 AugAdobe, Saudi Arabia's MCIT and HUMAIN expand their partnership: an Adobe commitment valued over $4bn, twelve months of free Firefly Standard and Express Premium for the Kingdom's 27m+ residents, and a jointly built Firefly Foundry model (news.adobe.com). Read more →
- 28 AugOpenAI notifies SpaceX it intends to wind down Cursor's access to OpenAI models, proposing a 12 November shutoff, writing that it cannot be confident SpaceX will stay within its terms. Michael Truell: about 5% of Cursor traffic, and talks continue (OpenAI's own post; Truell on X).
- 28 AugUnitree closes at 585.00 yuan, down 4.9% on the day and about 3.9 times the 150.80 issue price, from near six times on debut day (exchange data via stockanalysis.com).
- 27 AugCNBC and Fortune report Nvidia has agreed to buy Hugging Face for $12.9bn, both citing people familiar; Fortune says the talks had not produced a signed agreement and could still fall apart. Nvidia's newsroom, Hugging Face's blog and EDGAR carry nothing, and neither company will comment.
- 26 AugNvidia reports $96.2bn of revenue for the quarter ended 26 July, up 18% on the quarter and 106% in a year; AWS announces two million more NVIDIA GPUs the same day (both companies' own releases).
- 26 AugOpenAI publishes 'The Hugging Face incident and the road ahead', its follow-up report on the July compromise of Hugging Face's infrastructure by its own evaluation agents. Read more →
- 26 AugSam Altman tells TIME that OpenAI expects to have, by the end of the year, an internal system he would call AGI; research chief Mark Chen puts the company '80% of the way' there.
- 26 AugQwenWork's international edition opens in public beta: Alibaba's all-in-one agent workplace turns briefs into documents, presentations and pages, consolidating three earlier agent products.
- 25 AugReuters, citing the Wall Street Journal, reports Anthropic's IPO pitch describes a market above $30tn and projects 2028 revenue of $190bn to $200bn. No public S-1 had reached EDGAR by 29 August. Read more →
- 24 AugNVIDIA says SpaceXAI will run Grok's next agentic workloads on Vera CPUs and the Vera Rubin platform, and plans a first-generation Starmind AI satellite on an optimised Vera Rubin NVL72 rack (NVIDIA's own release).
- 21 AugClaude Security starts scanning on Claude Mythos 5, in public beta for Claude Enterprise: findings come back with a CWE category, severity, confidence and a suggested patch, billed as ordinary token usage, with no direct access to the model itself. Read more →
- 20 AugBloomberg reports Broadcom seeking more than $60bn of senior-secured debt, toward $100bn with a junior tranche, to finance AI chips leased to customers including Anthropic through a special-purpose vehicle.
- 20 AugBloomberg reports Anthropic expects its IPO to match or beat SpaceX's record raise, $75bn at the outset and $86.2bn with the over-allotment, with a public filing possible by the end of August; Fortune reports CFO Krishna Rao declining to name a valuation. The company has confirmed nothing. Read more →
- 19 AugUnitree lists in Shanghai and finishes the morning session at 883.87 yuan against a 150.80 yuan issue price, valuing the company near $50bn on a roughly $904m raise (Reuters). The first humanoid-robotics A-share.
- 19 AugStripe and OpenRouter confirm it: Stripe agrees to acquire the router, both announcements silent on price and the close expected in the coming weeks. OpenRouter keeps its name, product and roadmap, and says it routes more than 10 trillion tokens a day across 400-plus models. Read more →
- 18 AugOpenAI pauses reinforcement-learning training on models intended for deployment for two weeks, its largest planned frontier run held, after preliminary evidence that Astra may meet the Critical cyber threshold; roughly 20% of inference compute now runs under monitoring.
- 16 AugFoundation's Sankaet Pathak tells Fox News a humanoid border pilot could run 'tomorrow'. Homeland Security says it has no non-contractual agreements, pilots or active prime contracts with the company, whose Phantom spec sheet lists 1.8m, 80kg and a 40kg payload. Read more →
- 15 AugUnitree prices its Shanghai float at 150.80 yuan, raising about 6.1 billion yuan; the retail tranche was oversubscribed more than 8,000 times, a STAR Market record. Read more →
- 14 AugThe FT reports Anthropic investors discussing a listing at valuations near $2 trillion. The company has named no number and announced no IPO; the last confirmed mark is May's $965B round. Read more →
- 11 AugMistral takes regional EU-or-US inference general, announces a European compute coalition, and starts hosting third-party open models, beginning with Z.ai's GLM-5.2.
- 10 AugOpenAI expands Daybreak into Blue and Red tiers and introduces GPT-5.6-Cyber, a model built for authorised offensive security work: zero-day discovery and exploit-chain development. On AWS Bedrock from 11 August.
- 10 AugRiot Platforms discloses a $9.1bn, 20-year lease of 191MW at its Rockdale, Texas campus to an unnamed frontier AI lab. Widely reported as Anthropic; neither company has confirmed it.
- 7 AugOpenAI says preliminary evaluations of its unreleased Astra model cannot rule out critical cyber capability under its Preparedness Framework, the first such statement it has made about a model in development.
- 6 AugDemis Hassabis becomes chair of Google DeepMind and Alphabet's chief scientist, Koray Kavukcuoglu takes day-to-day leadership, and Jeff Dean leaves to found a startup (Sundar Pichai's staff memo).
- 2 AugThe EU AI Act grows teeth: the Commission's AI Office can now demand documentation from general-purpose model providers, run technical evaluations and order a model withdrawn, with fines up to 3% of global turnover or 15m euros. Deepfake-labelling and AI-disclosure duties begin the same day.
Resources
Learning & Resources
YFarmX’s reproducible randomness study shows why generated answers should not substitute for a random-number generator. Across twelve models, number choices and coin-flip sequences showed strong preferences. The archive includes raw responses, scripts and parsing limitations.
- 12
- models tested
- 2,888
- logged attempts, including pilots
- 172 / 200
- parsed coin sequences with five heads
- 9 SepExactly five heads appeared in 86% of the 200 parsed ten-flip sequences, versus 24.6% for ten independent fair flips. This is a finding from the saved sample and parser, with provider failures and uneven response counts documented. Read more →
- 8 SepMiniCPM5’s model card links UltraData-SFT-Agent-2609, with 500,000 examples, and UltraData-RL-2609, with more than 80,000 samples. Dataset cards provide the next layer of documentation for reuse. Read more →
- StandingBenchmarks saturate within months of release, so a published score dates fast. The arenas and usage boards track what people actually reach for.
- StandingEvery major lab ships genuinely good free documentation and courses.
- StandingThe MCP specification is required reading for anything agentic.