The AI Index
AI Knowledge Hub
62 entries · sector notes verified 24 Jul 2026
llmsLarge language models are the engine room of modern AI: reasoning, writing, coding, search and agent control all run through them. The frontier now competes on long-horizon reliability, tool use and price rather than raw benchmark scores.9 entries
◈ Where things stand: The frontier is getting cheaper by the week. Alibaba's Qwen3.8-Max went official on 3 August: 2.4 trillion parameters, a 1M-token window, $2 in and $6 out per million tokens, and a promise to publish the weights, the first Max-tier Qwen to open up. DeepSeek refreshed V4 Flash on 31 July with a stronger agent-tuned build at the same address. They are chasing a leaderboard Anthropic now tops twice over: Claude Opus 5 (24 July, $5 and $25, half what Fable 5 costs) leads the Artificial Analysis index outright, Fable 5 sits level behind it and holds first on the Arena text board, and Kimi K3's 1.4TB of open weights make it the strongest open model anyone can download. OpenAI answered on price on 30 July, cutting GPT-5.6 Luna by 80% to $0.20 and $1.20 and Terra by 20% to $2 and $12, with Sol unchanged at $5 and $30. The everyday bargain is Sonnet 5: introductory pricing of $2 in and $10 out runs until 31 August, with the standard $3 and $15 returning from 1 September.
- Claude
Anthropic's frontier family
Opus 5 (24 July) leads the Artificial Analysis index at the same $5/$25 as the Opus 4.8 it retires, half Fable 5's price; Fable 5 sits level behind it and tops the Arena text board. Sonnet 5 stays the everyday default with a 1M-token window, on introductory $2/$10 pricing until 31 August.
claude.ai ↗ - GPT-5.6
OpenAI's flagship tiers
GPT-5.6 went general on 9 July 2026 in three tiers, Sol, Terra and Luna, after a brief government-brokered preview. On 30 July OpenAI cut Luna by 80% to $0.20 and $1.20 per million tokens and Terra by 20% to $2 and $12, with Sol unchanged. GPT-5.5 Instant stays the everyday default inside ChatGPT.
chatgpt.com ↗ - Gemini 3.1 Pro
Google DeepMind's generalist
Gemini 3.1 Pro holds Google's Pro tier while 3.5 Pro stays in partner testing; the Flash seat moved on to Gemini 3.6 Flash on 21 July 2026.
gemini.google.com ↗ - Grok 4.5
xAI's real-time model
Grok 4.5 shipped on 8 July 2026 at $2 in and $6 out per million tokens with a 500K context; the older Grok 4.3 stays on at $1.25 and $2.50 as the budget line.
grok.com ↗ - Kimi K2.6 / K3
Moonshot AI's open flagship
Kimi K3, a 2.8-trillion-parameter flagship, launched on 16 July; its full weights landed on 27 July as the largest open-weight release yet, roughly 1.4TB even at four-bit precision, and the strongest open model on the Artificial Analysis index, under a bespoke licence with a revenue clause for model-as-a-service hosts.
kimi.com ↗ - Qwen 3.8 Max
Alibaba's flagship line
Alibaba's flagship: 2.4 trillion parameters with 95 billion active, a 1M-token window, $2/$6 per million tokens, launched 3 August 2026. Built for long autonomous runs rather than single answers, and the first Max-tier Qwen whose weights Alibaba says it will publish.
chat.qwen.ai ↗ - GLM-5 series
Z.ai's agent-endurance models
The series flagship is GLM-5.2, released 13 June 2026 with open weights under an MIT licence and a 1M-token context; NIST's CAISI published an assessment of it in July.
z.ai ↗ - DeepSeek
The price-floor setter
The lab that made frontier reasoning cheap. V4-Flash-0731 (31 July 2026) swapped a stronger agent-tuned build in at the same API address and price, and its open releases keep resetting the floor the market has to answer.
deepseek.com ↗ - Muse Spark
Meta's proprietary frontier
Meta Superintelligence Labs' first Muse model (8 April 2026) reasons with parallel agents in Contemplating mode; Muse Spark 1.1 (9 July) added a 1M context, tool and computer use, and a paid API at $1.25 in and $4.25 out.
ai.meta.com ↗
image generationImage models have split into two camps: arena-topping generators for ideas, and production systems with real editing, typography and brand control. The right pick depends on whether the image is the destination or a step in a workflow.8 entries
◈ Where things stand: Checked on 6 August, OpenAI's GPT Image 2 still leads the arena by a record margin, and it plans or searches the web before it draws rather than guessing at what you asked for. Google's Nano Banana Pro owns precise editing and native 4K, with Nano Banana 2 Lite (30 June) now the fast tier at roughly $0.034 an image. Midjourney's V8.2 shipped on 24 July and took the default seat from V8.1, with its Personalization profiles the headline upgrade, FLUX.2 remains the strongest open-weight family, and Adobe's Firefly Image 5 is the pick wherever licensing decides the tool.
- GPT Image 2
OpenAI's reasoning image model
Top of the image arena by the largest lead recorded, and the first mainstream model that reasons about a brief, checks references and self-corrects before rendering.
chatgpt.com ↗ - Nano Banana Pro
Google's editing powerhouse
Gemini 3 Pro's image engine: native 4K, best-in-class edits and character consistency. Nano Banana 2 covers the fast-and-cheap end of the family.
gemini.google.com ↗ - Midjourney V8.2
The art direction standard
Still the strongest pure art direction. V8.2 took the default seat on 24 July with Personalization profiles as the headline upgrade, on top of V8.1's native 2048px output; Niji 7 handles anime styles.
midjourney.com ↗ - FLUX.2
Black Forest Labs' open family
The open-weight family to beat: Pro, Flex, Dev and Klein tiers, camera-accurate lighting and up to ten reference images for character consistency.
bfl.ai ↗ - Ideogram
Text rendering specialist
The reliable choice when the image must carry words: posters, logos, packaging and layout-aware design work.
ideogram.ai ↗ - Recraft V4
Design-first generation
Built for designers rather than prompters: vector output, exact brand colours and print-ready renders straight from the model.
recraft.ai ↗ - Seedream
ByteDance's value line
ByteDance's image line delivers 4K output at some of the lowest per-image prices on the market, with strong bilingual text rendering.
seed.bytedance.com ↗ - Firefly Image 5
Adobe's licensed-data model
Trained on licensed data with indemnification for enterprise use, and now a hub that also fronts partner models inside Creative Cloud.
adobe.com ↗
video generationAI video crossed from novelty to production this year: native audio, multi-shot control and 4K output are now table stakes at the top of the market.10 entries
◈ Where things stand: ByteDance opened Seedance 2.5's API on 16 July with native 30-second single-pass clips, the longest one-shot generation anyone offers. The boards split above it: Google's Gemini Omni Flash tops text-to-video while Seedance 2.0 keeps the image-to-video and editing crowns. Kling 3.0 is the value pick at native 4K and 60fps, and Veo 3.1 the safest Western choice with native dialogue and SynthID provenance. One date to diarise: Sora 2's API switches off on 24 September 2026. Open weights are genuinely usable now.
- Seedance 2.0
ByteDance's board leader
Top of the image-to-video and editing boards. One model generates picture and sound together, so lips, foley and music actually match the frame.
seed.bytedance.com ↗ - Veo 3.1
Google DeepMind's flagship
Native 48kHz dialogue, 4K output and SynthID provenance marks. The default choice where brand safety and clean licensing decide the choice.
labs.google ↗ - Kling 3.0
Kuaishou's value flagship
Native 4K at 60fps, 15-second multi-shot storyboards and lip-sync in five languages, at roughly $0.10 a second.
klingai.com ↗ - Sora 2
OpenAI · being retired
Being retired: app and web access closed on 26 April and the API ends on 24 September 2026. Plan migrations now rather than building on it.
openai.com ↗ - Runway Gen-4.5
The filmmaker's control surface
No longer top of the leaderboards but still the best control surface in the business: motion brushes, keyframes and a real editing timeline.
runwayml.com ↗ - Ray3 · Dream Machine
Luma's HDR pioneer
First native 16-bit HDR video model, with Ray3 Modify for video-to-video work inside the Dream Machine 2.0 studio.
lumalabs.ai ↗ - Wan 2.7
Alibaba's open video family
The strongest fully open video family, Apache-licensed with first-and-last-frame control for planning shots precisely.
wan.video ↗ - LTX-2.3
Lightricks' open 4K model
A 22B open model producing native 4K at 50fps with stereo audio; free for commercial use below a revenue threshold.
lightricks.com ↗ - Hailuo 2.3
MiniMax's motion specialist
MiniMax's line is prized for expressive, physical motion: the model animators reach for when movement has to feel alive.
hailuoai.video ↗ - Gemini Omni Flash
conversational video generation
Google's conversational video model (30 June 2026) generates and edits ten-second clips through the Gemini API and now tops the text-to-video arena.
deepmind.google ↗
agentic frameworks / routingAgent frameworks sit between models and outcomes: routing, tools, memory and the guardrails that decide whether an agent is useful or dangerous. Reliability and control now count for more than raw capability.10 entries
◈ Where things stand: Hermes Agent found its voice on 3 August: the Herald release adds real-time conversational speech you can interrupt mid-sentence, on-device wake words, signed outbound webhooks and version 1.0 of an agent-to-agent protocol, with around 650 contributors behind it. That lands three weeks after Quicksilver cut first-turn time to first token by roughly 80%, and its defining trick still stands: finish a task and it writes itself a reusable skill, so the agent you run in December is more capable than the one you installed. Nous Research, which builds it, is in talks at a $1.5B valuation. OpenRouter joined the harness game on 4 August with Ori, a CLI that tunes Claude Code, Codex, OpenCode and Hermes to run through its gateway. OpenClaw remains the fastest-growing open-source project in GitHub history; Anthropic's Cowork leads desktop agents, Google's Gemini Spark runs around the clock on cloud VMs, and OpenAI's ChatGPT Work ships finished documents, spreadsheets and web apps. MCP is now the universal tool standard. The warning label stands: July's Claw Chain disclosure chained four OpenClaw flaws across roughly 245,000 reachable servers, all since patched.
- OpenClaw
Self-hosted agent gateway
The breakout self-hosted gateway: one daemon, twenty-plus messaging channels, skills and scheduled jobs, any model behind it. Since April it needs API billing rather than a Claude subscription, and it should be hardened before exposure.
openclaw.ai ↗ - Hermes Agent
Self-improving open agent
Nous Research's open-source, self-hosted agent, and the one that changes as you use it: finishing a task writes a reusable skill back into the agent. The Herald release (3 August) added real-time voice with barge-in, wake words and an agent-to-agent protocol; Quicksilver (20 July) had already cut first-turn time to first token by roughly 80%.
hermes-agent.nousresearch.com ↗ - Claude Cowork
Anthropic's desktop agent
Anthropic's desktop agent for real file-and-app work, with domain packs for legal and operations tasks. Its launch is what markets nicknamed the SaaSpocalypse.
claude.com ↗ - Gemini Spark
Google's always-on agent
Google's always-on agent living in a cloud VM: deep Gmail, Docs and Calendar access, MCP connectors from day one, and it keeps working while your laptop sleeps.
gemini.google.com ↗ - ChatGPT Agent
OpenAI's consumer agent
Agent mode inside ChatGPT with its own virtual computer for browsing, files and multi-step tasks under supervision.
chatgpt.com ↗ - Claude Agent SDK
Anthropic's builder stack
The loop behind Claude Code, embeddable in your own product, alongside a managed runtime (beta) where Anthropic hosts the agent for you.
docs.claude.com ↗ - OpenAI Agents SDK
Handoffs and guardrails
OpenAI's production framework with explicit handoffs, guardrails and tracing; MCP support landed early this year.
openai.github.io ↗ - Model Context Protocol
The connector standard
The USB-C of agent tooling: one protocol that lets any client use any tool or data server. Every major platform now speaks it.
modelcontextprotocol.io ↗ - OpenRouter
One API, every model
One API over every major model with live rankings drawn from real usage: the quickest way to A/B models behind an agent without rewrites. Its Ori CLI (4 August 2026) tunes Claude Code, Codex, OpenCode and Hermes to run through the gateway.
openrouter.ai ↗ - ChatGPT Work
OpenAI's finished-work agent
OpenAI's agent for long-running jobs (9 July 2026): powered by GPT-5.6, it works for hours and returns finished documents, spreadsheets, slides and web apps.
openai.com ↗
ai coding & dev toolsAI coding assistants went from autocomplete to delegation: the leading tools now plan, edit across a repo, run tests and open pull requests on their own.8 entries
◈ Where things stand: Cursor gave its agents keys to the office on 3 August: Google Workspace plugins let them search Gmail and Drive, draft mail and manage the calendar directly. It is tied with Claude Code at roughly 18% adoption each while Copilot's larger installed base has stopped growing. Claude Code's default model changed on 24 July, when Opus 5 landed at the same price as the model it replaced; Fable 5 remains the tier above it for the hardest work, and the product is already the fastest-growing developer tool on record, at a reported $2.5B revenue run-rate inside nine months. Codex runs on the GPT-5.6 Sol family, and Windsurf re-emerged in June as Devin Desktop under Cognition.
- Claude Code
Anthropic's terminal agent
Terminal-native and repo-aware, with Dynamic Workflows for long unattended runs; Opus 5 became its default on 24 July, with Fable 5 the tier above it. Included with a Claude Pro plan.
claude.com ↗ - Cursor
Anysphere's AI-native editor
The AI-native editor: Tab completions plus Composer cloud agents that work branches in parallel VMs. Model-agnostic across Claude, GPT, Gemini and Grok.
cursor.com ↗ - GitHub Copilot
The installed-base leader
Still the default inside VS Code and enterprises, with a free tier, IP indemnity and new flexible usage billing from 1 June.
github.com ↗ - Codex
OpenAI's cloud engineer
OpenAI's asynchronous cloud engineer: give it an issue and it returns a tested pull request. Now powered by the coding-tuned GPT-5.6 Sol.
openai.com ↗ - Antigravity 2.0
Google's agent-first IDE
Google's agent-first IDE built around the Gemini 3.5 line, with a browser-testing loop and a generous free tier via a Google account.
antigravity.google ↗ - Devin Desktop
Windsurf's successor
Windsurf's next chapter after the Cognition deal: the Devin autonomous engineer wrapped in the familiar editor, relaunched in June.
cognition.ai ↗ - Kiro
AWS's spec-driven agent IDE
Amazon's take: spec-first development where agents implement against written requirements, aimed at teams that want process over improvisation.
kiro.dev ↗ - Replit Agent
Prompt to deployed app
From prompt to deployed app in one place: environment, database and hosting handled for you. The fastest route for non-specialists.
replit.com ↗
ai companiesThe market is now a platform race: strong models are table stakes, and distribution, chips, agents and developer ecosystems decide who compounds.9 entries
◈ Where things stand: Brussels switched its enforcement powers on on 2 August: the EU's AI Office can now demand documentation from general-purpose model providers, run technical evaluations and order a model off the market, with fines reaching 3% of global turnover. Unitree's Shanghai IPO is days from pricing, with the exchange confirming a roughly $623m raise on 31 July and subscriptions opening 10 August, the first humanoid-robotics A-share. Anthropic priced against itself on 24 July, releasing Claude Opus 5 at half what it charges for its own Fable 5, two months after closing a $65B round at a $965B valuation on a run-rate near $47B; it appointed Mariano-Florentino Cuellar as chief global affairs officer on 4 August. OpenAI took GPT-5.6 general on 9 July and launched ChatGPT Work beside it. xAI is now SpaceXAI, and Chinese labs keep closing the gap: Kimi K3's weights are open, and Alibaba says Qwen's Max tier is next. Nous Research, maker of the Hermes agent, is in talks at a $1.5B valuation.
- Anthropic
Maker of Claude
Fable 5 and the Mythos tier up top, Opus 5 (24 July) delivering near-Fable scores at half the price, Sonnet 5 as the default, Claude Code and Cowork carrying the product story. May's round valued it at $965B.
anthropic.com ↗ - OpenAI
Maker of ChatGPT
Took GPT-5.6 general on 9 July 2026 and launched ChatGPT Work beside it, a GPT-5.6 agent that runs for hours and returns finished documents, spreadsheets and web apps.
openai.com ↗ - Google DeepMind
The full-stack player
Gemini 3.1 into 3.5, Nano Banana for images, Veo for video, Spark for agents, and TPUs underneath it all.
deepmind.google ↗ - SpaceXAI (xAI)
Maker of Grok
Rebranded SpaceXAI on 6 July 2026 after February's merger into SpaceX and a June Nasdaq listing; Grok 4.5 shipped two days later on 8 July.
x.ai ↗ - Moonshot AI
Maker of Kimi
Kimi K3, a 2.8-trillion-parameter flagship, launched 16 July 2026 with open weights promised by the 27th, extending Moonshot's run at the front of the open field.
moonshot.ai ↗ - Alibaba · Qwen
China's broadest stack
Ships the open Qwen line, and with Qwen3.8-Max on 3 August 2026 says it will open the Max tier too, reversing the closed, API-only Qwen3.7-Max; China's anthropomorphic-agent rules took effect on 15 July 2026, taking Qwen's human-like companions offline.
qwen.ai ↗ - ByteDance
The media-model power
Dominant in media models: Seedance 2.0 leads video quality boards and Seedream keeps 4K images cheap.
seed.bytedance.com ↗ - DeepSeek
The price-collapse lab
The lab whose open releases forced the whole market to justify its margins; its next-generation signals keep that pressure on.
deepseek.com ↗ - Meta Superintelligence Labs
Meta's frontier lab
Moved Meta's frontier work from open-weight Llama to the proprietary Muse family, and began charging developers for Muse Spark 1.1 through the Meta Model API on 9 July 2026.
ai.meta.com ↗
learning & resourcesThe stack changes monthly; the way to stay current is primary documentation plus live leaderboards, not year-old tutorials.8 entries
◈ Where things stand: Benchmarks now saturate within months of release, so a static score is a snapshot rather than a ranking, and the arenas and usage boards tell you more about what people actually reach for. Every major lab ships genuinely good free documentation and courses, and the MCP specification is required reading for anything agentic.
- Claude Docs & Academy
Anthropic's guides and courses
Model guides, prompting best practice and the Agent SDK reference, plus free certification courses.
docs.claude.com ↗ - OpenAI Cookbook
Runnable API recipes
Hands-on, runnable examples for the API, agents and evals; the fastest way from idea to working call.
cookbook.openai.com ↗ - Google AI Studio
Free Gemini playground
Free Gemini playground with generous limits: prototype prompts, tune settings and export working code.
aistudio.google.com ↗ - Hugging Face
Home of open models
The home of open weights: models, datasets, Spaces demos and the free training courses behind them.
huggingface.co ↗ - MCP Documentation
The connector standard's docs
The spec, SDKs and server gallery for the connector standard every agent platform adopted.
modelcontextprotocol.io ↗ - Artificial Analysis
Independent model benchmarks
Independent, continuously updated comparisons of quality, speed and price across every major model.
artificialanalysis.ai ↗ - LMArena
Human preference leaderboard
Blind head-to-head votes from millions of users; the leaderboard that reflects taste rather than test sets.
arena.ai ↗ - Stanford AI Index
The annual state of AI
The definitive yearly measurement of the field: capability trends, economics and policy in one report.
hai.stanford.edu ↗
Nothing matches that search.