The AI Index
AI Knowledge Hub
57 entries · sector notes current as of 6 Jul 2026
llmsLarge language models are the engine room of modern AI: reasoning, writing, coding, search and agent control all run through them. The frontier now competes on long-horizon reliability, tool use and price rather than raw benchmark scores.8 entries
◈ Where things stand: Anthropic opened a new tier above Opus with Claude Fable 5 in June, then made Sonnet 5 its default on 30 June with a 1M-token context. OpenAI's GPT-5.5 leads several agentic benchmarks while the GPT-5.6 preview sits behind a vetted access list. Google's Gemini 3.1 Pro heads multimodal work with 3.5 cleared for release this month, and open-weight labs now land within striking distance at a fraction of the price.
- Claude
Anthropic's frontier family
Fable 5 tops the hardest reasoning and coding evals; Sonnet 5 (30 June) is the new everyday default with a 1M-token window. Opus 4.8 remains the long-horizon agentic workhorse.
claude.ai ↗ - GPT-5.5
OpenAI's flagship line
OpenAI's flagship leads terminal and tool-use benchmarks, with GPT-5.5 Pro strongest on hard maths. The GPT-5.6 family is in a gated preview limited to vetted organisations.
chatgpt.com ↗ - Gemini 3.1 Pro
Google DeepMind's generalist
The strongest multimodal generalist, with a 1M-token context and deep Workspace hooks. Gemini 3.5 Pro has been cleared for general release this month.
gemini.google.com ↗ - Grok 4.3
xAI's real-time model
Sharp reasoning with live X data and the most permissive defaults of the big labs. API pricing undercuts every Western rival at $1.25 in and $2.50 out per million tokens.
grok.com ↗ - Kimi K2.6
Moonshot AI's open flagship
The open-weight model many agent stacks now run first: strong coding, long context and near-frontier scores at a fraction of closed-model prices.
kimi.com ↗ - Qwen 3.7 Max
Alibaba's flagship line
Frontier-grade maths and coding from Alibaba's flagship line. The 3.7 tier is closed-weight; self-hosters get the earlier 3.6 openly.
chat.qwen.ai ↗ - GLM-5 series
Z.ai's agent-endurance models
Purpose-built for long autonomous runs: GLM-5.1 tops OpenRouter's agent-endurance boards with multi-hour unattended sessions.
z.ai ↗ - DeepSeek
The price-floor setter
The lab that made frontier reasoning cheap. Its open releases keep resetting the price floor the rest of the market has to answer.
deepseek.com ↗
image generationImage models have split into two camps: arena-topping generators for ideas, and production systems with real editing, typography and brand control. The right pick depends on whether the image is the destination or a step in a workflow.8 entries
◈ Where things stand: OpenAI's GPT Image 2 (April) leads the arena by a record margin and plans or web-searches before it draws. Google's Nano Banana Pro owns precise editing and native 4K, Midjourney moved to V8.1 as its default in June, and FLUX.2 remains the strongest open-weight family. Adobe's Firefly Image 5 is the pick where licensing questions decide the tool.
- GPT Image 2
OpenAI's reasoning image model
Top of the image arena by the largest lead recorded, and the first mainstream model that reasons about a brief, checks references and self-corrects before rendering.
chatgpt.com ↗ - Nano Banana Pro
Google's editing powerhouse
Gemini 3 Pro's image engine: native 4K, best-in-class edits and character consistency. Nano Banana 2 covers the fast-and-cheap end of the family.
gemini.google.com ↗ - Midjourney V8.1
The art direction standard
Still the strongest pure art direction. V8.1 became the default in June with native 2048px output and a big speed jump; Niji 7 handles anime styles.
midjourney.com ↗ - FLUX.2
Black Forest Labs' open family
The open-weight family to beat: Pro, Flex, Dev and Klein tiers, camera-accurate lighting and up to ten reference images for character consistency.
bfl.ai ↗ - Ideogram
Text rendering specialist
The reliable choice when the image must carry words: posters, logos, packaging and layout-aware design work.
ideogram.ai ↗ - Recraft V4
Design-first generation
Built for designers rather than prompters: vector output, exact brand colours and print-ready renders straight from the model.
recraft.ai ↗ - Seedream
ByteDance's value line
ByteDance's image line delivers 4K output at some of the lowest per-image prices on the market, with strong bilingual text rendering.
seed.bytedance.com ↗ - Firefly Image 5
Adobe's licensed-data model
Trained on licensed data with indemnification for enterprise use, and now a hub that also fronts partner models inside Creative Cloud.
adobe.com ↗
video generationAI video crossed from novelty to production this year: native audio, multi-shot control and 4K output are now table stakes at the top of the market.9 entries
◈ Where things stand: ByteDance's Seedance 2.0 leads the independent quality boards with a unified audio-video architecture, with Kling 3.0 close behind on value at native 4K and 60fps. Veo 3.1 remains the safest Western pick with native dialogue and SynthID watermarking. The big caution: OpenAI has retired Sora 2, with the API switching off on 24 September 2026. Open weights are genuinely usable now.
- Seedance 2.0
ByteDance's board leader
Top of the independent quality boards. One model generates picture and sound together, so lips, foley and music actually match the frame.
seed.bytedance.com ↗ - Veo 3.1
Google DeepMind's flagship
Native 48kHz dialogue, 4K output and SynthID provenance marks. The default choice where brand safety and clean licensing matter.
labs.google ↗ - Kling 3.0
Kuaishou's value flagship
Native 4K at 60fps, 15-second multi-shot storyboards and lip-sync in five languages, at roughly $0.10 a second.
klingai.com ↗ - Sora 2
OpenAI · being retired
Being retired: app and web access closed on 26 April and the API ends on 24 September 2026. Plan migrations now rather than building on it.
openai.com ↗ - Runway Gen-4.5
The filmmaker's control surface
No longer top of the leaderboards but still the best control surface in the business: motion brushes, keyframes and a real editing timeline.
runwayml.com ↗ - Ray3 · Dream Machine
Luma's HDR pioneer
First native 16-bit HDR video model, with Ray3 Modify for video-to-video work inside the Dream Machine 2.0 studio.
lumalabs.ai ↗ - Wan 2.7
Alibaba's open video family
The strongest fully open video family, Apache-licensed with first-and-last-frame control for planning shots precisely.
wan.video ↗ - LTX-2.3
Lightricks' open 4K model
A 22B open model producing native 4K at 50fps with stereo audio; free for commercial use below a revenue threshold.
lightricks.com ↗ - Hailuo 2.3
MiniMax's motion specialist
MiniMax's line is prized for expressive, physical motion: the model animators reach for when movement has to feel alive.
hailuoai.video ↗
agentic frameworks / routingAgent frameworks sit between models and outcomes: routing, tools, memory and the guardrails that decide whether an agent is useful or dangerous. Reliability and control now count for more than raw capability.8 entries
◈ Where things stand: OpenClaw became the fastest-growing open-source project in GitHub history and now anchors a whole ecosystem, including NVIDIA's NemoClaw build. Anthropic's Cowork leads desktop agents, Google answered in May with Gemini Spark running around the clock on cloud VMs, and MCP has settled in as the universal tool standard. Treat the space as powerful but young: a critical OpenClaw vulnerability in March and malicious community skills are the warning labels.
- OpenClaw
Self-hosted agent gateway
The breakout self-hosted gateway: one daemon, twenty-plus messaging channels, skills and scheduled jobs, any model behind it. Since April it needs API billing rather than a Claude subscription, and it should be hardened before exposure.
openclaw.ai ↗ - Claude Cowork
Anthropic's desktop agent
Anthropic's desktop agent for real file-and-app work, with domain packs for legal and operations tasks. Its launch is what markets nicknamed the SaaSpocalypse.
claude.com ↗ - Gemini Spark
Google's always-on agent
Google's always-on agent living in a cloud VM: deep Gmail, Docs and Calendar access, MCP connectors from day one, and it keeps working while your laptop sleeps.
gemini.google.com ↗ - ChatGPT Agent
OpenAI's consumer agent
Agent mode inside ChatGPT with its own virtual computer for browsing, files and multi-step tasks under supervision.
chatgpt.com ↗ - Claude Agent SDK
Anthropic's builder stack
The loop behind Claude Code, embeddable in your own product, alongside a managed runtime (beta) where Anthropic hosts the agent for you.
docs.claude.com ↗ - OpenAI Agents SDK
Handoffs and guardrails
OpenAI's production framework with explicit handoffs, guardrails and tracing; MCP support landed early this year.
openai.github.io ↗ - Model Context Protocol
The connector standard
The USB-C of agent tooling: one protocol that lets any client use any tool or data server. Every major platform now speaks it.
modelcontextprotocol.io ↗ - OpenRouter
One API, every model
One API over every major model with live rankings drawn from real usage: the quickest way to A/B models behind an agent without rewrites.
openrouter.ai ↗
ai coding & dev toolsAI coding assistants went from autocomplete to delegation: the leading tools now plan, edit across a repo, run tests and open pull requests on their own.8 entries
◈ Where things stand: Copilot keeps the biggest installed base but growth has stalled; Cursor and Claude Code are tied at roughly 18% adoption each and rising. Claude Code reached a reported $2.5B revenue run-rate inside nine months, the fastest developer product on record, while Cursor's Composer added parallel cloud agents. Codex now runs on the GPT-5.6 Sol family, and Windsurf re-emerged in June as Devin Desktop under Cognition.
- Claude Code
Anthropic's terminal agent
Terminal-native and repo-aware, with Dynamic Workflows on Opus 4.8 for long unattended runs. Included with a Claude Pro plan.
claude.com ↗ - Cursor
Anysphere's AI-native editor
The AI-native editor: Tab completions plus Composer cloud agents that work branches in parallel VMs. Model-agnostic across Claude, GPT, Gemini and Grok.
cursor.com ↗ - GitHub Copilot
The installed-base leader
Still the default inside VS Code and enterprises, with a free tier, IP indemnity and new flexible usage billing from 1 June.
github.com ↗ - Codex
OpenAI's cloud engineer
OpenAI's asynchronous cloud engineer: give it an issue and it returns a tested pull request. Now powered by the coding-tuned GPT-5.6 Sol.
openai.com ↗ - Antigravity 2.0
Google's agent-first IDE
Google's agent-first IDE built around the Gemini 3.5 line, with a browser-testing loop and a generous free tier via a Google account.
antigravity.google ↗ - Devin Desktop
Windsurf's successor
Windsurf's next chapter after the Cognition deal: the Devin autonomous engineer wrapped in the familiar editor, relaunched in June.
cognition.ai ↗ - Kiro
AWS's spec-driven agent IDE
Amazon's take: spec-first development where agents implement against written requirements, aimed at teams that want process over improvisation.
kiro.dev ↗ - Replit Agent
Prompt to deployed app
From prompt to deployed app in one place: environment, database and hosting handled for you. The fastest route for non-specialists.
replit.com ↗
ai companiesThe market is now a platform race: models matter, but distribution, chips, agents and developer ecosystems decide who compounds.8 entries
◈ Where things stand: Anthropic raised at a $965B valuation in May on a reported run-rate near $47B and opened the Mythos tier. OpenAI splits its line between the public GPT-5.5 and a vetted GPT-5.6 preview. Google is shipping across every layer at once, and Chinese labs closed most of the gap on open weights even as new rules push Doubao and Qwen to switch off human-like companion agents by 15 July.
- Anthropic
Maker of Claude
Fable 5 and the Mythos tier up top, Sonnet 5 as the default, Claude Code and Cowork carrying the product story. May's round valued it at $965B.
anthropic.com ↗ - OpenAI
Maker of ChatGPT
ChatGPT remains the biggest consumer surface; GPT-5.5 leads public benchmarks while 5.6 trials a vetted-access model. Sora's retirement marks a sharper focus on core models and agents.
openai.com ↗ - Google DeepMind
The full-stack player
Gemini 3.1 into 3.5, Nano Banana for images, Veo for video, Spark for agents, and TPUs underneath it all.
deepmind.google ↗ - xAI
Maker of Grok
Grok 4.3 pairs frontier reasoning with the lowest big-lab API prices and live X data; Grok 4.5 entered private beta in late June.
x.ai ↗ - Moonshot AI
Maker of Kimi
The open-weight standard-bearer: Kimi K2.6 is the model agent builders benchmark everything else against on price for performance.
moonshot.ai ↗ - Alibaba · Qwen
China's broadest stack
Qwen 3.7 Max tops maths boards and the wider family spans text, vision and video, even as new agent rules bite at home.
qwen.ai ↗ - ByteDance
The media-model power
Quietly dominant in media models: Seedance 2.0 leads video quality boards and Seedream keeps 4K images cheap.
seed.bytedance.com ↗ - DeepSeek
The price-collapse lab
The lab whose open releases forced the whole market to justify its margins; its next-generation signals keep that pressure on.
deepseek.com ↗
learning & resourcesThe stack changes monthly; the way to stay current is primary documentation plus live leaderboards, not year-old tutorials.8 entries
◈ Where things stand: Benchmarks now saturate within months of release, so treat static scores as snapshots: the arenas and usage boards tell you more. Every major lab ships genuinely good documentation and free courses, and the MCP docs are required reading for anything agentic.
- Claude Docs & Academy
Anthropic's guides and courses
Model guides, prompting best practice and the Agent SDK reference, plus free certification courses.
docs.claude.com ↗ - OpenAI Cookbook
Runnable API recipes
Hands-on, runnable examples for the API, agents and evals; the fastest way from idea to working call.
cookbook.openai.com ↗ - Google AI Studio
Free Gemini playground
Free Gemini playground with generous limits: prototype prompts, tune settings and export working code.
aistudio.google.com ↗ - Hugging Face
Home of open models
The home of open weights: models, datasets, Spaces demos and the free training courses behind them.
huggingface.co ↗ - MCP Documentation
The connector standard's docs
The spec, SDKs and server gallery for the connector standard every agent platform adopted.
modelcontextprotocol.io ↗ - Artificial Analysis
Independent model benchmarks
Independent, continuously updated comparisons of quality, speed and price across every major model.
artificialanalysis.ai ↗ - LMArena
Human preference leaderboard
Blind head-to-head votes from millions of users; the leaderboard that reflects taste rather than test sets.
lmarena.ai ↗ - Stanford AI Index
The annual state of AI
The definitive yearly measurement of the field: capability trends, economics and policy in one report.
hai.stanford.edu ↗
Nothing matches that search.