AI · Reference
AI Index
The structured reference behind the AI desk: models, agents, tools, frameworks and companies, searchable, each with a source on the card.
- 64
- entries
- 7
- areas
- 23 Sept 2026
- verified
Search and browse
Every entry in the Index
64 entries · sector notes verified 23 Sept 2026
LLMsLarge language models are the engine room of modern AI: reasoning, writing, coding, search and agent control all run through them. The frontier now competes on long-horizon reliability, tool use and price rather than raw benchmark scores.9
Overview
Where things standAnthropic released Claude Opus 5.5 on 22 September at $4 and $20 per million tokens, 20% below Opus 5, and made it the model its documentation tells developers to start with; the same day OpenAI released GPT-6 Sol at $2 and $10 and GPT-6 Luna at $0.10 and $0.50, half what the GPT-5.6 versions cost. xAI's Grok 4.7 arrived a day earlier at $2 and $6. The frontier text picture sits in the dated log below.
- Claude
Anthropic's frontier family
Opus 5.5 (22 September) is Anthropic's new default: $4 in and $20 out per million, 20% below Opus 5, with cache reads at $0.20, and Anthropic says it matches Fable 5.1 on most work at 40% less than Opus 5 on typical jobs. Fable 5.1 ($10/$50) stays the top tier and shares the top of the Artificial Analysis index with GPT-6 Astra on 53. Sonnet 5 stays the everyday default at $2/$10, and Sonnet 5.5 and Haiku 5.5 follow in the coming weeks.
Our guide →claude.ai ↗ - GPT-6
OpenAI's new generation: Astra, Sol and Luna
GPT-6 Astra began rolling out on 3 September 2026 at $10 in and $50 out per million with a 1,050,000-token context, the first model OpenAI classifies Critical for cybersecurity. GPT-6 Sol ($2/$10) and GPT-6 Luna ($0.10/$0.50) followed on 22 September at half the GPT-5.6 prices, with Luna reaching free ChatGPT users in the desktop app. GPT-5.6 Terra stays on sale at $2/$12.
Our guide →chatgpt.com ↗ - Gemini 3.1 Pro
Google DeepMind's generalist
Gemini 3.1 Pro holds Google's Pro tier while 3.5 Pro stays in partner testing. The Flash seat moved again on 2 September 2026 to Gemini 3.8 Flash, the third Flash in six weeks, at the same introductory $0.75 and $3.75 per million until 31 December; 3.8 Flash Cyber, its vulnerability-hunting sibling, is gated to vetted defenders through the Fairwind Program.
Our guide →gemini.google.com ↗ - Grok 4.7
xAI's flagship line
Grok 4.7, released 21 September 2026, holds the 500K context at $2 in and $6 out per million tokens, the same price as Grok 4.6. A prompt past 200,000 tokens re-prices the whole request at the long-context tier. On xAI's own seven-benchmark launch table it beats Grok 4.6 on every row and takes the best score on two, against four for Fable 5.1 Max at $10 and $50.
Our guide →grok.com ↗ - Kimi K2.6 / K3
Moonshot AI's open flagship
Kimi K3, a 2.8-trillion-parameter flagship, launched on 16 July; its full weights landed on 27 July as the largest open-weight release yet, roughly 1.4TB even at four-bit precision, and the strongest open model on the Artificial Analysis index, under a bespoke licence with a revenue clause for model-as-a-service hosts. Kimi K2.8 Preview (11 September) rolled out across Kimi Code with the model ID unchanged, close to K3 on Moonshot's own description.
Our guide →kimi.com ↗ - Qwen 3.8 Max
Alibaba's flagship line
Alibaba's flagship: 2.4 trillion parameters with 95 billion active, a 1M-token window, $2/$6 per million tokens, launched 3 August 2026. The promised open weights landed on the official Qwen Hugging Face pages around 13 August, in standard and FP8 form, with a 27B sibling alongside.
Our guide →chat.qwen.ai ↗ - GLM-5 series
Z.ai's agent-endurance models
GLM-5.3 (14 August 2026) runs the GLM-5.2 base with heavier post-training, and its full 753B weights landed on Hugging Face on 28 August under a custom glm-5.3 licence rather than MIT. GLM-5.3-Flash (26 August, MIT, 320B with 18B active, 1M context) is the open workhorse, and the model Z.ai had been testing anonymously as ox-alpha.
Our guide →z.ai ↗ - DeepSeek
The price-floor setter
DeepSeek-V4.1-Flash arrived on 10 September 2026 under an MIT licence: 552 billion backbone parameters, 8 billion active reading a prompt and 16 billion writing one, a 1M-token context and native image reading. Its global KV cache holds 890 bytes per token, and output runs $0.6 per million off-peak against $1.2 at peak. V4-Pro-0813 remains on the API at its own rates.
Our guide →deepseek.com ↗ - Muse Spark
Meta's proprietary frontier
Meta Superintelligence Labs' first Muse model (8 April 2026) reasons with parallel agents in Contemplating mode; Muse Spark 1.1 (9 July) added a 1M context, tool and computer use, and a paid API. On 11 August 2026 Meta opened Muse Glimmer under Apache 2.0, dropping the Llama licence rules.
Our guide →ai.meta.com ↗
Image generationImage models have split into two camps: arena-topping generators for ideas, and production systems with real editing, typography and brand control. The right pick depends on whether the image is the destination or a step in a workflow.9
Overview
Where things standOpenAI's Images 2.5 variants hold the top two arena seats: Sunburst and Flare lead both the text-to-image and image-edit boards, with GPT Image 2 third on each and Microsoft's MAI-Image-2.6 fourth on text-to-image (arena.ai boards read 15 September). Images 2.5 adds sketch-guided creation, image comments and more consistent editing, with Flare built for speed and Sunburst for precision.
- GPT Image 2
OpenAI's reasoning image line
OpenAI's image line holds the arena's top three seats, with the Images 2.5 Sunburst and Flare variants first and second and GPT Image 2 third. The first mainstream model family that reasons about a brief, checks references and self-corrects before rendering.
Our guide →chatgpt.com ↗ - Nano Banana Pro
Google's editing powerhouse
Gemini 3 Pro's image engine: native 4K, best-in-class edits and character consistency. Nano Banana 2 covers the fast-and-cheap end of the family.
Our guide →gemini.google.com ↗ - Midjourney V8.2
The art direction standard
Still the strongest pure art direction. V8.2 took the default seat on 24 July with Personalization profiles as the headline upgrade, on top of V8.1's native 2048px output; Niji 7 handles anime styles.
Our guide →midjourney.com ↗ - FLUX.2
Black Forest Labs' open family
The open-weight family to beat: Pro, Flex, Dev and Klein tiers, camera-accurate lighting and up to ten reference images for character consistency.
Our guide →bfl.ai ↗ - Ideogram
Text rendering specialist
The reliable choice when the image must carry words: posters, logos, packaging and layout-aware design work.
Our guide →ideogram.ai ↗ - Recraft V4
Design-first generation
Built for designers rather than prompters: vector output, exact brand colours and print-ready renders straight from the model.
Our guide →recraft.ai ↗ - Seedream
ByteDance's value line
ByteDance's image line delivers 4K output at some of the lowest per-image prices on the market, with strong bilingual text rendering.
Our guide →seed.bytedance.com ↗ - MAI-Image-2.6
Microsoft AI's image model
Microsoft AI's in-house image model took second place on the arena text-to-image board at its 10 August launch, ahead of Google, Meta and xAI, with text rendering up 91 Elo on Microsoft's own figures and a focus on product, branding and photoreal work.
microsoft.ai ↗ - Firefly Image 5
Adobe's licensed-data model
Trained on licensed data with indemnification for enterprise use, and now a hub that also fronts partner models inside Creative Cloud.
Our guide →adobe.com ↗
Video generationAI video crossed from novelty to production this year: native audio, multi-shot control and 4K output are now table stakes at the top of the market.11
Overview
Where things standNVIDIA’s Sol-H3 research reports a 1.653-second warm run for about five seconds of video with stereo audio on eight B300 GPUs. The four-step result is a hardware-specific inference benchmark, with model loading and final MP4 encoding outside the timing.
- Seedance 2.0 / 2.5
ByteDance's board contender
Seedance 2.5 sits second on the video-editing board, four points behind Wan 3.0, and fourth on image-to-video. One model generates picture and sound together in 30-second single passes, so lips, foley and music actually match the frame.
Our guide →seed.bytedance.com ↗ - Veo 3.1
Google DeepMind's flagship
Native 48kHz dialogue, 4K output and SynthID provenance marks. The default choice where brand safety and clean licensing decide the choice.
Our guide →labs.google ↗ - Kling 3.0
Kuaishou's value flagship
Native 4K at 60fps, 15-second multi-shot storyboards and lip-sync in five languages, at roughly $0.10 a second.
Our guide →klingai.com ↗ - Sora 2
OpenAI · being retired
Being retired: app and web access closed on 26 April and the API ends on 24 September 2026. Plan migrations now rather than building on it.
Our guide →openai.com ↗ - Runway Gen-4.5
The filmmaker's control surface
No longer top of the leaderboards but still the best control surface in the business: motion brushes, keyframes and a real editing timeline.
Our guide →runwayml.com ↗ - Ray3 · Dream Machine
Luma's HDR pioneer
First native 16-bit HDR video model, with Ray3 Modify for video-to-video work inside the Dream Machine 2.0 studio.
Our guide →lumalabs.ai ↗ - Wan 2.7 / 3.0
Alibaba's video family
Wan 3.0 (24 August 2026) generates 30-second single-pass clips, takes documents, slides and web pages as input, and tops the video-editing arena; its weights are not yet published. Wan 2.7 remains the strongest fully open video family, Apache-licensed with first-and-last-frame control.
Our guide →wan.video ↗ - LTX-2.3
Lightricks' open 4K model
A 22B open model producing native 4K at 50fps with stereo audio; free for commercial use below a revenue threshold.
Our guide →lightricks.com ↗ - Hailuo · MiniMax-H3
MiniMax's motion specialist
MiniMax's line is prized for expressive, physical motion: the model animators reach for when movement has to feel alive. Its H3 model now tops the image-to-video arena.
Our guide →hailuoai.video ↗ - FLUX 3 Video
Black Forest Labs' video debut
Black Forest Labs' first video family, rolled out in early August, sits level with Wan 3.0 on the text-to-video arena behind Google's two Omni Flash models, and in the top seven for image-to-video within weeks of release.
Our guide →bfl.ai ↗ - Gemini Omni 1.1 Flash
conversational video generation
Google's conversational video line. Omni 1.1 Flash (27 August 2026) adds end frames, scene extension to a cumulative 40 seconds, a 360p draft mode at a third of the cost, and 1080p or 4K export; it tops the text-to-video arena and sits second on image-to-video.
Our guide →deepmind.google ↗
Agentic frameworks / routingAgent frameworks sit between models and outcomes: routing, tools, memory and the guardrails that decide whether an agent is useful or dangerous. Reliability and control now count for more than raw capability.10
Overview
Where things standAnthropic retired the Cowork name on 16 September: 'Claude Cowork and chat are now one Claude', with the file-and-app agent folded into ordinary chat and Claude Docs and Slides in beta. Sakana's Fugu Max and Fugu Ultra v2 (11 September) route tasks across open-weight and specialist models, Fugu Max priced at $2/$6 per million tokens, and Grok Build added passive project memory on the 16th.
- OpenClaw
Self-hosted agent gateway
The breakout self-hosted gateway: one daemon, twenty-plus messaging channels, skills and scheduled jobs, any model behind it. Since April it needs API billing rather than a Claude subscription, and it should be hardened before exposure.
Our guide →openclaw.ai ↗ - Hermes Agent
Self-improving open agent
Nous Research's open-source, self-hosted agent, and the one that changes as you use it: finishing a task writes a reusable skill back into the agent. Version 0.20.6 (27 August) added consent-gated real-profile browsing from a managed copy of your Chrome profile, off by default; the Herald release (3 August) added real-time voice and an agent-to-agent protocol.
Our guide →hermes-agent.nousresearch.com ↗ - Claude Cowork
Merged into Claude
Folded into one Claude on 16 September 2026, rollout starting with Pro and Max: the file-and-app agentic work Cowork carried now runs in ordinary Claude chat, with Claude Docs and Slides launched in beta alongside. Its original launch is what markets nicknamed the SaaSpocalypse.
claude.com ↗ - Gemini Spark
Google's always-on agent
Google's always-on agent living in a cloud VM: deep Gmail, Docs and Calendar access, MCP connectors from day one, and it keeps working while your laptop sleeps.
gemini.google.com ↗ - ChatGPT Agent
OpenAI's consumer agent
Agent mode inside ChatGPT with its own virtual computer for browsing, files and multi-step tasks under supervision.
chatgpt.com ↗ - Claude Agent SDK
Anthropic's builder stack
The loop behind Claude Code, embeddable in your own product, alongside a managed runtime (beta) where Anthropic hosts the agent for you.
docs.claude.com ↗ - OpenAI Agents SDK
Handoffs and guardrails
OpenAI's production framework with explicit handoffs, guardrails and tracing; MCP support landed early this year.
openai.github.io ↗ - Model Context Protocol
The connector standard
The USB-C of agent tooling: one protocol that lets any client use any tool or data server. Every major platform now speaks it.
Our guide →modelcontextprotocol.io ↗ - OpenRouter
One API, every model
One API over every major model with live rankings drawn from real usage: the quickest way to A/B models behind an agent without rewrites. Its Ori CLI (4 August 2026) tunes Claude Code, Codex, OpenCode and Hermes to run through the gateway. Stripe agreed to acquire it on 19 August; name, product and roadmap stay.
openrouter.ai ↗ - ChatGPT Work
OpenAI's finished-work agent
OpenAI's agent for long-running jobs (9 July 2026): powered by GPT-5.6, it works for hours and returns finished documents, spreadsheets, slides and web apps.
openai.com ↗
AI coding & dev toolsAI coding assistants went from autocomplete to delegation: the leading tools now plan, edit across a repo, run tests and open pull requests on their own.8
Overview
Where things standMicrosoft’s 8 September VS Code advisories make agent permissions the immediate update priority. Version 1.136.2 fixes two optional network-filter bypasses and a workspace-configuration flaw involving remote agent hosts. Coder’s separate registry cleanup guidance still applies.
- Claude Code
Anthropic's terminal agent
Terminal-native and repo-aware, with Dynamic Workflows for long unattended runs; Opus 5 became its default on 24 July, with Fable 5 the tier above it. Included with a Claude Pro plan.
Our guide →claude.com ↗ - Cursor
SpaceX's AI-native editor
The AI-native editor, a SpaceX company since 14 August 2026: Tab completions plus Composer cloud agents that work branches in parallel VMs, running Claude, Gemini and Grok, with OpenAI models due to leave on 12 November unless the wind-down is resolved.
Our guide →cursor.com ↗ - GitHub Copilot
The installed-base leader
Still the default inside VS Code and enterprises, with a free tier, IP indemnity and new flexible usage billing from 1 June.
github.com ↗ - Codex
OpenAI's cloud engineer
OpenAI's asynchronous cloud engineer: give it an issue and it returns a tested pull request. Now powered by the coding-tuned GPT-5.6 Sol.
Our guide →openai.com ↗ - Antigravity 2.0
Google's agent-first IDE
Google's agent-first IDE built around the Gemini 3.5 line, with a browser-testing loop and a generous free tier via a Google account.
antigravity.google ↗ - Devin Desktop
Windsurf's successor
Windsurf's next chapter after the Cognition deal: the Devin autonomous engineer wrapped in the familiar editor, relaunched in June.
Our guide →cognition.ai ↗ - Kiro
AWS's spec-driven agent IDE
Amazon's take: spec-first development where agents implement against written requirements, aimed at teams that want process over improvisation.
kiro.dev ↗ - Replit Agent
Prompt to deployed app
From prompt to deployed app in one place: environment, database and hosting handled for you. The fastest route for non-specialists.
replit.com ↗
AI companiesThe market is now a platform race: strong models are table stakes, and distribution, chips, agents and developer ecosystems decide who compounds.9
Overview
Where things standAnthropic's R&D Automation Index, published 17 September, reports Claude leading 26% of the company's own research and development at its AL4 tier, up from under 1% in February, with about 30,000 agents at work at any time in August; the figures are Anthropic's own. The dated funding and acquisition announcements below distinguish signed deals from completed transactions.
- Anthropic
Maker of Claude
Fable 5 and the Mythos tier up top, Opus 5 (24 July) delivering near-Fable scores at half the price, Sonnet 5 as the default, Claude Code and, since the 16 September merge, one Claude carrying the product story. A Life Sciences Verification Program opened on 17 September, swapping real-time blocking for usage monitoring at vetted organisations. May's round valued it at $965B.
anthropic.com ↗ - OpenAI
Maker of ChatGPT
Took GPT-5.6 general on 9 July 2026 and launched ChatGPT Work beside it, a GPT-5.6 agent that runs for hours and returns finished documents, spreadsheets and web apps. Astra for Law followed on 17 September 2026: GPT-6 Astra over a 230-million-URL case-law index, with Sullivan & Cromwell, Ropes & Gray and Cooley named as early builders.
openai.com ↗ - Google DeepMind
The full-stack player
Gemini 3.1 into 3.5, Nano Banana for images, Veo and Omni Flash for video, Spark for agents, and TPUs underneath it all. Since early August, Demis Hassabis chairs while Koray Kavukcuoglu leads day to day.
deepmind.google ↗ - SpaceXAI (xAI)
Maker of Grok
Rebranded SpaceXAI on 6 July 2026 after February's merger into SpaceX and a June Nasdaq listing. SpaceX's $60bn all-stock purchase of Anysphere, maker of Cursor, closed on 14 August, and NVIDIA says Grok's next agentic workloads will run on its Vera Rubin platform.
x.ai ↗ - Moonshot AI
Maker of Kimi
Kimi K3, a 2.8-trillion-parameter flagship, launched 16 July 2026 with open weights promised by the 27th, extending Moonshot's run at the front of the open field.
moonshot.ai ↗ - Alibaba · Qwen
China's broadest stack
Ships the open Qwen line, and delivered on the Max-tier promise: Qwen3.8-Max's 2.4-trillion-parameter weights landed on Hugging Face on 8 August 2026. QwenWork, its all-in-one agent workplace, opened an international beta on 26 August; China's anthropomorphic-agent rules took Qwen's human-like companions offline in July.
qwen.ai ↗ - ByteDance
The media-model power
Dominant in media models: Seedance 2.0 leads video quality boards and Seedream keeps 4K images cheap.
seed.bytedance.com ↗ - DeepSeek
The price-collapse lab
The lab whose open releases forced the whole market to justify its margins; its next-generation signals keep that pressure on.
deepseek.com ↗ - Meta Superintelligence Labs
Meta's frontier lab
Moved Meta's frontier work from open-weight Llama to the proprietary Muse family, and began charging developers for Muse Spark 1.1 through the Meta Model API on 9 July 2026.
ai.meta.com ↗
Learning & resourcesThe stack changes monthly; the way to stay current is primary documentation plus live leaderboards, not year-old tutorials.8
Overview
Where things standYFarmX’s reproducible randomness study shows why generated answers should not substitute for a random-number generator. Across twelve models, number choices and coin-flip sequences showed strong preferences. The archive includes raw responses, scripts and parsing limitations.
- Claude Docs & Academy
Anthropic's guides and courses
Model guides, prompting best practice and the Agent SDK reference, plus free certification courses.
docs.claude.com ↗ - OpenAI Cookbook
Runnable API recipes
Hands-on, runnable examples for the API, agents and evals; the fastest way from idea to working call.
cookbook.openai.com ↗ - Google AI Studio
Free Gemini playground
Free Gemini playground with generous limits: prototype prompts, tune settings and export working code.
aistudio.google.com ↗ - Hugging Face
Home of open models
The home of open weights: models, datasets, Spaces demos and the free training courses behind them.
huggingface.co ↗ - MCP Documentation
The connector standard's docs
The spec, SDKs and server gallery for the connector standard every agent platform adopted.
Our guide →modelcontextprotocol.io ↗ - Artificial Analysis
Independent model benchmarks
Independent, continuously updated comparisons of quality, speed and price across every major model.
artificialanalysis.ai ↗ - LMArena
Human preference leaderboard
Blind head-to-head votes from millions of users; the leaderboard that reflects taste rather than test sets.
arena.ai ↗ - Stanford AI Index
The annual state of AI
The definitive yearly measurement of the field: capability trends, economics and policy in one report.
hai.stanford.edu ↗
Nothing matches that search.



