AI · Reference

AI Index

The structured reference behind the AI desk: models, agents, tools, frameworks and companies, searchable, each with a source on the card.

64
entries
7
areas
30 Aug 2026
verified

Search and browse

Every entry in the Index

64 entries · sector notes verified 30 Aug 2026

LLMsLarge language models are the engine room of modern AI: reasoning, writing, coding, search and agent control all run through them. The frontier now competes on long-horizon reliability, tool use and price rather than raw benchmark scores.9

Where things standZ.ai closed August by opening up. The full GLM-5.3 weights, 753 billion parameters across 756GB, landed on Hugging Face on 28 August under a custom glm-5.3 licence rather than the MIT of the 5.2 line, ending the fortnight safety hold. Two days earlier it had launched GLM-5.3-Flash, a 320B open MoE with 18B active, MIT-licensed, multimodal and 1M-context, and confirmed it was the model it had been testing anonymously as ox-alpha since 20 August, the stealth model YFarmX fingerprinted the week before. Two more open-weight releases arrived in the same week. Tencent published its new flagship, Hy4 preview, on 27 August under Apache 2.0: 770B total parameters with 49B active per token on Tencent's own figures, 78 layers and a 1M context. Alibaba opened Qwen3.8-Flash-Next on 26 August, 125B with about 6B active at $0.15 in and $0.47 out, which Qwen calls an experimental preview of the architecture Qwen4 will be built on. OpenAI moved on price on 21 August: GPT-5.6 Sol's API rate fell to $4 in and $20 out per million, a promotional cut running to at least 21 November, with Terra and Luna unchanged. Behind that, the mid-August wave stands: Grok 4.6 (12 August), Gemini 3.7 Flash at an introductory $0.75 and $3.75 that doubles on 1 January 2027 (13 August), DeepSeek V4 Pro under MIT with peak and off-peak rates up to 4.5 times the old flat figure, though weekends run entirely off-peak, and GLM-5.3 itself (14 August). The frame holds: Claude Opus 5 (24 July, $5 and $25) leads the Artificial Analysis index with Fable 5 level behind it and first on the Arena text board, Kimi K3's 1.4TB of open weights remain the strongest open model anyone can download, Qwen3.8-Max's promised weights are live on Hugging Face, and Sonnet 5's $2 and $10 is permanent.

  • Claude

    Anthropic's frontier family

    Opus 5 (24 July) leads the Artificial Analysis index at the same $5/$25 as the Opus 4.8 it retires, half Fable 5's price; Fable 5 sits level behind it and tops the Arena text board. Sonnet 5 stays the everyday default with a 1M-token window at $2/$10, pricing made permanent on 10 August.

    Frontier1M contextOur guide →claude.ai ↗
  • GPT-5.6

    OpenAI's flagship tiers

    GPT-5.6 went general on 9 July 2026 in three tiers, Sol, Terra and Luna, after a brief government-brokered preview. On 30 July OpenAI cut Luna by 80% to $0.20 and $1.20 per million tokens and Terra by 20% to $2 and $12. Sol followed on 21 August: $4 and $20 per million on a promotion running to at least 21 November, from $5 and $30. GPT-5.5 Instant stays the everyday default inside ChatGPT.

    FrontierOur guide →chatgpt.com ↗
  • Gemini 3.1 Pro

    Google DeepMind's generalist

    Gemini 3.1 Pro holds Google's Pro tier while 3.5 Pro stays in partner testing. The Flash seat moved to Gemini 3.7 Flash on 13 August 2026: the fastest model Artificial Analysis tracks, at an introductory $0.75 and $3.75 per million tokens that doubles on 1 January 2027.

    Multimodal1M contextOur guide →gemini.google.com ↗
  • Grok 4.6

    xAI's flagship line

    Grok 4.6, released 12 August 2026, holds the 500K context at $2 in and $6 out per million tokens. Cached input rose from $0.30 to $0.50, and a prompt past 200,000 tokens re-prices the whole request at the long-context tier. On xAI's own ten-benchmark launch table it beats Grok 4.5 everywhere and wins three rows against the field.

    FrontierReal-timeOur guide →grok.com ↗
  • Kimi K2.6 / K3

    Moonshot AI's open flagship

    Kimi K3, a 2.8-trillion-parameter flagship, launched on 16 July; its full weights landed on 27 July as the largest open-weight release yet, roughly 1.4TB even at four-bit precision, and the strongest open model on the Artificial Analysis index, under a bespoke licence with a revenue clause for model-as-a-service hosts.

    Open weightsOur guide →kimi.com ↗
  • Qwen 3.8 Max

    Alibaba's flagship line

    Alibaba's flagship: 2.4 trillion parameters with 95 billion active, a 1M-token window, $2/$6 per million tokens, launched 3 August 2026. The promised open weights landed on the official Qwen Hugging Face pages around 13 August, in standard and FP8 form, with a 27B sibling alongside.

    FrontierOpen weightsOur guide →chat.qwen.ai ↗
  • GLM-5 series

    Z.ai's agent-endurance models

    GLM-5.3 (14 August 2026) runs the GLM-5.2 base with heavier post-training, and its full 753B weights landed on Hugging Face on 28 August under a custom glm-5.3 licence rather than MIT. GLM-5.3-Flash (26 August, MIT, 320B with 18B active, 1M context) is the open workhorse, and the model Z.ai had been testing anonymously as ox-alpha.

    Open weightsAgentsOur guide →z.ai ↗
  • DeepSeek

    The price-floor setter

    DeepSeek-V4-Pro-0813 went general on 13 August 2026 under an MIT licence with a 1M-token context and big agentic gains on DeepSeek's own harness. The flat pricing ended 16 August: output runs $1.98 off-peak and $3.96 at peak, from a flat $0.87, with the peak hours applying on weekdays only.

    Open weightsValueOur guide →deepseek.com ↗
  • Muse Spark

    Meta's proprietary frontier

    Meta Superintelligence Labs' first Muse model (8 April 2026) reasons with parallel agents in Contemplating mode; Muse Spark 1.1 (9 July) added a 1M context, tool and computer use, and a paid API. On 11 August 2026 Meta opened Muse Glimmer under Apache 2.0, dropping the Llama licence rules.

    MultimodalAgentsOur guide →ai.meta.com ↗
Image generationImage models have split into two camps: arena-topping generators for ideas, and production systems with real editing, typography and brand control. The right pick depends on whether the image is the destination or a step in a workflow.9

Where things standChecked on 30 August, OpenAI's GPT Image 2 still leads the arena, and it plans or searches the web before it draws rather than guessing at what you asked for, with its lead back out to 51 points over the chaser: Microsoft's MAI-Image-2.6, launched 10 August, holds second, with Grok Imagine Image 2.0 third, Reve 2.1 fourth and Meta's Muse Image fifth, and Google's Nano Banana 2 down in seventh. Muse Image itself reached Meta's own developer API and OpenRouter on 26 August at a cent an image, priced for production volume. Nano Banana Pro still owns precise editing and native 4K, with Nano Banana 2 Lite (30 June) the fast tier at roughly $0.034 an image. Midjourney's V8.2 shipped on 24 July and took the default seat from V8.1, with its Personalization profiles the headline upgrade, FLUX.2 remains the strongest open-weight family, and Adobe's Firefly Image 5 is the pick wherever licensing decides the tool.

  • GPT Image 2

    OpenAI's reasoning image model

    Top of the image arena by the largest lead recorded, and the first mainstream model that reasons about a brief, checks references and self-corrects before rendering.

    Arena #1Our guide →chatgpt.com ↗
  • Nano Banana Pro

    Google's editing powerhouse

    Gemini 3 Pro's image engine: native 4K, best-in-class edits and character consistency. Nano Banana 2 covers the fast-and-cheap end of the family.

    4KEditingOur guide →gemini.google.com ↗
  • Midjourney V8.2

    The art direction standard

    Still the strongest pure art direction. V8.2 took the default seat on 24 July with Personalization profiles as the headline upgrade, on top of V8.1's native 2048px output; Niji 7 handles anime styles.

    AestheticsOur guide →midjourney.com ↗
  • FLUX.2

    Black Forest Labs' open family

    The open-weight family to beat: Pro, Flex, Dev and Klein tiers, camera-accurate lighting and up to ten reference images for character consistency.

    Open weightsOur guide →bfl.ai ↗
  • Ideogram

    Text rendering specialist

    The reliable choice when the image must carry words: posters, logos, packaging and layout-aware design work.

    TypographyOur guide →ideogram.ai ↗
  • Recraft V4

    Design-first generation

    Built for designers rather than prompters: vector output, exact brand colours and print-ready renders straight from the model.

    DesignVectorsOur guide →recraft.ai ↗
  • Seedream

    ByteDance's value line

    ByteDance's image line delivers 4K output at some of the lowest per-image prices on the market, with strong bilingual text rendering.

    Value4KOur guide →seed.bytedance.com ↗
  • MAI-Image-2.6

    Microsoft AI's image model

    Microsoft AI's in-house image model took second place on the arena text-to-image board at its 10 August launch, ahead of Google, Meta and xAI, with text rendering up 91 Elo on Microsoft's own figures and a focus on product, branding and photoreal work.

    Arena #2microsoft.ai ↗
  • Firefly Image 5

    Adobe's licensed-data model

    Trained on licensed data with indemnification for enterprise use, and now a hub that also fronts partner models inside Creative Cloud.

    IP-safeOur guide →adobe.com ↗
Video generationAI video crossed from novelty to production this year: native audio, multi-shot control and 4K output are now table stakes at the top of the market.11

Where things standThe board leadership changed twice in a week. Google shipped Gemini Omni 1.1 Flash on 27 August, adding end frames, scene extension to a cumulative 40 seconds and a cheaper 360p draft mode, and by the 29 August check it topped text-to-video with its predecessor second and Black Forest Labs' FLUX 3 Video pushed down to third. Alibaba's Wan 3.0, released wide on 24 August with 30-second single-pass clips and documents, slides and web pages as input, took the video-editing board from Seedance 2.5 by four points. MiniMax's H3 still holds image-to-video, with Omni 1.1 Flash straight in at second and Seedance 2.5 down to third (arena.ai, 29 August). Seedance 2.5 keeps its own claim: one model generating picture and sound together in 30-second passes, on an API open since 16 July. Kling 3.0 is the value pick at native 4K and 60fps, and Veo 3.1 the safest Western choice with native dialogue and SynthID provenance. One date to diarise: Sora 2's API switches off on 24 September 2026. Open weights are genuinely usable now.

  • Seedance 2.0 / 2.5

    ByteDance's board contender

    Seedance 2.5 sits second on the video-editing board, four points behind Wan 3.0, and third on image-to-video. One model generates picture and sound together in 30-second single passes, so lips, foley and music actually match the frame.

    AudioEditingOur guide →seed.bytedance.com ↗
  • Veo 3.1

    Google DeepMind's flagship

    Native 48kHz dialogue, 4K output and SynthID provenance marks. The default choice where brand safety and clean licensing decide the choice.

    AudioProvenanceOur guide →labs.google ↗
  • Kling 3.0

    Kuaishou's value flagship

    Native 4K at 60fps, 15-second multi-shot storyboards and lip-sync in five languages, at roughly $0.10 a second.

    4K 60fpsValueOur guide →klingai.com ↗
  • Sora 2

    OpenAI · being retired

    Being retired: app and web access closed on 26 April and the API ends on 24 September 2026. Plan migrations now rather than building on it.

    SunsetOur guide →openai.com ↗
  • Runway Gen-4.5

    The filmmaker's control surface

    No longer top of the leaderboards but still the best control surface in the business: motion brushes, keyframes and a real editing timeline.

    ControlOur guide →runwayml.com ↗
  • Ray3 · Dream Machine

    Luma's HDR pioneer

    First native 16-bit HDR video model, with Ray3 Modify for video-to-video work inside the Dream Machine 2.0 studio.

    HDROur guide →lumalabs.ai ↗
  • Wan 2.7 / 3.0

    Alibaba's video family

    Wan 3.0 (24 August 2026) generates 30-second single-pass clips, takes documents, slides and web pages as input, and tops the video-editing arena; its weights are not yet published. Wan 2.7 remains the strongest fully open video family, Apache-licensed with first-and-last-frame control.

    Open weightsArena #1 editingOur guide →wan.video ↗
  • LTX-2.3

    Lightricks' open 4K model

    A 22B open model producing native 4K at 50fps with stereo audio; free for commercial use below a revenue threshold.

    Open weights4KOur guide →lightricks.com ↗
  • Hailuo · MiniMax-H3

    MiniMax's motion specialist

    MiniMax's line is prized for expressive, physical motion: the model animators reach for when movement has to feel alive. Its H3 model now tops the image-to-video arena.

    MotionArena #1 i2vOur guide →hailuoai.video ↗
  • FLUX 3 Video

    Black Forest Labs' video debut

    Black Forest Labs' first video family, rolled out in early August, sits third on the text-to-video arena behind Google's two Omni Flash models, and in the top seven for image-to-video within weeks of release.

    New entrantOur guide →bfl.ai ↗
  • Gemini Omni 1.1 Flash

    conversational video generation

    Google's conversational video line. Omni 1.1 Flash (27 August 2026) adds end frames, scene extension to a cumulative 40 seconds, a 360p draft mode at a third of the cost, and 1080p or 4K export; it tops the text-to-video arena and sits second on image-to-video.

    VideoArena #1 t2vOur guide →deepmind.google ↗
Agentic frameworks / routingAgent frameworks sit between models and outcomes: routing, tools, memory and the guardrails that decide whether an agent is useful or dangerous. Reliability and control now count for more than raw capability.10

Where things standHermes Agent can now browse as you. Version 0.20.6 (27 August) adds consent-gated real-profile browsing: the agent drives a managed copy of your Chrome profile, logins and all, off by default and described in its own documentation as a convenience rather than an isolation boundary. The same release added fifty-plus vendor-hosted remote MCP servers and opt-in keychain encryption for stored secrets, on top of the Herald voice work of 3 August and Quicksilver's 80% cut to first-turn latency in July; its defining trick still stands, writing itself a reusable skill each time it finishes a task. Nous Research, which builds it, is in talks at a $1.5B valuation. xAI's Grok Bot stays in early beta: up to fifty persistent 'AI teammates' on a shared cloud computer, sold through SuperGrok Heavy and Cursor's top plans at $200 a month for Ultra and $120 a seat for Teams Premium. OpenRouter joined the harness game on 4 August with Ori, a CLI that tunes Claude Code, Codex, OpenCode and Hermes to run through its gateway, and agreed on 19 August to join Stripe, with the close still pending. OpenClaw remains the fastest-growing open-source project in GitHub history; Anthropic's Cowork leads desktop agents, Google's Gemini Spark runs around the clock on cloud VMs, and OpenAI's ChatGPT Work ships finished documents, spreadsheets and web apps. MCP is now the universal tool standard. The warning label stands: July's Claw Chain disclosure chained four OpenClaw flaws across roughly 245,000 reachable servers, all since patched.

  • OpenClaw

    Self-hosted agent gateway

    The breakout self-hosted gateway: one daemon, twenty-plus messaging channels, skills and scheduled jobs, any model behind it. Since April it needs API billing rather than a Claude subscription, and it should be hardened before exposure.

    Open sourceSelf-hostedOur guide →openclaw.ai ↗
  • Hermes Agent

    Self-improving open agent

    Nous Research's open-source, self-hosted agent, and the one that changes as you use it: finishing a task writes a reusable skill back into the agent. Version 0.20.6 (27 August) added consent-gated real-profile browsing from a managed copy of your Chrome profile, off by default; the Herald release (3 August) added real-time voice and an agent-to-agent protocol.

    AgentsOpen sourceSelf-hostedOur guide →hermes-agent.nousresearch.com ↗
  • Claude Cowork

    Anthropic's desktop agent

    Anthropic's desktop agent for real file-and-app work, with domain packs for legal and operations tasks. Its launch is what markets nicknamed the SaaSpocalypse.

    Desktopclaude.com ↗
  • Gemini Spark

    Google's always-on agent

    Google's always-on agent living in a cloud VM: deep Gmail, Docs and Calendar access, MCP connectors from day one, and it keeps working while your laptop sleeps.

    Always-ongemini.google.com ↗
  • ChatGPT Agent

    OpenAI's consumer agent

    Agent mode inside ChatGPT with its own virtual computer for browsing, files and multi-step tasks under supervision.

    Consumerchatgpt.com ↗
  • Claude Agent SDK

    Anthropic's builder stack

    The loop behind Claude Code, embeddable in your own product, alongside a managed runtime (beta) where Anthropic hosts the agent for you.

    Buildersdocs.claude.com ↗
  • OpenAI Agents SDK

    Handoffs and guardrails

    OpenAI's production framework with explicit handoffs, guardrails and tracing; MCP support landed early this year.

    Buildersopenai.github.io ↗
  • Model Context Protocol

    The connector standard

    The USB-C of agent tooling: one protocol that lets any client use any tool or data server. Every major platform now speaks it.

    StandardOur guide →modelcontextprotocol.io ↗
  • OpenRouter

    One API, every model

    One API over every major model with live rankings drawn from real usage: the quickest way to A/B models behind an agent without rewrites. Its Ori CLI (4 August 2026) tunes Claude Code, Codex, OpenCode and Hermes to run through the gateway. Stripe agreed to acquire it on 19 August; name, product and roadmap stay.

    Routingopenrouter.ai ↗
  • ChatGPT Work

    OpenAI's finished-work agent

    OpenAI's agent for long-running jobs (9 July 2026): powered by GPT-5.6, it works for hours and returns finished documents, spreadsheets, slides and web apps.

    AgentGPT-5.6openai.com ↗
AI coding & dev toolsAI coding assistants went from autocomplete to delegation: the leading tools now plan, edit across a repo, run tests and open pull requests on their own.8

Where things standCursor changed hands and then got notice from a supplier. SpaceX's $60bn all-stock acquisition of Anysphere closed on 14 August, and on 28 August OpenAI said it intends to wind down Cursor's access to OpenAI models, proposing a 12 November shutoff and writing that it cannot be confident SpaceX will stay within its terms of service. Michael Truell says OpenAI models carry about 5% of Cursor traffic and that the two sides are talking. Behind that, Cursor's agents got direct Google Workspace access on 3 August, and it stays tied with Claude Code at roughly 18% adoption each while Copilot's larger installed base has stopped growing. Claude Code's default model changed on 24 July, when Opus 5 landed at the same price as the model it replaced; Fable 5 remains the tier above it for the hardest work, and the product is already the fastest-growing developer tool on record, at a reported $2.5B revenue run-rate inside nine months. Codex runs on the GPT-5.6 Sol family, cut to $4 and $20 per million on a promotion to 21 November, and Windsurf re-emerged in June as Devin Desktop under Cognition.

  • Claude Code

    Anthropic's terminal agent

    Terminal-native and repo-aware, with Dynamic Workflows for long unattended runs; Opus 5 became its default on 24 July, with Fable 5 the tier above it. Included with a Claude Pro plan.

    TerminalAgentOur guide →claude.com ↗
  • Cursor

    SpaceX's AI-native editor

    The AI-native editor, a SpaceX company since 14 August 2026: Tab completions plus Composer cloud agents that work branches in parallel VMs, running Claude, Gemini and Grok, with OpenAI models due to leave on 12 November unless the wind-down is resolved.

    IDEOur guide →cursor.com ↗
  • GitHub Copilot

    The installed-base leader

    Still the default inside VS Code and enterprises, with a free tier, IP indemnity and new flexible usage billing from 1 June.

    Enterprisegithub.com ↗
  • Codex

    OpenAI's cloud engineer

    OpenAI's asynchronous cloud engineer: give it an issue and it returns a tested pull request. Now powered by the coding-tuned GPT-5.6 Sol.

    Cloud agentOur guide →openai.com ↗
  • Antigravity 2.0

    Google's agent-first IDE

    Google's agent-first IDE built around the Gemini 3.5 line, with a browser-testing loop and a generous free tier via a Google account.

    Free tierantigravity.google ↗
  • Devin Desktop

    Windsurf's successor

    Windsurf's next chapter after the Cognition deal: the Devin autonomous engineer wrapped in the familiar editor, relaunched in June.

    AutonomousOur guide →cognition.ai ↗
  • Kiro

    AWS's spec-driven agent IDE

    Amazon's take: spec-first development where agents implement against written requirements, aimed at teams that want process over improvisation.

    Spec-drivenkiro.dev ↗
  • Replit Agent

    Prompt to deployed app

    From prompt to deployed app in one place: environment, database and hosting handled for you. The fastest route for non-specialists.

    App builderreplit.com ↗
AI companiesThe market is now a platform race: strong models are table stakes, and distribution, chips, agents and developer ecosystems decide who compounds.9

Where things standThe week's biggest story is a report, not a deal: CNBC and Fortune wrote on 27 August that Nvidia has agreed to buy Hugging Face for $12.9bn, both citing people familiar, with Fortune adding that the talks had not produced a signed agreement. Nvidia's newsroom, Hugging Face's blog and EDGAR all carry nothing, and neither company will comment; Nvidia reported $96.2bn of quarterly revenue the day before, up 106% in a year. What is confirmed is OpenAI turning on a customer: on 28 August it told SpaceX it intends to wind down Cursor's access to OpenAI models, proposing a 12 November shutoff, a fortnight after SpaceX's $60bn purchase of Anysphere closed. Anthropic's listing plans have still not produced a public filing: nothing had reached EDGAR by 29 August, past the end-of-month window Bloomberg reported, and Reuters, citing the Wall Street Journal, reported on 25 August a pitch describing a market above $30tn and 2028 revenue projections of $190bn to $200bn; Anthropic has confirmed none of it beyond May's $965B round. OpenAI's reinforcement-learning pause of 18 August has not been formally lifted, with the largest planned frontier run still held and roughly a fifth of inference compute under monitoring, and on 26 August it published its follow-up report on the July compromise of Hugging Face's infrastructure by its own evaluation agents. Stripe's OpenRouter acquisition, announced 19 August, had not closed by the 29th and neither company has named a price. Brussels' enforcement powers stand: since 2 August the EU's AI Office can demand documentation, run technical evaluations and order a model off the market, with fines reaching 3% of global turnover. Unitree's shares keep sliding: 585.00 yuan at the 28 August close, about 3.9 times the 150.80 issue price, from near six times on debut day. Google DeepMind changed hands at the top in early August, Demis Hassabis stepping up to chair with Koray Kavukcuoglu running it day to day and Jeff Dean leaving to found a startup. The Chinese labs keep closing the gap: Kimi K3's weights are open, Qwen's Max-tier weights landed on 8 August, and QwenWork, Alibaba's all-in-one agent workplace, opened an international beta on 26 August. Nous Research, maker of the Hermes agent, is in talks at a $1.5B valuation.

  • Anthropic

    Maker of Claude

    Fable 5 and the Mythos tier up top, Opus 5 (24 July) delivering near-Fable scores at half the price, Sonnet 5 as the default, Claude Code and Cowork carrying the product story. May's round valued it at $965B.

    Frontieranthropic.com ↗
  • OpenAI

    Maker of ChatGPT

    Took GPT-5.6 general on 9 July 2026 and launched ChatGPT Work beside it, a GPT-5.6 agent that runs for hours and returns finished documents, spreadsheets and web apps.

    Frontieropenai.com ↗
  • Google DeepMind

    The full-stack player

    Gemini 3.1 into 3.5, Nano Banana for images, Veo and Omni Flash for video, Spark for agents, and TPUs underneath it all. Since early August, Demis Hassabis chairs while Koray Kavukcuoglu leads day to day.

    Full stackdeepmind.google ↗
  • SpaceXAI (xAI)

    Maker of Grok

    Rebranded SpaceXAI on 6 July 2026 after February's merger into SpaceX and a June Nasdaq listing. SpaceX's $60bn all-stock purchase of Anysphere, maker of Cursor, closed on 14 August, and NVIDIA says Grok's next agentic workloads will run on its Vera Rubin platform.

    Real-timex.ai ↗
  • Moonshot AI

    Maker of Kimi

    Kimi K3, a 2.8-trillion-parameter flagship, launched 16 July 2026 with open weights promised by the 27th, extending Moonshot's run at the front of the open field.

    Open weightsmoonshot.ai ↗
  • Alibaba · Qwen

    China's broadest stack

    Ships the open Qwen line, and delivered on the Max-tier promise: Qwen3.8-Max's 2.4-trillion-parameter weights landed on Hugging Face on 8 August 2026. QwenWork, its all-in-one agent workplace, opened an international beta on 26 August; China's anthropomorphic-agent rules took Qwen's human-like companions offline in July.

    Open weightsqwen.ai ↗
  • ByteDance

    The media-model power

    Dominant in media models: Seedance 2.0 leads video quality boards and Seedream keeps 4K images cheap.

    Mediaseed.bytedance.com ↗
  • DeepSeek

    The price-collapse lab

    The lab whose open releases forced the whole market to justify its margins; its next-generation signals keep that pressure on.

    Open weightsdeepseek.com ↗
  • Meta Superintelligence Labs

    Meta's frontier lab

    Moved Meta's frontier work from open-weight Llama to the proprietary Muse family, and began charging developers for Muse Spark 1.1 through the Meta Model API on 9 July 2026.

    FrontierMuseai.meta.com ↗
Learning & resourcesThe stack changes monthly; the way to stay current is primary documentation plus live leaderboards, not year-old tutorials.8

Where things standBenchmarks now saturate within months of release, so a static score is a snapshot rather than a ranking, and the arenas and usage boards tell you more about what people actually reach for. Every major lab ships genuinely good free documentation and courses, and the MCP specification is required reading for anything agentic.

  • Claude Docs & Academy

    Anthropic's guides and courses

    Model guides, prompting best practice and the Agent SDK reference, plus free certification courses.

    DocsCoursesdocs.claude.com ↗
  • OpenAI Cookbook

    Runnable API recipes

    Hands-on, runnable examples for the API, agents and evals; the fastest way from idea to working call.

    Recipescookbook.openai.com ↗
  • Google AI Studio

    Free Gemini playground

    Free Gemini playground with generous limits: prototype prompts, tune settings and export working code.

    Playgroundaistudio.google.com ↗
  • Hugging Face

    Home of open models

    The home of open weights: models, datasets, Spaces demos and the free training courses behind them.

    Open modelshuggingface.co ↗
  • MCP Documentation

    The connector standard's docs

    The spec, SDKs and server gallery for the connector standard every agent platform adopted.

    AgentsOur guide →modelcontextprotocol.io ↗
  • Artificial Analysis

    Independent model benchmarks

    Independent, continuously updated comparisons of quality, speed and price across every major model.

    Benchmarksartificialanalysis.ai ↗
  • LMArena

    Human preference leaderboard

    Blind head-to-head votes from millions of users; the leaderboard that reflects taste rather than test sets.

    Leaderboardarena.ai ↗
  • Stanford AI Index

    The annual state of AI

    The definitive yearly measurement of the field: capability trends, economics and policy in one report.

    Reporthai.stanford.edu ↗