AI Desk · Model Identity

Tokenizer Atlas

Explore the tokenizer-family structure underlying the AI model ecosystem: which models read text with the same vocabulary, where the families sit, and which entries share ancestry their labels never mention.

Models measured
158
Distinct fingerprints
46
Tokenizer groups
35
Groups with 2+ models
21
As at
22 August 2026

How to read this page: a tokenizer is the fixed vocabulary a model chops text with, and it cannot change without retraining, so it works like a fingerprint. Models in one group below chop all fifty of our test strings identically, which means they share a vocabulary, and usually an ancestry.

The constellation

One star cluster per tokenizer group; size follows membership

Hover a cluster for its family; click to jump to its members. The full listing below carries everything the canvas shows, so nothing depends on it.

Family explorer

One card per shared vocabulary; every model listed on a card chops text identically
Z.aifg_1d05586d · 4 models · 2 signatures
OpenAI cl100kfg_ba918336 · 3 models · 1 signature
Llama4fg_20b9d005 · 3 models · 2 signatures
Llama2fg_5348f9f6 · 2 models · 2 signatures
Upstagefg_865d84ce · 2 models · 1 signature
Novafg_76e424dd · 2 models · 1 signature
Mistralfg_ddc193e9 · 2 models · 1 signature
Gemmafg_ba03436b · 2 models · 1 signature
Z.aifg_3eaaf3e0 · 2 models · 1 signature

One of a kind, so far

A unique EXACT fingerprint does not take a model out of its family: the nearest relative, and how close it comes, is shown on each row
Qwen: Qwen3.8 27Bqwen/qwen3.8-27bno close relative among the models measured yetUnique so farCohere: North Mini Code (free)cohere/north-mini-code:freeno close relative among the models measured yetUnique so farNVIDIA: Nemotron 3.5 Content Safety (free)nvidia/nemotron-3.5-content-safety:freeno close relative among the models measured yetUnique so farLiquidAI: LFM2.5-2.6B (free)liquid/lfm-2.5-2.6b:freeno close relative among the models measured yetUnique so farCohere: Command R7B (12-2024)cohere/command-r7b-12-2024closest relative: Cohere: Command R (08-2024) (49 of 49 strings agree)Unique so farPoolside: Laguna XS 2.1poolside/laguna-xs-2.1no close relative among the models measured yetUnique so farPoolside: Laguna S 2.1poolside/laguna-s-2.1no close relative among the models measured yetUnique so farReka Edgerekaai/reka-edgeno close relative among the models measured yetUnique so farMeta: Llama Guard 4 12Bmeta-llama/llama-guard-4-12bno close relative among the models measured yetUnique so farArcee AI: Trinity Large Thinkingarcee-ai/trinity-large-thinkingno close relative among the models measured yetUnique so farTencent: Hunyuan A13B Instructtencent/hunyuan-a13b-instructno close relative among the models measured yetUnique so farAnthropic: Claude 3 Haikuanthropic/claude-3-haikuno close relative among the models measured yetUnique so farMeituan: LongCat 2.0meituan/longcat-2.0no close relative among the models measured yetUnique so farAmazon: Nova 2 Liteamazon/nova-2-lite-v1no close relative among the models measured yetUnique so far

A model lands here when its exact 50-number fingerprint has no twin yet, which usually means a vocabulary revision inside its own family (Qwen3.8 sits close to the Qwen line without matching it exactly) or relatives that are still unmeasured.

Reading the Atlas

A group gathers every model whose measured fingerprints are the same vocabulary: exact signature matches, plus signatures one uniform template shift apart. Group membership is lineage evidence, not release identification, and the methodology records exactly how much the fingerprints can distinguish. The Atlas's working face is the unknowns queue: what is still unplaced, and what measurement has already placed.