Tokenizer Atlas
Explore the tokenizer-family structure underlying the AI model ecosystem: which models read text with the same vocabulary, where the families sit, and which entries share ancestry their labels never mention.
- Models measured
- 158
- Distinct fingerprints
- 46
- Tokenizer groups
- 35
- Groups with 2+ models
- 21
- As at
- 22 August 2026
How to read this page: a tokenizer is the fixed vocabulary a model chops text with, and it cannot change without retraining, so it works like a fingerprint. Models in one group below chop all fifty of our test strings identically, which means they share a vocabulary, and usually an ancestry.
The constellation
One star cluster per tokenizer group; size follows membershipHover a cluster for its family; click to jump to its members. The full listing below carries everything the canvas shows, so nothing depends on it.
Family explorer
One card per shared vocabulary; every model listed on a card chops text identically- ByteDance: UI-TARS 7B
- Dots Studio: Dots3-Note Preview (free)
- Qwen2.5 72B Instruct
- Qwen: Qwen2.5 7B Instruct
- Qwen: Qwen-Plus
- Qwen: Qwen Plus 0728
- Qwen: Qwen3 14B
- Qwen: Qwen3 235B A22B Instruct 2507
- Mistral: Codestral 2508
- Mistral: Ministral 3 14B 2512
- Mistral: Ministral 3 3B 2512
- Mistral: Ministral 3 8B 2512
- Mistral: Mistral Nemo
- Mistral: Saba
- Mistral: Mistral Small 4
- Mistral: Mistral Small 3.1 24B
- DeepSeek: DeepSeek V3
- DeepSeek: DeepSeek V3 0324
- DeepSeek: DeepSeek V3.1
- DeepSeek: R1 Distill Llama 70B
- DeepSeek: DeepSeek V3.1 Terminus
- DeepSeek: DeepSeek V3.2
- DeepSeek: DeepSeek V3.2 Exp
- DeepSeek: DeepSeek V4 Flash 0423
- Meta: Llama 3.1 70B Instruct
- Meta: Llama 3.1 8B Instruct
- Meta: Llama 3.2 1B Instruct
- Meta: Llama 3.2 3B Instruct
- Meta: Llama 3.3 70B Instruct
- Nous: Hermes 3 405B Instruct
- Nous: Hermes 3 70B Instruct
- Nous: Hermes 4 70B
- Nex AGI: Nex-N2-Mini
- Nex AGI: Nex-N2-Pro
- Perceptron: Perceptron Mk1
- Qwen: Qwen3.5-27B
- Qwen: Qwen3.5-35B-A3B
- Qwen: Qwen3.5-9B
- Qwen: Qwen3.5-Flash
- Qwen: Qwen3.5 Plus 2026-02-15
- OpenAI: GPT-4.1 Nano
- OpenAI: GPT-4o-mini
- OpenAI: GPT-4o-mini (2024-07-18)
- OpenAI: GPT-5 Nano
- OpenAI: GPT-5.1-Codex-Mini
- OpenAI: GPT-5.4 Nano
- OpenAI: GPT-5.6 Luna
- OpenAI: gpt-oss-120b
One of a kind, so far
A unique EXACT fingerprint does not take a model out of its family: the nearest relative, and how close it comes, is shown on each rowA model lands here when its exact 50-number fingerprint has no twin yet, which usually means a vocabulary revision inside its own family (Qwen3.8 sits close to the Qwen line without matching it exactly) or relatives that are still unmeasured.
Reading the Atlas
A group gathers every model whose measured fingerprints are the same vocabulary: exact signature matches, plus signatures one uniform template shift apart. Group membership is lineage evidence, not release identification, and the methodology records exactly how much the fingerprints can distinguish. The Atlas's working face is the unknowns queue: what is still unplaced, and what measurement has already placed.
