YFarmX
AI Desk · Model Identity

Tokenizer Atlas

Explore the tokenizer-family structure underlying the AI model ecosystem: which models read text with the same vocabulary, where the families sit, and which entries share ancestry their labels never mention.

Models measured
280
Complete fingerprints
262
Distinct fingerprints
65
Tokenizer groups
51
Groups with 2+ models
26
As at
18 September 2026

How to read this page: a tokenizer is the fixed vocabulary a model chops text with, and it cannot change without retraining, so it works like a fingerprint. Of the 280 models measured, 262 returned all fifty rows clean; only those carry an exact signature and appear in the groups below. A group gathers the signatures the evidence shows to be one tokenizer.

The constellation

One star cluster per tokenizer group; size follows membership

Tap or hover a cluster for its family; click to jump to its members. Position and distance on the canvas carry no meaning: cluster size follows membership, and only the grouping itself is data. The full listing below carries everything the canvas shows, so nothing depends on it.

The fingerprint matrix

One column per major group, one row per test string. A bright row is a string that tells the vocabularies apart; a faded row is one they nearly agree on.
Qwen AGPTMistral AGemini AQwen BDeepSeekLlama3fg_c018f9f2en-proseen-prose · Qwen A: 14 tokensen-prose · GPT: 14 tokensen-prose · Mistral A: 14 tokensen-prose · Gemini A: 14 tokensen-prose · Qwen B: 14 tokensen-prose · DeepSeek: 14 tokensen-prose · Llama3: 14 tokensen-prose · fg_c018f9f2: 13 tokensen-longen-long · Qwen A: 16 tokensen-long · GPT: 17 tokensen-long · Mistral A: 18 tokensen-long · Gemini A: 16 tokensen-long · Qwen B: 16 tokensen-long · DeepSeek: 16 tokensen-long · Llama3: 16 tokensen-long · fg_c018f9f2: 15 tokensspaces-20spaces-20 · Qwen A: 3 tokensspaces-20 · GPT: 3 tokensspaces-20 · Mistral A: 3 tokensspaces-20 · Gemini A: 3 tokensspaces-20 · Qwen B: 3 tokensspaces-20 · DeepSeek: 3 tokensspaces-20 · Llama3: 3 tokensspaces-20 · fg_c018f9f2: 3 tokensspaces-60spaces-60 · Qwen A: 3 tokensspaces-60 · GPT: 3 tokensspaces-60 · Mistral A: 3 tokensspaces-60 · Gemini A: 4 tokensspaces-60 · Qwen B: 3 tokensspaces-60 · DeepSeek: 3 tokensspaces-60 · Llama3: 3 tokensspaces-60 · fg_c018f9f2: 3 tokenstabs-20tabs-20 · Qwen A: 3 tokenstabs-20 · GPT: 3 tokenstabs-20 · Mistral A: 4 tokenstabs-20 · Gemini A: 3 tokenstabs-20 · Qwen B: 3 tokenstabs-20 · DeepSeek: 4 tokenstabs-20 · Llama3: 3 tokenstabs-20 · fg_c018f9f2: 3 tokensnewlines-20newlines-20 · Qwen A: 4 tokensnewlines-20 · GPT: 4 tokensnewlines-20 · Mistral A: 7 tokensnewlines-20 · Gemini A: 3 tokensnewlines-20 · Qwen B: 4 tokensnewlines-20 · DeepSeek: 4 tokensnewlines-20 · Llama3: 4 tokensnewlines-20 · fg_c018f9f2: 4 tokensmixed-wsmixed-ws · Qwen A: 5 tokensmixed-ws · GPT: 6 tokensmixed-ws · Mistral A: 7 tokensmixed-ws · Gemini A: 9 tokensmixed-ws · Qwen B: 5 tokensmixed-ws · DeepSeek: 7 tokensmixed-ws · Llama3: 5 tokensmixed-ws · fg_c018f9f2: 5 tokensdigits-9digits-9 · Qwen A: 9 tokensdigits-9 · GPT: 3 tokensdigits-9 · Mistral A: 9 tokensdigits-9 · Gemini A: 9 tokensdigits-9 · Qwen B: 9 tokensdigits-9 · DeepSeek: 3 tokensdigits-9 · Llama3: 3 tokensdigits-9 · fg_c018f9f2: 5 tokensdigits-12digits-12 · Qwen A: 12 tokensdigits-12 · GPT: 4 tokensdigits-12 · Mistral A: 12 tokensdigits-12 · Gemini A: 12 tokensdigits-12 · Qwen B: 12 tokensdigits-12 · DeepSeek: 4 tokensdigits-12 · Llama3: 4 tokensdigits-12 · fg_c018f9f2: 7 tokensdigits-30digits-30 · Qwen A: 30 tokensdigits-30 · GPT: 10 tokensdigits-30 · Mistral A: 30 tokensdigits-30 · Gemini A: 30 tokensdigits-30 · Qwen B: 30 tokensdigits-30 · DeepSeek: 10 tokensdigits-30 · Llama3: 10 tokensdigits-30 · fg_c018f9f2: 10 tokensdigits-sepdigits-sep · Qwen A: 12 tokensdigits-sep · GPT: 7 tokensdigits-sep · Mistral A: 12 tokensdigits-sep · Gemini A: 12 tokensdigits-sep · Qwen B: 12 tokensdigits-sep · DeepSeek: 7 tokensdigits-sep · Llama3: 7 tokensdigits-sep · fg_c018f9f2: 9 tokensfloat-longfloat-long · Qwen A: 22 tokensfloat-long · GPT: 9 tokensfloat-long · Mistral A: 22 tokensfloat-long · Gemini A: 22 tokensfloat-long · Qwen B: 22 tokensfloat-long · DeepSeek: 9 tokensfloat-long · Llama3: 9 tokensfloat-long · fg_c018f9f2: 14 tokenszh-commonzh-common · Qwen A: 10 tokenszh-common · GPT: 12 tokenszh-common · Mistral A: 16 tokenszh-common · Gemini A: 10 tokenszh-common · Qwen B: 9 tokenszh-common · DeepSeek: 9 tokenszh-common · Llama3: 12 tokenszh-common · fg_c018f9f2: 9 tokenszh-longzh-long · Qwen A: 12 tokenszh-long · GPT: 17 tokenszh-long · Mistral A: 20 tokenszh-long · Gemini A: 12 tokenszh-long · Qwen B: 10 tokenszh-long · DeepSeek: 10 tokenszh-long · Llama3: 18 tokenszh-long · fg_c018f9f2: 10 tokenszh-rarezh-rare · Qwen A: 17 tokenszh-rare · GPT: 20 tokenszh-rare · Mistral A: 20 tokenszh-rare · Gemini A: 25 tokenszh-rare · Qwen B: 18 tokenszh-rare · DeepSeek: 18 tokenszh-rare · Llama3: 18 tokenszh-rare · fg_c018f9f2: 18 tokensja-kanaja-kana · Qwen A: 8 tokensja-kana · GPT: 10 tokensja-kana · Mistral A: 11 tokensja-kana · Gemini A: 7 tokensja-kana · Qwen B: 7 tokensja-kana · DeepSeek: 9 tokensja-kana · Llama3: 11 tokensja-kana · fg_c018f9f2: 10 tokensja-kanjija-kanji · Qwen A: 13 tokensja-kanji · GPT: 13 tokensja-kanji · Mistral A: 13 tokensja-kanji · Gemini A: 9 tokensja-kanji · Qwen B: 10 tokensja-kanji · DeepSeek: 11 tokensja-kanji · Llama3: 12 tokensja-kanji · fg_c018f9f2: 14 tokenskoko · Qwen A: 17 tokensko · GPT: 11 tokensko · Mistral A: 13 tokensko · Gemini A: 11 tokensko · Qwen B: 13 tokensko · DeepSeek: 13 tokensko · Llama3: 14 tokensko · fg_c018f9f2: 17 tokensruru · Qwen A: 16 tokensru · GPT: 11 tokensru · Mistral A: 12 tokensru · Gemini A: 12 tokensru · Qwen B: 12 tokensru · DeepSeek: 13 tokensru · Llama3: 14 tokensru · fg_c018f9f2: 13 tokensarar · Qwen A: 8 tokensar · GPT: 11 tokensar · Mistral A: 9 tokensar · Gemini A: 10 tokensar · Qwen B: 10 tokensar · DeepSeek: 10 tokensar · Llama3: 12 tokensar · fg_c018f9f2: 13 tokenshehe · Qwen A: 9 tokenshe · GPT: 10 tokenshe · Mistral A: 10 tokenshe · Gemini A: 13 tokenshe · Qwen B: 14 tokenshe · DeepSeek: 14 tokenshe · Llama3: 24 tokenshe · fg_c018f9f2: 24 tokenshihi · Qwen A: 29 tokenshi · GPT: 11 tokenshi · Mistral A: 14 tokenshi · Gemini A: 8 tokenshi · Qwen B: 15 tokenshi · DeepSeek: 18 tokenshi · Llama3: 15 tokenshi · fg_c018f9f2: 27 tokensthth · Qwen A: 12 tokensth · GPT: 11 tokensth · Mistral A: 17 tokensth · Gemini A: 10 tokensth · Qwen B: 8 tokensth · DeepSeek: 11 tokensth · Llama3: 14 tokensth · fg_c018f9f2: 24 tokenselel · Qwen A: 28 tokensel · GPT: 14 tokensel · Mistral A: 13 tokensel · Gemini A: 13 tokensel · Qwen B: 13 tokensel · DeepSeek: 16 tokensel · Llama3: 13 tokensel · fg_c018f9f2: 14 tokensemoji-basicemoji-basic · Qwen A: 5 tokensemoji-basic · GPT: 8 tokensemoji-basic · Mistral A: 20 tokensemoji-basic · Gemini A: 5 tokensemoji-basic · Qwen B: 9 tokensemoji-basic · DeepSeek: 10 tokensemoji-basic · Llama3: 10 tokensemoji-basic · fg_c018f9f2: 5 tokensemoji-skinemoji-skin · Qwen A: 10 tokensemoji-skin · GPT: 13 tokensemoji-skin · Mistral A: 40 tokensemoji-skin · Gemini A: 10 tokensemoji-skin · Qwen B: 30 tokensemoji-skin · DeepSeek: 14 tokensemoji-skin · Llama3: 30 tokensemoji-skin · fg_c018f9f2: 10 tokensemoji-zwj-familyemoji-zwj-family · Qwen A: 10 tokensemoji-zwj-family · GPT: 11 tokensemoji-zwj-family · Mistral A: 19 tokensemoji-zwj-family · Gemini A: 7 tokensemoji-zwj-family · Qwen B: 18 tokensemoji-zwj-family · DeepSeek: 11 tokensemoji-zwj-family · Llama3: 15 tokensemoji-zwj-family · fg_c018f9f2: 7 tokensemoji-zwj-x3emoji-zwj-x3 · Qwen A: 30 tokensemoji-zwj-x3 · GPT: 33 tokensemoji-zwj-x3 · Mistral A: 57 tokensemoji-zwj-x3 · Gemini A: 21 tokensemoji-zwj-x3 · Qwen B: 54 tokensemoji-zwj-x3 · DeepSeek: 33 tokensemoji-zwj-x3 · Llama3: 45 tokensemoji-zwj-x3 · fg_c018f9f2: 21 tokensemoji-flagsemoji-flags · Qwen A: 8 tokensemoji-flags · GPT: 16 tokensemoji-flags · Mistral A: 32 tokensemoji-flags · Gemini A: 8 tokensemoji-flags · Qwen B: 24 tokensemoji-flags · DeepSeek: 16 tokensemoji-flags · Llama3: 24 tokensemoji-flags · fg_c018f9f2: 24 tokensemoji-profemoji-prof · Qwen A: 14 tokensemoji-prof · GPT: 20 tokensemoji-prof · Mistral A: 35 tokensemoji-prof · Gemini A: 11 tokensemoji-prof · Qwen B: 29 tokensemoji-prof · DeepSeek: 19 tokensemoji-prof · Llama3: 26 tokensemoji-prof · fg_c018f9f2: 15 tokensmathmath · Qwen A: 26 tokensmath · GPT: 31 tokensmath · Mistral A: 33 tokensmath · Gemini A: 26 tokensmath · Qwen B: 29 tokensmath · DeepSeek: 22 tokensmath · Llama3: 31 tokensmath · fg_c018f9f2: 27 tokensboxdrawboxdraw · Qwen A: 17 tokensboxdraw · GPT: 27 tokensboxdraw · Mistral A: 30 tokensboxdraw · Gemini A: 17 tokensboxdraw · Qwen B: 28 tokensboxdraw · DeepSeek: 27 tokensboxdraw · Llama3: 23 tokensboxdraw · fg_c018f9f2: 28 tokenscombiningcombining · Qwen A: 5 tokenscombining · GPT: 11 tokenscombining · Mistral A: 14 tokenscombining · Gemini A: 10 tokenscombining · Qwen B: 5 tokenscombining · DeepSeek: 10 tokenscombining · Llama3: 11 tokenscombining · fg_c018f9f2: 10 tokenscjk-ext-bcjk-ext-b · Qwen A: 12 tokenscjk-ext-b · GPT: 13 tokenscjk-ext-b · Mistral A: 16 tokenscjk-ext-b · Gemini A: 16 tokenscjk-ext-b · Qwen B: 13 tokenscjk-ext-b · DeepSeek: 16 tokenscjk-ext-b · Llama3: 13 tokenscjk-ext-b · fg_c018f9f2: 13 tokenssurrogatessurrogates · Qwen A: 22 tokenssurrogates · GPT: 40 tokenssurrogates · Mistral A: 53 tokenssurrogates · Gemini A: 27 tokenssurrogates · Qwen B: 40 tokenssurrogates · DeepSeek: 33 tokenssurrogates · Llama3: 40 tokenssurrogates · fg_c018f9f2: 40 tokenszalgozalgo · Qwen A: 36 tokenszalgo · GPT: 35 tokenszalgo · Mistral A: 36 tokenszalgo · Gemini A: 22 tokenszalgo · Qwen B: 36 tokenszalgo · DeepSeek: 36 tokenszalgo · Llama3: 35 tokenszalgo · fg_c018f9f2: 22 tokensrtl-mixrtl-mix · Qwen A: 7 tokensrtl-mix · GPT: 5 tokensrtl-mix · Mistral A: 6 tokensrtl-mix · Gemini A: 5 tokensrtl-mix · Qwen B: 6 tokensrtl-mix · DeepSeek: 6 tokensrtl-mix · Llama3: 8 tokensrtl-mix · fg_c018f9f2: 8 tokenspy-codepy-code · Qwen A: 25 tokenspy-code · GPT: 25 tokenspy-code · Mistral A: 25 tokenspy-code · Gemini A: 28 tokenspy-code · Qwen B: 26 tokenspy-code · DeepSeek: 25 tokenspy-code · Llama3: 25 tokenspy-code · fg_c018f9f2: 24 tokenspy-indentpy-indent · Qwen A: 15 tokenspy-indent · GPT: 15 tokenspy-indent · Mistral A: 15 tokenspy-indent · Gemini A: 18 tokenspy-indent · Qwen B: 18 tokenspy-indent · DeepSeek: 15 tokenspy-indent · Llama3: 15 tokenspy-indent · fg_c018f9f2: 15 tokensjsonjson · Qwen A: 25 tokensjson · GPT: 20 tokensjson · Mistral A: 25 tokensjson · Gemini A: 28 tokensjson · Qwen B: 25 tokensjson · DeepSeek: 20 tokensjson · Llama3: 20 tokensjson · fg_c018f9f2: 20 tokenshtmlhtml · Qwen A: 16 tokenshtml · GPT: 16 tokenshtml · Mistral A: 16 tokenshtml · Gemini A: 14 tokenshtml · Qwen B: 16 tokenshtml · DeepSeek: 16 tokenshtml · Llama3: 16 tokenshtml · fg_c018f9f2: 15 tokensregexregex · Qwen A: 44 tokensregex · GPT: 44 tokensregex · Mistral A: 49 tokensregex · Gemini A: 44 tokensregex · Qwen B: 44 tokensregex · DeepSeek: 47 tokensregex · Llama3: 44 tokensregex · fg_c018f9f2: 44 tokenscamelcamel · Qwen A: 5 tokenscamel · GPT: 6 tokenscamel · Mistral A: 6 tokenscamel · Gemini A: 6 tokenscamel · Qwen B: 5 tokenscamel · DeepSeek: 7 tokenscamel · Llama3: 5 tokenscamel · fg_c018f9f2: 5 tokenssnakesnake · Qwen A: 6 tokenssnake · GPT: 6 tokenssnake · Mistral A: 7 tokenssnake · Gemini A: 11 tokenssnake · Qwen B: 6 tokenssnake · DeepSeek: 8 tokenssnake · Llama3: 6 tokenssnake · fg_c018f9f2: 6 tokensrare-word-x5rare-word-x5 · Qwen A: 31 tokensrare-word-x5 · GPT: 31 tokensrare-word-x5 · Mistral A: 36 tokensrare-word-x5 · Gemini A: 25 tokensrare-word-x5 · Qwen B: 31 tokensrare-word-x5 · DeepSeek: 26 tokensrare-word-x5 · Llama3: 30 tokensrare-word-x5 · fg_c018f9f2: 30 tokensrepeat-tokrepeat-tok · Qwen A: 22 tokensrepeat-tok · GPT: 22 tokensrepeat-tok · Mistral A: 41 tokensrepeat-tok · Gemini A: 21 tokensrepeat-tok · Qwen B: 22 tokensrepeat-tok · DeepSeek: 41 tokensrepeat-tok · Llama3: 21 tokensrepeat-tok · fg_c018f9f2: 21 tokensbase64base64 · Qwen A: 36 tokensbase64 · GPT: 34 tokensbase64 · Mistral A: 39 tokensbase64 · Gemini A: 34 tokensbase64 · Qwen B: 36 tokensbase64 · DeepSeek: 34 tokensbase64 · Llama3: 35 tokensbase64 · fg_c018f9f2: 34 tokenshexhex · Qwen A: 17 tokenshex · GPT: 11 tokenshex · Mistral A: 19 tokenshex · Gemini A: 16 tokenshex · Qwen B: 17 tokenshex · DeepSeek: 12 tokenshex · Llama3: 11 tokenshex · fg_c018f9f2: 14 tokensurlurl · Qwen A: 19 tokensurl · GPT: 17 tokensurl · Mistral A: 19 tokensurl · Gemini A: 23 tokensurl · Qwen B: 19 tokensurl · DeepSeek: 17 tokensurl · Llama3: 17 tokensurl · fg_c018f9f2: 17 tokensuuiduuid · Qwen A: 36 tokensuuid · GPT: 27 tokensuuid · Mistral A: 36 tokensuuid · Gemini A: 36 tokensuuid · Qwen B: 36 tokensuuid · DeepSeek: 27 tokensuuid · Llama3: 27 tokensuuid · fg_c018f9f2: 28 tokens

Each cell is the marginal token cost of that string under that group's exemplar, shaded within its row from the row's lowest count to its highest. Exact values sit in the cell tooltips, and two models are compared string by string on thebench.

Family explorer

One card per inferred tokenizer group with two or more models; unique signatures sit in their own section beneath. Signatures are grouped when the evidence shows one vocabulary behind different serving templates.
Z.aifg_3eaaf3e0 · 7 models · 1 measured signature · 1 inferred tokenizer
OpenAI cl100kfg_ba918336 · 4 models · 1 measured signature · 1 inferred tokenizer
Tencent · group Afg_2b541395 · 4 models · 1 measured signature · 1 inferred tokenizer
Nova · group Afg_76e424dd · 3 models · 1 measured signature · 1 inferred tokenizer
Mistral · group Bfg_ddc193e9 · 3 models · 1 measured signature · 1 inferred tokenizer
Moonshot AI · group Afg_ea6068bf · 3 models · 1 measured signature · 1 inferred tokenizer
Grok · group Afg_a17bbc1d · 3 models · 1 measured signature · 1 inferred tokenizer
Upstagefg_865d84ce · 2 models · 1 measured signature · 1 inferred tokenizer
Cohere · group Afg_2729e470 · 2 models · 1 measured signature · 1 inferred tokenizer
Gemmafg_ba03436b · 2 models · 1 measured signature · 1 inferred tokenizer
Claudefg_2d06485c · 2 models · 1 measured signature · 1 inferred tokenizer
Aion Labs · group Afg_d3c6ae80 · 2 models · 1 measured signature · 1 inferred tokenizer

When one family name appears on several groups (Qwen has 4; Mistral has 4; Tencent has 4; Nova has 3; Moonshot AI has 3; Gemini has 2; Grok has 2; Cohere has 2; Aion Labs has 2; Poolside has 2), those are distinct measured vocabularies from the same broader model family: tokenizer generations, lettered group A, B and onwards in size order. The letters are ours; the evidence does not name the generations.

One of a kind, so far

A unique EXACT fingerprint does not take a model out of its family: the nearest relative, and how close it comes, is shown on each row
Qwen: Qwen3.8 27Bqwen/qwen3.8-27bclosest relative: Qwen: Qwen3.5-122B-A10B (44 of 48 strings agree)Unique so farCohere: North Mini Code (free)cohere/north-mini-code:freeno close relative among the models measured yetUnique so farNVIDIA: Nemotron 3.5 Content Safety (free)nvidia/nemotron-3.5-content-safety:freeno close relative among the models measured yetUnique so farLiquidAI: LFM2.5-2.6B (free)liquid/lfm-2.5-2.6b:freeno close relative among the models measured yetUnique so farPoolside: Laguna XS 2.1poolside/laguna-xs-2.1no close relative among the models measured yetUnique so farPoolside: Laguna S 2.1poolside/laguna-s-2.1no close relative among the models measured yetUnique so farReka Edgerekaai/reka-edgeno close relative among the models measured yetUnique so farMeta: Llama Guard 4 12Bmeta-llama/llama-guard-4-12bno close relative among the models measured yetUnique so farArcee AI: Trinity Large Thinkingarcee-ai/trinity-large-thinkingno close relative among the models measured yetUnique so farTencent: Hunyuan A13B Instructtencent/hunyuan-a13b-instructclosest relative: Tencent: Hy-MT2-7B (49 of 50 strings agree)Unique so farMeituan: LongCat 2.0meituan/longcat-2.0no close relative among the models measured yetUnique so farAmazon: Nova 2 Liteamazon/nova-2-lite-v1no close relative among the models measured yetUnique so farTheDrummer: UnslopNemo 12Bthedrummer/unslopnemo-12bno close relative among the models measured yetUnique so farBaidu: ERNIE 4.5 VL 424B A47B baidu/ernie-4.5-vl-424b-a47bno close relative among the models measured yetUnique so farWizardLM-2 8x22Bmicrosoft/wizardlm-2-8x22bno close relative among the models measured yetUnique so farMoonshotAI: Kimi K2 0711moonshotai/kimi-k2closest relative: MoonshotAI: Kimi K2.5 (38 of 48 strings agree)Unique so farGoogle: Gemma 2 27Bgoogle/gemma-2-27b-itno close relative among the models measured yetUnique so farMoonshotAI: Kimi K2 0905moonshotai/kimi-k2-0905closest relative: MoonshotAI: Kimi K2.6 (49 of 50 strings agree)Unique so farAionLabs: Aion-RP 1.0 (8B)aion-labs/aion-rp-llama-3.1-8bno close relative among the models measured yetUnique so farSpaceXAI: Grok 4.5x-ai/grok-4.5closest relative: SpaceXAI: Grok 4.6 (49 of 49 strings agree)Unique so farAmazon: Nova Premier 1.0amazon/nova-premier-v1closest relative: Amazon: Nova Micro 1.0 (40 of 50 strings agree)Unique so farIBM: Granite 4.2 8Bibm-granite/granite-4.2-8bclosest relative: IBM: Granite 4.0 Micro (35 of 50 strings agree)Unique so farTencent: Hy-MT2-7Btencent/hy-mt2-7bclosest relative: Tencent: Hunyuan A13B Instruct (49 of 50 strings agree)Unique so farTencent: Hy4 previewtencent/hy4-previewno close relative among the models measured yetUnique so farTypeSafe: Jev 1.13typesafe/jev-1.13no close relative among the models measured yetUnique so far

A model lands here when its exact 50-number fingerprint has no twin yet, which usually means a vocabulary revision inside its own family (Qwen3.8 sits close to the Qwen line without matching it exactly) or relatives that are still unmeasured.

Reading the Atlas

A group gathers every model whose measured fingerprints are the same vocabulary: exact signature matches, plus signatures one uniform template shift apart. Group membership is lineage evidence, not release identification, and the methodology records exactly how much the fingerprints can distinguish. The Atlas's working face is the unknowns queue: what is still unplaced, and what measurement has already placed.