AI Desk · Model Identity · Evidence object

Tokenizer signature tk_2d06485c

An exact tokenizer signature: every model listed below returned the same marginal prompt-token count on all 50 test strings. Sharing a signature is strong evidence of a shared vocabulary. This signature sits in the Claude group.

Members
1
First observed
21 August 2026
Last observed
21 August 2026
Group
Claude

Models with this signature

The measured vector

Marginal prompt-token cost per string, as measured on Anthropic: Claude 3 Haiku
All 50 rows

English and whitespace

en-prose15
en-long17
spaces-204
spaces-606
tabs-204
newlines-203
mixed-ws9

Digits

digits-94
digits-125
digits-3011
digits-sep8
float-long10

CJK

zh-common19
zh-long25
zh-rare21
ja-kana12
ja-kanji17
ko20

Other scripts

ru14
ar21
he18
hi23
th25
el22

Emoji

emoji-basic11
emoji-skin26
emoji-zwj-family15
emoji-zwj-x343
emoji-flags25
emoji-prof28

Rare Unicode

math47
boxdraw29
combining6
cjk-ext-b17
surrogates55
zalgo36
rtl-mix10

Code

py-code30
py-indent19
json23
html17
regex52
camel7
snake11

Repetition and encodings

rare-word-x541
repeat-tok21
base6443
hex14
url22
uuid28

Method: tokenizer fingerprinting, v1-50. Raw data: fingerprints.json · index.json · test-strings.json.