AI Desk · Model Identity · Evidence object

Tokenizer signature tk_d1411c9d

An exact tokenizer signature: every model listed below returned the same marginal prompt-token count on all 50 test strings. Sharing a signature is strong evidence of a shared vocabulary. This signature sits in the GPT group.

Members
10
First observed
21 August 2026
Last observed
21 August 2026
Group
GPT

Models with this signature

The measured vector

Marginal prompt-token cost per string, as measured on OpenAI: GPT-4.1 Nano
All 50 rows

English and whitespace

en-prose14
en-long17
spaces-203
spaces-603
tabs-203
newlines-204
mixed-ws6

Digits

digits-93
digits-124
digits-3010
digits-sep7
float-long9

CJK

zh-common12
zh-long17
zh-rare20
ja-kana10
ja-kanji13
ko11

Other scripts

ru11
ar11
he10
hi11
th11
el14

Emoji

emoji-basic8
emoji-skin13
emoji-zwj-family11
emoji-zwj-x333
emoji-flags16
emoji-prof20

Rare Unicode

math31
boxdraw27
combining11
cjk-ext-b13
surrogates40
zalgo35
rtl-mix5

Code

py-code25
py-indent15
json20
html16
regex44
camel6
snake6

Repetition and encodings

rare-word-x531
repeat-tok22
base6434
hex11
url17
uuid27

Method: tokenizer fingerprinting, v1-50. Raw data: fingerprints.json · index.json · test-strings.json.