AI Desk · Model Identity · Evidence object

Tokenizer signature tk_73c4f0dc

An exact tokenizer signature: every model listed below returned the same marginal prompt-token count on all 50 test strings. Sharing a signature is strong evidence of a shared vocabulary. This signature sits in the DeepSeek group.

Members
10
First observed
20 August 2026
Last observed
21 August 2026
Group
DeepSeek

Models with this signature

Template variants of the same tokenizer

These signatures differ from this one by a uniform one-token shift on the affected strings: the same vocabulary behind a different chat template.

tk_869dc1cf · 2 models

The measured vector

Marginal prompt-token cost per string, as measured on DeepSeek: DeepSeek V3
All 50 rows

English and whitespace

en-prose14
en-long16
spaces-203
spaces-603
tabs-204
newlines-204
mixed-ws7

Digits

digits-93
digits-124
digits-3010
digits-sep7
float-long9

CJK

zh-common9
zh-long10
zh-rare18
ja-kana9
ja-kanji11
ko13

Other scripts

ru13
ar10
he14
hi18
th11
el16

Emoji

emoji-basic10
emoji-skin14
emoji-zwj-family11
emoji-zwj-x333
emoji-flags16
emoji-prof19

Rare Unicode

math22
boxdraw27
combining10
cjk-ext-b16
surrogates33
zalgo36
rtl-mix6

Code

py-code25
py-indent15
json20
html16
regex47
camel7
snake8

Repetition and encodings

rare-word-x526
repeat-tok41
base6434
hex12
url17
uuid27

Method: tokenizer fingerprinting, v1-50. Raw data: fingerprints.json · index.json · test-strings.json.