AI Desk · Model Identity · Evidence object

Tokenizer signature tk_03ecdf53

An exact tokenizer signature: every model listed below returned the same marginal prompt-token count on all 50 test strings. Sharing a signature is strong evidence of a shared vocabulary. This signature sits in the Llama2 group.

Members
1
First observed
21 August 2026
Last observed
21 August 2026
Group
Llama2

Models with this signature

Template variants of the same tokenizer

These signatures differ from this one by a uniform one-token shift on the affected strings: the same vocabulary behind a different chat template.

tk_5348f9f6 · 1 model

The measured vector

Marginal prompt-token cost per string, as measured on Mancer: Weaver (alpha)
All 50 rows

English and whitespace

en-prose16
en-long18
spaces-204
spaces-606
tabs-2022
newlines-2022
mixed-ws9

Digits

digits-99
digits-1212
digits-3030
digits-sep12
float-long22

CJK

zh-common22
zh-long34
zh-rare27
ja-kana16
ja-kanji16
ko29

Other scripts

ru14
ar29
he26
hi31
th26
el33

Emoji

emoji-basic20
emoji-skin40
emoji-zwj-family19
emoji-zwj-x357
emoji-flags32
emoji-prof35

Rare Unicode

math26
boxdraw17
combining10
cjk-ext-b16
surrogates54
zalgo35
rtl-mix13

Code

py-code28
py-indent19
json28
html17
regex52
camel8
snake12

Repetition and encodings

rare-word-x545
repeat-tok40
base6441
hex20
url23
uuid36

Method: tokenizer fingerprinting, v1-50. Raw data: fingerprints.json · index.json · test-strings.json.