YFarmX

AI Desk · Model Identity · Compare

DeepSeek: DeepSeek V4 Pro 0813 · Tencent: Hy4 preview

Identity similarity: lowDivergent fingerprints · 28/50These models tokenise differently. Divergence on the high-information strings (CJK, emoji, digits) separates unrelated vocabularies within a handful of samples.

DeepSeek: DeepSeek V4 Pro 0813 recordTencent: Hy4 preview recordOpen in the interactive bench

Tokenizer fingerprint

28 of 50 mutually clean strings agree
Test stringDeepSeek: DeepSeek V4 Pro 0813Tencent: Hy4 previewVerdict
en-prose1414match
en-long1617differs
spaces-2034differs
spaces-6036differs
tabs-20411differs
newlines-2047differs
mixed-ws77match
digits-933match
digits-1244match
digits-301010match
digits-sep77match
float-long99match
zh-common98differs
zh-long1010match
zh-rare1818match
ja-kana910differs
ja-kanji1111match
ko1312differs
ru1313match
ar1012differs
he1414match
hi1817differs
th1114differs
el1620differs
emoji-basic1010match
emoji-skin1414match
emoji-zwj-family1111match
emoji-zwj-x33333match
emoji-flags1616match
emoji-prof1919match
math2227differs
boxdraw2728differs
combining1010match
cjk-ext-b1616match
surrogates3341differs
zalgo3636match
rtl-mix67differs
py-code2525match
py-indent1516differs
json2020match
html1616match
regex4746differs
camel76differs
snake86differs
rare-word-x52631differs
repeat-tok4141match
base643435differs
hex1212match
url1717match
uuid2727match

Values are the marginal prompt-token cost of each string with the model's fixed overhead subtracted. An x marks a row excluded as corrupt for that model.

API surface

The declared serving contracts differ
FieldDeepSeek: DeepSeek V4 Pro 0813Tencent: Hy4 previewVerdict
Context length10485761048576✓ match
Max output tokens39321664000× differs
Supported parametersfrequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_pinclude_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, stop, structured_outputs, temperature, tool_choice, tools× differs
Defaults{"temperature":1,"top_p":1}{}× differs
Reasoning contract{"mandatory":false,"supported_efforts":["max","high","low"],"default_effort":"high"}{"mandatory":false,"default_enabled":true,"supported_efforts":["high","low","none"],"default_effort":"high"}× differs
Modalitiestext->texttext->text✓ match

Method: tokenizer fingerprinting, v1, 50 strings. This page is static and dated; the interactive bench compares any two of the 486 catalogue models.