AI Desk · Model Identity · Compare

DeepSeek: DeepSeek V4 Flash 0423 · DeepSeek: DeepSeek V4 Pro 0813

Identity similarity: very highExact tokenizer agreement · 50/50Exact agreement across every measured string is strong evidence of shared tokenizer lineage. It shows shared ancestry, and does not on its own establish that these are the same release.

DeepSeek: DeepSeek V4 Flash 0423 recordDeepSeek: DeepSeek V4 Pro 0813 recordOpen in the interactive bench

Tokenizer fingerprint

50 of 50 mutually clean strings agree
Test stringDeepSeek: DeepSeek V4 Flash 0423DeepSeek: DeepSeek V4 Pro 0813Verdict
en-prose1414match
en-long1616match
spaces-2033match
spaces-6033match
tabs-2044match
newlines-2044match
mixed-ws77match
digits-933match
digits-1244match
digits-301010match
digits-sep77match
float-long99match
zh-common99match
zh-long1010match
zh-rare1818match
ja-kana99match
ja-kanji1111match
ko1313match
ru1313match
ar1010match
he1414match
hi1818match
th1111match
el1616match
emoji-basic1010match
emoji-skin1414match
emoji-zwj-family1111match
emoji-zwj-x33333match
emoji-flags1616match
emoji-prof1919match
math2222match
boxdraw2727match
combining1010match
cjk-ext-b1616match
surrogates3333match
zalgo3636match
rtl-mix66match
py-code2525match
py-indent1515match
json2020match
html1616match
regex4747match
camel77match
snake88match
rare-word-x52626match
repeat-tok4141match
base643434match
hex1212match
url1717match
uuid2727match

Values are the marginal prompt-token cost of each string with the model's fixed overhead subtracted. An x marks a row excluded as corrupt for that model.

API surface

The declared serving contracts differ
FieldDeepSeek: DeepSeek V4 Flash 0423DeepSeek: DeepSeek V4 Pro 0813Verdict
Context length10485761048576✓ match
Max output tokens384000× differs
Supported parametersfrequency_penalty, include_reasoning, logit_bias, logprobs, max_completion_tokens, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_a, top_k, top_logprobs, top_pfrequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p× differs
Defaults{}{"temperature":1,"top_p":1}× differs
Reasoning contract{"mandatory":false,"supported_efforts":["xhigh","high"],"default_effort":"high"}{"mandatory":false,"supported_efforts":["max","high","low"],"default_effort":"high"}× differs
Modalitiestext->texttext->text✓ match

Method: tokenizer fingerprinting, v1, 50 strings. Raw data on the data page. This page is static and dated; the interactive bench compares any two of the 422 catalogue models.