DeepSeek: DeepSeek V3.1

deepseek/deepseek-chat-v3.1 · DeepSeekDeclared and measured agree
Tokenizer group (measured)
DeepSeek
Evidence confidence
High
Why high: an exact fingerprint shared with 1 model that declares the same family, without a second signal of a different kind yet.
Declared tokenizer tag
DeepSeek
Entered catalogue
21 August 2025
Last tested
21 August 2026
Measurement
Measured
Evidence
  • exact tokenizer signature shared with 1 model

Identity statement

DeepSeek: DeepSeek V3.1 is DeepSeek's model, and its measured tokenizer sits in DeepSeek's own signature group alongside its sibling models. Measured models from StepFun carry the same vocabulary, which reads as reuse of DeepSeek's tokenizer line and questions nobody's authorship; they are listed under "Shares this fingerprint" below.

Evidence stack

Signals of different kinds, weighed together
SignalResultReference
Tokenizer signatureExact match with 1 other modeltk_869dc1cf
API surfaceNo other entry shares this exact contractapi_13422a94
Context and output163,840 context · 32,768 max outputdeclared
Reasoning contractOptionaldeclared
Serving providers observedSiliconFlow, Googlemeasured 21 August 2026

Shares this fingerprint

Exact signature first; template-boundary shifts of the same signature beneath

Closest measured models

Agreement across the mutually clean test strings

Click a row to open the full comparison.

The measured fingerprint

Marginal prompt-token cost of each test string, grouped by script
All 50 rows

English and whitespace

en-prose14
en-long16
spaces-203
spaces-603
tabs-204
newlines-204
mixed-ws7

Digits

digits-93
digits-124
digits-3010
digits-sep7
float-long9

CJK

zh-common9
zh-long10
zh-rare18
ja-kana9
ja-kanji11
ko13

Other scripts

ru13
ar10
he14
hi18
th11
el16

Emoji

emoji-basic10
emoji-skin14
emoji-zwj-family11
emoji-zwj-x333
emoji-flags16
emoji-prof19

Rare Unicode

math22
boxdraw27
combining10
cjk-ext-b16
surrogates33
zalgo36
rtl-mix6

Code

py-code25
py-indent15
json20
html16
regex47
camel7
snake8

Repetition and encodings

rare-word-x525
repeat-tok40
base6434
hex12
url17
uuid27

An x marks a row excluded as corrupt (caching or provider interference detected during measurement). Overhead subtracted: 7 prompt tokens.

Declared record

What the catalogue claims about this model

DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context...

Modalities
text->text
Declared tokenizer
DeepSeek
Prompt price
$0.25 / M tokens
Completion price
$0.95 / M tokens

Supported parameters

frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p

defaults: {}

Catalogue entryWeights on Hugging Face

History

Every observation, kept as taken
  • 21 August 2025Enters the OpenRouter catalogue declared as DeepSeek.
  • 21 August 2026Fingerprinted in the YFarmX catalogue sweep · 50 of 50 strings measured clean.
  • 21 August 2026YFarmX assessment: consistent with the DeepSeek family, high confidence.

Cite this page as the evidence record for DeepSeek: DeepSeek V3.1: the URL is stable, measurements are dated, and revisions append to the history above. Method: tokenizer fingerprinting, v1, 50 strings. The raw data behind every figure is on the data page.