YFarmX logoYFarmX

DeepSeek: R1

deepseek/deepseek-r1 · DeepSeekNot yet established
Tokenizer group (measured)
DeepSeek
Declared tokenizer tag
DeepSeek
Entered catalogue
20 January 2025
Last tested
5 October 2026
Evidence confidence
Moderate

Why moderate: near-exact agreement (47 of 50 strings) with deepseek/deepseek-r1-distill-llama-70b, short of a full exact signature.

Evidence
  • ✓ near-exact agreement (47/50) with deepseek/deepseek-r1-distill-llama-70b

Identity statement

DeepSeek: R1 is DeepSeek's model, and its measured tokenizer sits in DeepSeek's own signature group alongside its sibling models. No other maker's measured model shares it so far.

What we checked

Four different kinds of evidence, and what each one showed
We checkedWhat we foundWhere it came from
How it counts tokensCounted cleanly, and no other model in the catalogue counts the same waytk_d6c0b2be
How its API is set upA combination of settings no other listing in the catalogue usesapi_2283b8cd
How much it can read and writeReads up to 64,000 tokens at once, writes up to 16,000its own listing
Whether it thinks before answeringAlways reasons, and cannot be turned offits own listing
Who served our requestsNovitawe measured it, 5 October 2026

Closest measured models

Agreement across the mutually clean test strings

Click a row to open the full comparison.

The measured fingerprint

What each test string cost this model, in tokens

We sent DeepSeek: R1 fifty short pieces of text and recorded what each one cost it in tokens. The bars below are those costs. Two models built on the same tokenizer produce the same bars; a model built on a different one produces a different set, which is what makes this a fingerprint.

All 50 rows

English and whitespace

en-prose14
en-long16
spaces-203
spaces-603
tabs-204
newlines-204
mixed-ws7

Digits

digits-93
digits-124
digits-3010
digits-sep7
float-long9

CJK

zh-common9
zh-long10
zh-rare18
ja-kana9
ja-kanji11
ko13

Other scripts

ru13
ar10
he14
hi18
th11
el18

Emoji

emoji-basic10
emoji-skin14
emoji-zwj-family11
emoji-zwj-x335
emoji-flags16
emoji-prof21

Rare Unicode

math22
boxdraw27
combining10
cjk-ext-b16
surrogates33
zalgo36
rtl-mix6

Code

py-code25
py-indent15
json20
html16
regex47
camel7
snake8

Repetition and encodings

rare-word-x525
repeat-tok40
base6434
hex12
url17
uuid27

Overhead subtracted: 3 prompt tokens.

Declared record

What the catalogue claims about this model

DeepSeek R1 is here: Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass....

Modalities
text->text
Declared tokenizer
DeepSeek
Prompt price
$0.70 / M tokens
Completion price
$2.50 / M tokens

Supported parameters

frequency_penaltyinclude_reasoningmax_tokenspresence_penaltyreasoningrepetition_penaltyresponse_formatseedstoptemperaturetool_choicetoolstop_ktop_p

defaults: {"temperature":null,"top_p":null,"top_k":null,"frequency_penalty":null,"presence_penalty":null,"repetition_penalty":null}

Catalogue entryWeights on Hugging Face

History

Every observation, kept as taken
  • 20 January 2025Enters the OpenRouter catalogue declared as DeepSeek.
  • 5 October 2026Fingerprinted in the YFarmX catalogue sweep · 50 of 50 strings measured clean.
  • 5 October 2026YFarmX assessment: consistent with the DeepSeek family, moderate confidence.