YFarmX logoYFarmX

DeepSeek: R1 Distill Llama 70B

deepseek/deepseek-r1-distill-llama-70b · DeepSeekPartial agreement
Tokenizer group (measured)
DeepSeek
Declared tokenizer tag
Llama3
Entered catalogue
23 January 2025
Last tested
21 August 2026
Evidence confidence
High

Why high: an exact fingerprint shared with 2 models that declare the same family, without a second signal of a different kind yet.

Evidence
  • ✓ exact tokenizer signature shared with 2 models

Identity statement

DeepSeek: R1 Distill Llama 70B is DeepSeek's model, and its measured tokenizer sits in DeepSeek's own signature group alongside its sibling models. Measured models from StepFun carry the same vocabulary, which reads as reuse of DeepSeek's tokenizer line and questions nobody's authorship; they are listed under "Shares this fingerprint" below.

One model, two answers

The split this record documents
The weights' own tokenizerLlama 3, verified offline against the model's repository: 50 of 50 strings, vocabulary 128,256
The endpoint's token accountingDeepSeek's vocabulary: the billed counts match DeepSeek-V3's tokenizer on every string once the one-token template boundary is folded, and the model's own tokenizer on only 18 of 50
Where the two meetThe bill follows the serving layer rather than the weights, so the catalogue's Llama3 label stays right about the model itself
Not established by thisThat the weights run a DeepSeek vocabulary; the repository tokenizer says they do not

Verification note

Checked against the primary artefacts, 21 August 2026

The mismatch was traced to the serving layer, not the label. The model's own repository tokenizer, run over the same 50 strings offline, is Llama 3's exactly (50 of 50, vocabulary 128,256), so the catalogue's Llama3 tag is right about the weights. The endpoint's reported token counts tell a different story: they match DeepSeek-V3's repository tokenizer on every string once the template boundary is folded, and the model's own tokenizer on only 18 of 50. The token accounting behind this endpoint uses a DeepSeek vocabulary the model does not read with, which sets what users are billed. Whether that is a billing-side tokenizer configuration or something else about the serving path cannot be established from token counts alone.

DeepSeek-R1-Distill-Llama-70B repository (config.json: vocab_size 128256, LlamaForCausalLM)DeepSeek-V3 repository (config.json: vocab_size 129280)

What we checked

Four different kinds of evidence, and what each one showed
We checkedWhat we foundWhere it came from
How it counts tokensCounts every one of the 50 test strings exactly like 2 other modelstk_869dc1cf
How its API is set upA combination of settings no other listing in the catalogue usesapi_3560b58d
How much it can read and writeReads up to 8,192 tokens at onceOpenRouter's listing shows 7,372 output tokens, which is 90% of the context window, the figure it carries where the provider declares no maximum. DeepSeek's own API reference caps max_tokens at 384K, 393,216 tokens, for deepseek-flash and deepseek-v4-pro.its own listing
Whether it thinks before answeringReasons when asked toits own listing
Who served our requestsNovitawe measured it, 21 August 2026

Shares this fingerprint

Exact signature first; template-boundary shifts of the same signature beneath

Closest measured models

Agreement across the mutually clean test strings

Click a row to open the full comparison.

The measured fingerprint

What each test string cost this model, in tokens

We sent DeepSeek: R1 Distill Llama 70B fifty short pieces of text and recorded what each one cost it in tokens. The bars below are those costs. Two models built on the same tokenizer produce the same bars; a model built on a different one produces a different set, which is what makes this a fingerprint.

All 50 rows

English and whitespace

en-prose14
en-long16
spaces-203
spaces-603
tabs-204
newlines-204
mixed-ws7

Digits

digits-93
digits-124
digits-3010
digits-sep7
float-long9

CJK

zh-common9
zh-long10
zh-rare18
ja-kana9
ja-kanji11
ko13

Other scripts

ru13
ar10
he14
hi18
th11
el16

Emoji

emoji-basic10
emoji-skin14
emoji-zwj-family11
emoji-zwj-x333
emoji-flags16
emoji-prof19

Rare Unicode

math22
boxdraw27
combining10
cjk-ext-b16
surrogates33
zalgo36
rtl-mix6

Code

py-code25
py-indent15
json20
html16
regex47
camel7
snake8

Repetition and encodings

rare-word-x525
repeat-tok40
base6434
hex12
url17
uuid27

Overhead subtracted: 3 prompt tokens.

Declared record

What the catalogue claims about this model

DeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), using outputs from [DeepSeek R1](/deepseek/deepseek-r1). The model combines advanced distillation techniques to achieve high performance across...

Modalities
text->text
Declared tokenizer
Llama3
Prompt price
$0.80 / M tokens
Completion price
$0.80 / M tokens

Supported parameters

frequency_penaltyinclude_reasoningmax_tokenspresence_penaltyreasoningrepetition_penaltyseedstoptemperaturetop_ktop_p

defaults: {}

Catalogue entryWeights on Hugging Face

History

Every observation, kept as taken
  • 23 January 2025Enters the OpenRouter catalogue declared as Llama3.
  • 21 August 2026Fingerprinted in the YFarmX catalogue sweep · 50 of 50 strings measured clean.
  • 21 August 2026YFarmX assessment: consistent with the DeepSeek family, high confidence.

* DeepSeek: R1 Distill Llama 70B no longer appears in the OpenRouter catalogue, first absent from the snapshot of 5 October 2026. The record stays because the measurements above were real when they were taken and this address has been published.