DeepSeek: R1 Distill Llama 70B
deepseek/deepseek-r1-distill-llama-70b · DeepSeekPartial agreement- Tokenizer group (measured)
- DeepSeek
- Evidence confidence
- High
- Declared tokenizer tag
- Llama3
- Entered catalogue
- 23 January 2025
- Last tested
- 21 August 2026
- Measurement
- Measured
Identity statement
DeepSeek: R1 Distill Llama 70B is DeepSeek's model, and its measured tokenizer sits in DeepSeek's own signature group, shared with its sibling models and nobody else measured so far.
Verification note
Checked against the primary artefacts, 21 August 2026The mismatch was traced to the serving layer, not the label. The model's own repository tokenizer, run over the same 50 strings offline, is Llama 3's exactly (50 of 50, vocabulary 128,256), so the catalogue's Llama3 tag is right about the weights. The endpoint's reported token counts tell a different story: they match DeepSeek-V3's repository tokenizer on every string once the template boundary is folded, and the model's own tokenizer on only 18 of 50. The token accounting behind this endpoint uses a DeepSeek vocabulary the model does not read with, which sets what users are billed. Whether that is a billing-side tokenizer configuration or something else about the serving path cannot be established from token counts alone.
DeepSeek-R1-Distill-Llama-70B repository (config.json: vocab_size 128256, LlamaForCausalLM)DeepSeek-V3 repository (config.json: vocab_size 129280)
Evidence stack
Independent signals, weighed together| Signal | Result | Reference |
|---|
| Tokenizer signature | Exact match with 1 other model | tk_869dc1cf |
| API surface | No other entry shares this exact contract | api_28524539 |
| Context and output | 8,192 context · 8,192 max output | declared |
| Reasoning contract | Optional | declared |
| Serving providers observed | Novita | measured 21 August 2026 |
Shares this fingerprint
Exact signature first; template-boundary shifts of the same signature beneathClosest measured models
Agreement across the mutually clean test stringsThe measured fingerprint
Marginal prompt-token cost of each test string, grouped by scriptEnglish and whitespace
en-prose14
en-long16
spaces-203
spaces-603
tabs-204
newlines-204
mixed-ws7
Digits
digits-93
digits-124
digits-3010
digits-sep7
float-long9
CJK
zh-common9
zh-long10
zh-rare18
ja-kana9
ja-kanji11
ko13
Emoji
emoji-basic10
emoji-skin14
emoji-zwj-family11
emoji-zwj-x333
emoji-flags16
emoji-prof19
Rare Unicode
math22
boxdraw27
combining10
cjk-ext-b16
surrogates33
zalgo36
rtl-mix6
Code
py-code25
py-indent15
json20
html16
regex47
camel7
snake8
Repetition and encodings
rare-word-x525
repeat-tok40
base6434
hex12
url17
uuid27
Declared record
What the catalogue claims about this modelDeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), using outputs from [DeepSeek R1](/deepseek/deepseek-r1). The model combines advanced distillation techniques to achieve high performance across...
- Modalities
- text->text
- Declared tokenizer
- Llama3
- Prompt price
- $0.80 / M tokens
- Completion price
- $0.80 / M tokens
History
Every observation, kept as taken- 23 January 2025Enters the OpenRouter catalogue declared as Llama3.
- 21 August 2026Fingerprinted in the YFarmX catalogue sweep · 50 of 50 strings measured clean.
- 21 August 2026YFarmX assessment: consistent with the DeepSeek family, high confidence.