AI Desk · Model Identity · Compare
DeepSeek: DeepSeek V3.1 · DeepSeek: R1 Distill Llama 70B
Identity similarity: very highExact tokenizer agreement · 50/50Exact agreement across every measured string is strong evidence of shared tokenizer lineage. It shows shared ancestry, and does not on its own establish that these are the same release.
DeepSeek: DeepSeek V3.1 recordDeepSeek: R1 Distill Llama 70B recordOpen in the interactive bench
Tokenizer fingerprint
50 of 50 mutually clean strings agree| Test string | DeepSeek: DeepSeek V3.1 | DeepSeek: R1 Distill Llama 70B | Verdict |
|---|---|---|---|
| en-prose | 14 | 14 | match |
| en-long | 16 | 16 | match |
| spaces-20 | 3 | 3 | match |
| spaces-60 | 3 | 3 | match |
| tabs-20 | 4 | 4 | match |
| newlines-20 | 4 | 4 | match |
| mixed-ws | 7 | 7 | match |
| digits-9 | 3 | 3 | match |
| digits-12 | 4 | 4 | match |
| digits-30 | 10 | 10 | match |
| digits-sep | 7 | 7 | match |
| float-long | 9 | 9 | match |
| zh-common | 9 | 9 | match |
| zh-long | 10 | 10 | match |
| zh-rare | 18 | 18 | match |
| ja-kana | 9 | 9 | match |
| ja-kanji | 11 | 11 | match |
| ko | 13 | 13 | match |
| ru | 13 | 13 | match |
| ar | 10 | 10 | match |
| he | 14 | 14 | match |
| hi | 18 | 18 | match |
| th | 11 | 11 | match |
| el | 16 | 16 | match |
| emoji-basic | 10 | 10 | match |
| emoji-skin | 14 | 14 | match |
| emoji-zwj-family | 11 | 11 | match |
| emoji-zwj-x3 | 33 | 33 | match |
| emoji-flags | 16 | 16 | match |
| emoji-prof | 19 | 19 | match |
| math | 22 | 22 | match |
| boxdraw | 27 | 27 | match |
| combining | 10 | 10 | match |
| cjk-ext-b | 16 | 16 | match |
| surrogates | 33 | 33 | match |
| zalgo | 36 | 36 | match |
| rtl-mix | 6 | 6 | match |
| py-code | 25 | 25 | match |
| py-indent | 15 | 15 | match |
| json | 20 | 20 | match |
| html | 16 | 16 | match |
| regex | 47 | 47 | match |
| camel | 7 | 7 | match |
| snake | 8 | 8 | match |
| rare-word-x5 | 25 | 25 | match |
| repeat-tok | 40 | 40 | match |
| base64 | 34 | 34 | match |
| hex | 12 | 12 | match |
| url | 17 | 17 | match |
| uuid | 27 | 27 | match |
Values are the marginal prompt-token cost of each string with the model's fixed overhead subtracted. An x marks a row excluded as corrupt for that model.
API surface
The declared serving contracts differ| Field | DeepSeek: DeepSeek V3.1 | DeepSeek: R1 Distill Llama 70B | Verdict |
|---|---|---|---|
| Context length | 163840 | 8192 | × differs |
| Max output tokens | 144900 | 7372 | × differs |
| Supported parameters | frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | frequency_penalty, include_reasoning, max_tokens, presence_penalty, reasoning, repetition_penalty, seed, stop, temperature, top_k, top_p | × differs |
| Defaults | {} | {} | ✓ match |
| Reasoning contract | {"mandatory":false} | {"mandatory":false} | ✓ match |
| Modalities | text->text | text->text | ✓ match |
Method: tokenizer fingerprinting, v1, 50 strings. This page is static and dated; the interactive bench compares any two of the 453 catalogue models.