AI Desk · AI Fingerprinting · Compare
Perplexity: Sonar Pro · Sao10K: Llama 3.3 Euryale 70B
Identity similarity: highExact tokenizer agreement · 50/50Exact agreement across every measured string is one strong signal of shared tokenizer lineage. The published settings differ, so the second signal the methodology needs for very high is absent here; the table below shows where they part.
Perplexity: Sonar Pro recordSao10K: Llama 3.3 Euryale 70B recordOpen in the interactive bench
Tokenizer fingerprint
50 of 50 mutually clean strings agree| Test string | Perplexity: Sonar Pro | Sao10K: Llama 3.3 Euryale 70B | Verdict |
|---|---|---|---|
| en-prose | 14 | 14 | match |
| en-long | 16 | 16 | match |
| spaces-20 | 3 | 3 | match |
| spaces-60 | 3 | 3 | match |
| tabs-20 | 3 | 3 | match |
| newlines-20 | 4 | 4 | match |
| mixed-ws | 5 | 5 | match |
| digits-9 | 3 | 3 | match |
| digits-12 | 4 | 4 | match |
| digits-30 | 10 | 10 | match |
| digits-sep | 7 | 7 | match |
| float-long | 9 | 9 | match |
| zh-common | 12 | 12 | match |
| zh-long | 18 | 18 | match |
| zh-rare | 18 | 18 | match |
| ja-kana | 11 | 11 | match |
| ja-kanji | 12 | 12 | match |
| ko | 14 | 14 | match |
| ru | 14 | 14 | match |
| ar | 12 | 12 | match |
| he | 24 | 24 | match |
| hi | 15 | 15 | match |
| th | 14 | 14 | match |
| el | 13 | 13 | match |
| emoji-basic | 10 | 10 | match |
| emoji-skin | 30 | 30 | match |
| emoji-zwj-family | 15 | 15 | match |
| emoji-zwj-x3 | 45 | 45 | match |
| emoji-flags | 24 | 24 | match |
| emoji-prof | 26 | 26 | match |
| math | 31 | 31 | match |
| boxdraw | 23 | 23 | match |
| combining | 11 | 11 | match |
| cjk-ext-b | 13 | 13 | match |
| surrogates | 40 | 40 | match |
| zalgo | 35 | 35 | match |
| rtl-mix | 8 | 8 | match |
| py-code | 25 | 25 | match |
| py-indent | 15 | 15 | match |
| json | 20 | 20 | match |
| html | 16 | 16 | match |
| regex | 44 | 44 | match |
| camel | 5 | 5 | match |
| snake | 6 | 6 | match |
| rare-word-x5 | 30 | 30 | match |
| repeat-tok | 21 | 21 | match |
| base64 | 35 | 35 | match |
| hex | 11 | 11 | match |
| url | 17 | 17 | match |
| uuid | 27 | 27 | match |
Values are the marginal prompt-token cost of each string with the model's fixed overhead subtracted. An x marks a row excluded as corrupt for that model.
API surface
The declared serving contracts differ| Field | Perplexity: Sonar Pro | Sao10K: Llama 3.3 Euryale 70B | Verdict |
|---|---|---|---|
| Context length | 200000 | 131072 | × differs |
| Max output tokens | 8000 | 16384 | × differs |
| Supported parameters | frequency_penalty, max_tokens, presence_penalty, temperature, top_k, top_p, web_search_options | frequency_penalty, logprobs, max_tokens, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, top_logprobs, top_p | × differs |
| Defaults | {} | {} | ✓ match |
| Reasoning contract | – | – | ✓ match |
| Modalities | text+image->text | text->text | × differs |
Method: tokenizer fingerprinting, v1, 50 strings. This page is static and dated; the interactive bench compares any two of the 337 measured models.