AI Desk · Model Identity · Compare

Mistral: Mistral Small 4 · NVIDIA: Nemotron 3 Super

Identity similarity: very highExact tokenizer agreement · 50/50Exact agreement across every measured string is strong evidence of shared tokenizer lineage. It shows shared ancestry, and does not on its own establish that these are the same release.

Mistral: Mistral Small 4 recordNVIDIA: Nemotron 3 Super recordOpen in the interactive bench

Tokenizer fingerprint

50 of 50 mutually clean strings agree
Test stringMistral: Mistral Small 4NVIDIA: Nemotron 3 SuperVerdict
en-prose1414match
en-long1818match
spaces-2033match
spaces-6033match
tabs-2044match
newlines-2077match
mixed-ws77match
digits-999match
digits-121212match
digits-303030match
digits-sep1212match
float-long2222match
zh-common1616match
zh-long2020match
zh-rare2020match
ja-kana1111match
ja-kanji1313match
ko1313match
ru1212match
ar99match
he1010match
hi1414match
th1717match
el1313match
emoji-basic2020match
emoji-skin4040match
emoji-zwj-family1919match
emoji-zwj-x35757match
emoji-flags3232match
emoji-prof3535match
math3333match
boxdraw3030match
combining1414match
cjk-ext-b1616match
surrogates5353match
zalgo3636match
rtl-mix66match
py-code2525match
py-indent1515match
json2525match
html1616match
regex4949match
camel66match
snake77match
rare-word-x53636match
repeat-tok4141match
base643939match
hex1919match
url1919match
uuid3636match

Values are the marginal prompt-token cost of each string with the model's fixed overhead subtracted. An x marks a row excluded as corrupt for that model.

API surface

The declared serving contracts differ
FieldMistral: Mistral Small 4NVIDIA: Nemotron 3 SuperVerdict
Context length2621441000000× differs
Max output tokens16384× differs
Supported parametersfrequency_penalty, include_reasoning, max_tokens, presence_penalty, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_pfrequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p× differs
Defaults{"temperature":null,"top_p":null,"top_k":null,"frequency_penalty":null,"presence_penalty":null,"repetition_penalty":null}{"temperature":1,"top_p":0.95,"top_k":null,"frequency_penalty":null,"presence_penalty":null,"repetition_penalty":null}× differs
Reasoning contract{"mandatory":false,"default_enabled":false,"supported_efforts":["high","none"],"default_effort":"high"}{"mandatory":false,"default_enabled":true,"supports_max_tokens":true,"supported_efforts":["medium","low"],"default_effort":"medium"}× differs
Modalitiestext+image->texttext->text× differs

Method: tokenizer fingerprinting, v1, 50 strings. Raw data on the data page. This page is static and dated; the interactive bench compares any two of the 422 catalogue models.