YFarmX logoYFarmX

AI Desk · AI Fingerprinting · Compare

Mistral: Mistral Large 4 · OpenAI: GPT-5 Nano

Identity similarity: lowDivergent fingerprints · 31/50These models tokenise differently. The rows that part here are the English prose, whitespace, CJK, non-Latin script and emoji strings. Divergence on high-information strings separates unrelated vocabularies within a handful of samples.

Mistral: Mistral Large 4 recordOpenAI: GPT-5 Nano recordOpen in the interactive bench

Tokenizer fingerprint

31 of 50 mutually clean strings agree
Test stringMistral: Mistral Large 4OpenAI: GPT-5 NanoVerdict
en-prose1414match
en-long1817differs
spaces-2033match
spaces-6033match
tabs-2053differs
newlines-2044match
mixed-ws76differs
digits-933match
digits-1244match
digits-301010match
digits-sep77match
float-long99match
zh-common1212match
zh-long1617differs
zh-rare1920differs
ja-kana1110differs
ja-kanji1313match
ko1111match
ru1111match
ar911differs
he910differs
hi1111match
th1811differs
el1414match
emoji-basic108differs
emoji-skin2013differs
emoji-zwj-family1111match
emoji-zwj-x33333match
emoji-flags1616match
emoji-prof2120differs
math2731differs
boxdraw1627differs
combining1111match
cjk-ext-b1313match
surrogates4140differs
zalgo3435differs
rtl-mix55match
py-code2525match
py-indent1515match
json2020match
html1616match
regex4344differs
camel66match
snake66match
rare-word-x53631differs
repeat-tok2222match
base643334differs
hex1111match
url1717match
uuid2727match

Values are the marginal prompt-token cost of each string with the model's fixed overhead subtracted. An x marks a row excluded as corrupt for that model.

API surface

The declared serving contracts differ
FieldMistral: Mistral Large 4OpenAI: GPT-5 NanoVerdict
Context length1048576400000× differs
Max output tokens262144128000× differs
Supported parametersfrequency_penalty, include_reasoning, logprobs, max_tokens, presence_penalty, reasoning, reasoning_effort, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_logprobs, top_pinclude_reasoning, max_completion_tokens, max_tokens, reasoning, reasoning_effort, response_format, seed, structured_outputs, tool_choice, tools, verbosity× differs
Defaults{}{"temperature":null,"top_p":null,"top_k":null,"frequency_penalty":null,"presence_penalty":null,"repetition_penalty":null}× differs
Reasoning contract{"mandatory":false,"default_enabled":true,"supported_efforts":["high","none"],"default_effort":"high"}{"mandatory":true,"supported_efforts":["high","medium","low","minimal"],"default_effort":"medium"}× differs
Modalitiestext+image->texttext+image+file->text× differs

Method: tokenizer fingerprinting, v1, 50 strings. This page is static and dated; the interactive bench compares any two of the 337 measured models.