YFarmX logoYFarmX

AI Desk · AI Fingerprinting · Compare

IBM: Granite 4.1 8B · Microsoft: Phi 4

Identity similarity: highExact tokenizer agreement · 50/50Exact agreement across every measured string is one strong signal of shared tokenizer lineage. The published settings differ, so the second signal the methodology needs for very high is absent here; the table below shows where they part.

IBM: Granite 4.1 8B recordMicrosoft: Phi 4 recordOpen in the interactive bench

Tokenizer fingerprint

50 of 50 mutually clean strings agree
Test stringIBM: Granite 4.1 8BMicrosoft: Phi 4Verdict
en-prose1414match
en-long1616match
spaces-2033match
spaces-6033match
tabs-2033match
newlines-2044match
mixed-ws55match
digits-933match
digits-1244match
digits-301010match
digits-sep77match
float-long99match
zh-common1818match
zh-long2525match
zh-rare2323match
ja-kana1414match
ja-kanji1717match
ko1919match
ru1616match
ar2525match
he2424match
hi3232match
th2525match
el2929match
emoji-basic1010match
emoji-skin3030match
emoji-zwj-family1818match
emoji-zwj-x35454match
emoji-flags2424match
emoji-prof2929match
math3232match
boxdraw2828match
combining1414match
cjk-ext-b1313match
surrogates4040match
zalgo3636match
rtl-mix1111match
py-code2525match
py-indent1515match
json2020match
html1616match
regex4444match
camel55match
snake66match
rare-word-x53131match
repeat-tok2222match
base643535match
hex1111match
url1717match
uuid2727match

Values are the marginal prompt-token cost of each string with the model's fixed overhead subtracted. An x marks a row excluded as corrupt for that model.

API surface

The declared serving contracts differ
FieldIBM: Granite 4.1 8BMicrosoft: Phi 4Verdict
Context length13107216384× differs
Max output tokens––✓ match
Supported parametersfrequency_penalty, logprobs, max_tokens, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_pfrequency_penalty, logit_bias, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, top_k, top_p× differs
Defaults{"temperature":null,"top_p":null,"top_k":null,"frequency_penalty":null,"presence_penalty":null,"repetition_penalty":null}{}× differs
Reasoning contract––✓ match
Modalitiestext->texttext->text✓ match

IBM: Granite 4.1 8B: OpenRouter's listing shows 117,964 output tokens, which is 90% of the context window, the figure it carries where the provider declares no maximum. Microsoft: Phi 4: OpenRouter's listing shows 14,745 output tokens, which is 90% of the context window, the figure it carries where the provider declares no maximum.

Method: tokenizer fingerprinting, v1, 50 strings. This page is static and dated; the interactive bench compares any two of the 337 measured models.