AI Desk · AI Fingerprinting · Compare
IBM: Granite 4.1 8B · Microsoft: Phi 4
Identity similarity: highExact tokenizer agreement · 50/50Exact agreement across every measured string is one strong signal of shared tokenizer lineage. The published settings differ, so the second signal the methodology needs for very high is absent here; the table below shows where they part.
IBM: Granite 4.1 8B recordMicrosoft: Phi 4 recordOpen in the interactive bench
Tokenizer fingerprint
50 of 50 mutually clean strings agree| Test string | IBM: Granite 4.1 8B | Microsoft: Phi 4 | Verdict |
|---|---|---|---|
| en-prose | 14 | 14 | match |
| en-long | 16 | 16 | match |
| spaces-20 | 3 | 3 | match |
| spaces-60 | 3 | 3 | match |
| tabs-20 | 3 | 3 | match |
| newlines-20 | 4 | 4 | match |
| mixed-ws | 5 | 5 | match |
| digits-9 | 3 | 3 | match |
| digits-12 | 4 | 4 | match |
| digits-30 | 10 | 10 | match |
| digits-sep | 7 | 7 | match |
| float-long | 9 | 9 | match |
| zh-common | 18 | 18 | match |
| zh-long | 25 | 25 | match |
| zh-rare | 23 | 23 | match |
| ja-kana | 14 | 14 | match |
| ja-kanji | 17 | 17 | match |
| ko | 19 | 19 | match |
| ru | 16 | 16 | match |
| ar | 25 | 25 | match |
| he | 24 | 24 | match |
| hi | 32 | 32 | match |
| th | 25 | 25 | match |
| el | 29 | 29 | match |
| emoji-basic | 10 | 10 | match |
| emoji-skin | 30 | 30 | match |
| emoji-zwj-family | 18 | 18 | match |
| emoji-zwj-x3 | 54 | 54 | match |
| emoji-flags | 24 | 24 | match |
| emoji-prof | 29 | 29 | match |
| math | 32 | 32 | match |
| boxdraw | 28 | 28 | match |
| combining | 14 | 14 | match |
| cjk-ext-b | 13 | 13 | match |
| surrogates | 40 | 40 | match |
| zalgo | 36 | 36 | match |
| rtl-mix | 11 | 11 | match |
| py-code | 25 | 25 | match |
| py-indent | 15 | 15 | match |
| json | 20 | 20 | match |
| html | 16 | 16 | match |
| regex | 44 | 44 | match |
| camel | 5 | 5 | match |
| snake | 6 | 6 | match |
| rare-word-x5 | 31 | 31 | match |
| repeat-tok | 22 | 22 | match |
| base64 | 35 | 35 | match |
| hex | 11 | 11 | match |
| url | 17 | 17 | match |
| uuid | 27 | 27 | match |
Values are the marginal prompt-token cost of each string with the model's fixed overhead subtracted. An x marks a row excluded as corrupt for that model.
API surface
The declared serving contracts differ| Field | IBM: Granite 4.1 8B | Microsoft: Phi 4 | Verdict |
|---|---|---|---|
| Context length | 131072 | 16384 | × differs |
| Max output tokens | – | – | ✓ match |
| Supported parameters | frequency_penalty, logprobs, max_tokens, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p | frequency_penalty, logit_bias, max_tokens, min_p, presence_penalty, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, top_k, top_p | × differs |
| Defaults | {"temperature":null,"top_p":null,"top_k":null,"frequency_penalty":null,"presence_penalty":null,"repetition_penalty":null} | {} | × differs |
| Reasoning contract | – | – | ✓ match |
| Modalities | text->text | text->text | ✓ match |
IBM: Granite 4.1 8B: OpenRouter's listing shows 117,964 output tokens, which is 90% of the context window, the figure it carries where the provider declares no maximum. Microsoft: Phi 4: OpenRouter's listing shows 14,745 output tokens, which is 90% of the context window, the figure it carries where the provider declares no maximum.
Method: tokenizer fingerprinting, v1, 50 strings. This page is static and dated; the interactive bench compares any two of the 337 measured models.