YFarmX logoYFarmX

AI model fingerprints

counting tokens across 333 models to see which ones share a lineage

54 min readLarge Language ModelsLast updated:

Editorial collage headed MODEL FINGERPRINTS with the subtitle 333 models measured, 68 signatures: one large inked thumbprint on newsprint whose ridges are formed from printed digits, a torn spec sheet carrying two rows of token counts, a pencilled note reading 465 catalogued, a blank index card, and pinned paper chips showing the OpenAI, Anthropic, Google and Qwen logos

Key facts

465OpenRouter, 5 Oct 2026
Models catalogued
333310 complete signatures
Measured
68grouping into 55 families
Distinct fingerprints
2both confirmed by their makers
Stealth models named

Every AI model chops text into tokens, and the chopping rules differ from one model family to the next. Feed the same 50 awkward strings to 333 models, record how many tokens each one charges for, and models that share a tokenizer answer identically. That is enough to sort a catalogue into families, to place models whose makers say nothing about what is inside them, and, measured again weeks later, to say whether the model behind a name has held. It is not enough to name an individual model, and this page is careful about the difference. Twice now the method has named a stealth model anyway, and both times the maker confirmed the call.

New attribution · measured 23 September 2026 Space Bunny Alpha is a MiniMax model. OpenRouter's anonymous stealth model charged the same token count as MiniMax's models on all 50 test strings, measured the day it was listed. Our read is the next MiniMax after M3. Read the case ↓ Try it yourself Compare any two models Pick two models and see their token counts side by side, string by string, with a verdict on whether they share a tokenizer.

What this page is

Every AI model reads text by chopping it into pieces called tokens. The rules for chopping live in a part of the model called the tokenizer, and they differ from one family to the next. So if two models chop text in exactly the same way, across strings chosen to make those rules disagree, they almost certainly share foundations.

Between 20 August and 5 October 2026, in three sweeps, we sent the same 50 strings to every model we could reach on OpenRouter, and recorded how many tokens each one charged. On 5 October we sent them again to the 66 most-used models on record, to see whether the readings held. This page reports what came back. It is a measurement, run at stated times on a stated catalogue, and everything below is traceable to it.

Read this first

A shared tokenizer shows a shared lineage: the same foundations, the same tooling, often the same base model. It does not show that one model is a copy of another. A fine-tune keeps the tokenizer of the model it was tuned from, so a family here can hold genuinely different models built on one starting point.

The idea in one minute

Ask a model to read the letter “a”, then twenty spaces, then the letter “b”. One model charges you three tokens for that. Another charges four. Neither is wrong. They were built with different rules about how to bundle runs of whitespace, and those rules are baked in long before the model is trained.

Repeat that with a family emoji, a line of Chinese, a twelve-digit number, a run of tabs, a mathematical expression. Each string is chosen because tokenizers tend to disagree about it. After fifty of them you have a column of numbers that acts like a fingerprint of the tokenizer.

The measurement is cheap, it needs no special access, and the model cannot decline to answer, because the token count is on the bill.

Called before the reveal · confirmed 18 September 2026 Union Alpha is Pareto 26.9. The stealth model named itself two days after listing, confirming the attribution we published while it was anonymous: Circuit & Chisel's Pareto, called from its pricing arithmetic, its paper trail and the three tokenizers behind its one endpoint. Read the case ↓

Seeing a fingerprint

These are real counts from the sweep. Read down a column: where two models show the same number on every row, their tokenizers agree.

Model a + 20 spaces + b 20 tabs Family emoji Chinese phrase 12 digits
OpenAI GPT-5-nano3311124
Meta Llama 3.3 70B3315124
Anthropic Claude Sonnet 54414184
Google Gemini 3.7 Flash3371012
DeepSeek V3.2341194
Z.ai GLM-5.333797
Mistral Medium 3.534191612
Alibaba Qwen3 Max Thinking33101012
Five of the fifty strings, and the tokens each model charged for them. The first two columns separate almost nothing; the family emoji separates nearly everything. That is why the set has fifty strings rather than five.

Notice how little the first column tells you. Seven of these eight models charge three tokens for the letter “a”, twenty spaces and the letter “b”. The emoji column splits the same eight models into six different answers. Most of the work is done by a handful of strings, which is exactly what a set of fifty is for: enough awkward cases that models which merely look alike come apart.

Fireworks' Ember-1 is built on Kimi K3

Research note · 24 September 2026

A dated supplement, measured on the day the model listed. It adds one record to the dataset, replaces the Kimi K3 reading with a cleaner one, and leaves the historical sweep totals below unchanged.

Fireworks, the AI inference company, listed fireworks/ember-1 on OpenRouter at 00:07 UTC on 24 September 2026 and describes it as a reasoning model “built on Kimi K3”, Moonshot AI’s open-weight model. Fireworks’ launch post says Ember-1 “delivers Kimi K3’s quality with 40% fewer tokens”. We measured it the same day, and its token counts match Kimi K3’s on all 50 strings.

OpenRouter's model page for Fireworks Ember-1, showing the identifier fireworks slash ember-1, a description calling it a specialised reasoning model from Fireworks Research built on Kimi K3 that uses roughly 40% fewer tokens than the base model, text and image in and text out, a price of $3 in and $15 out per million tokens, a 1M context and a release date of 24 September 2026
The listing on 24 September 2026. Source: OpenRouter.

It matches Kimi K3 on all 50 strings

Ember-1 charged the same number of tokens as Kimi K3 on every one of the 50 test strings, measured on 24 September 2026 with both models served by the same provider, Fireworks. Scored against the other 288 records in the dataset, every model that reaches 50 of 50 is a Moonshot Kimi.

Model Maker Agreement
Kimi K3, measured on Fireworks the same day Moonshot AI 50 of 50
Kimi K2.6, K2.7 Code and K2 Thinking Moonshot AI 50 of 50
Hy3 Tencent 35 of 50
MiniMax M3 MiniMax 29 of 50
Median of all 288 14 of 50

The two models also agree on all 18 special-token probes, the strings that spell out other families’ control tokens, and both add the same 88 tokens of chat template to every request. A model trained on top of another keeps its base’s vocabulary, so this is the reading a Kimi K3 derivative gives. Its record page carries the full set of counts.

Bar chart headed Fireworks' Ember-1 runs on Kimi K3's tokenizer. Kimi K3, measured the same day on the same host, matches 50 of 50 token counts; Kimi K2.6, K2.7 Code and K2 Thinking 50 of 50; Tencent Hy3 35; MiniMax M3 29; the typical model 14. A footer notes the two also match on 18 of 18 special-token checks and share an 88-token chat template
Exact token-count agreement between Ember-1 and the closest models in the dataset, out of 50 strings.

What Fireworks trained on top

Fireworks trained Kimi K3 to reason in fewer tokens. Its launch post, dated 23 September 2026, reports more than 50 training experiments and over 200 evaluations, and says Kimi K3’s reasoning “could be shortened by 35–50% without sacrificing accuracy” across seven benchmarks and two customers’ production traffic. Those quality figures are Fireworks’ own evaluations. The price is the same as Kimi K3 from Moonshot on OpenRouter: $3 per million input tokens and $15 per million output, with the same 1,048,576-token context, so the saving Fireworks describes comes from writing fewer tokens per answer.

Every request carries a hidden “think at max” line

Ember-1 reads 89 tokens when it is sent a single letter, and 67 of them are an instruction the caller never sees. Kimi K3’s chat code, which Moonshot AI publishes on Hugging Face, inserts a system message whenever a reasoning effort is set, ending “Now the system is invoked with thinking_effort=max”. We ran that code and K3’s tokenizer on our own machine on 24 September 2026: they predict Ember-1’s 89-token baseline exactly, and all 50 of its string counts. OpenRouter’s default effort for Ember-1 is max, so the prompt still asks for the longest thinking, and the shorter reasoning Fireworks sells comes from the retrained weights.

Part of a one-letter request Tokens
Hidden effort instruction 67
Chat wrapping around each message 21
The letter itself 1
Total, as Ember-1 counts it 89
Chart headed Ember-1 is Kimi K3, retrained to think less. A bar splits the 89 tokens Ember-1 reads for a one-letter message into 67 for a hidden think-at-max instruction, 21 for chat wrapping and 1 for the letter. A second chart shows Fireworks' figures for output tokens per customer coding task: Kimi K3 49.3 thousand at a score of 0.751, Ember-1 29.9 thousand at 0.753
The 89-token split comes from Moonshot's published Kimi K3 code, run offline. The token and score figures are Fireworks' own, from its launch post.

The engine underneath is Kimi K3’s

Ember-1 inherits the architecture Moonshot publishes in Kimi K3’s configuration file, a hybrid that mixes fast linear-attention layers with full-attention ones.

Part Kimi K3, and so Ember-1
Layers 93: 69 Kimi Delta Attention (linear) and 24 full attention
Experts 896, with 16 used for each token plus 2 shared
Context 1,048,576 tokens
Weight format MXFP4

Kimi K3’s vocabulary file is byte-identical to the ones in Kimi K2, K2.5 and K2.6, so Moonshot has kept one tokenizer across four releases. Fireworks serves Ember-1 as checkpoint ember-1-20260923 beside its own Kimi K3 endpoint, with the same limits, settings and price, and calls it a two-week research preview. The weights are private.

We fixed our Kimi K3 record

Measuring Kimi K3 as the control corrected its record in this dataset. The 21 August sweep read Kimi K3 while OpenRouter spread the calls across seven providers, and eight rows came back exactly 3 tokens low, the gap between two providers’ fixed overheads. Measured on one provider on 24 September 2026, Kimi K3 matches Kimi K2.6, K2.7 Code and K2 Thinking on all 50 strings: Moonshot kept one vocabulary from Kimi K2 to K3. Kimi K3’s record now carries the single-provider reading and a Moonshot AI family placement. Two runs of 138 calls cost $0.15 in total, and the raw measurements are in the desk's research files for Ember-1 .

Space Bunny Alpha is a MiniMax model

Research note · 23 September 2026

A dated supplement, measured on the day the model listed. It adds one record to the dataset and leaves the historical sweep totals below unchanged.

OpenRouter listed stealth/space-bunny-alpha at 14:48 UTC on 23 September 2026: free, a 1M-token context, text, image and video in, and no maker named. We measured it the same afternoon. Its token counts match MiniMax’s on all 50 strings, the same score our reproducibility control returns when a model is measured against itself.

OpenRouter's model page for Space Bunny Alpha, showing the identifier stealth slash space-bunny-alpha, a description calling it an anonymous large model with fast inference and native multimodal input, modalities of text, image and video in and text out, a price of Free, a 1M context and a release date of 23 September 2026
The listing on 23 September 2026. Source: OpenRouter.

It matches MiniMax on all 50 strings

Space Bunny Alpha charged the same number of tokens as MiniMax’s models on every one of the 50 test strings, measured on 23 September 2026 and scored against 287 records in the dataset. The next closest family reaches 29.

Family Closest model Agreement
MiniMax all eight, MiniMax-01 to M3 50 of 50
Moonshot Kimi Kimi K2.7 Code 29 of 50
xAI Grok Grok Build 0.1 27 of 50
Tencent HY-MT2 1.8B 27 of 50
Meta Llama 4 Llama 4 Maverick 25 of 50
OpenAI GPT-5.6 Luna 24 of 50
Median of all 287 11 of 50

A shared vocabulary reads above 90% in this corpus, and a complete match is the strongest reading the method gives. The dataset’s own clustering agrees without being asked: it filed Space Bunny Alpha into the existing MiniMax signature, which now holds nine models, and opened no new group for it. Its record page carries the full set of counts.

Bar chart headed Space Bunny Alpha is a MiniMax model. MiniMax, all eight models from 01 to M3, matches 50 of 50 token counts; Moonshot Kimi K2 29; xAI Grok 4 27; Meta Llama 4 25; the typical model 11. A footer gives our read: the next MiniMax after M3
Exact token-count agreement between Space Bunny Alpha and the closest model in each family, out of 50 strings.

The run was clean

Two full passes on 23 September 2026 returned 50 clean rows each, and they agree on every count. The fixed overhead was 143 tokens before the strings and 143 after, with one provider, Stealth, answering every call. Union Alpha’s three tokenizers showed up as disagreements between repeated readings of the same request, so an endpoint of that kind would have shown itself here. This one returns one vocabulary. The model is free, so the 138 calls cost nothing.

Which MiniMax it could be

MiniMax has kept one vocabulary across all eight of its models on OpenRouter, from MiniMax-01 in January 2025 to M3 on 31 May 2026, so the counts name the family and stop there. The listing narrows it further. M3 is the only MiniMax model on OpenRouter that takes text, images and video with a 1M-token context, and Space Bunny Alpha has exactly that shape. Our read is a MiniMax model newer than M3. That read comes from the listing; the token counts carry the family.

It calls itself GPT-5

Asked who built it, Space Bunny Alpha answered “I was created by OpenAI, and my model name is GPT-5.” Models trained partly on other models’ output often repeat that output’s description of itself, and the counts settle the question: the closest OpenAI model scores 24 of 50, less than half, and Space Bunny Alpha sits inside MiniMax’s signature. Asked more carefully, it declined to name a maker at all.

Union Alpha is Pareto 26.9

On 16 September 2026 OpenRouter listed an anonymous model, stealth/union-alpha: free, a 262,144-token context, text and images, no owner named. Two days later the account behind it gave the answer itself: Pareto 26.9, the blended system Circuit & Chisel sells through its Unbiased platform. Between those two dates sit the measurements, the arithmetic and the attribution this page published while the endpoint was still anonymous. The working stays below, in the order it happened.

OpenRouter's model page for Union Alpha, showing the identifier stealth slash union-alpha, modalities of text and image in and text out, a price of Free, a 262K context and a release date of 16 September 2026
OpenRouter's listing on the day it appeared: free, 262K of context, text and image in, and a provider called stealth. Captured 17 September 2026 from openrouter.ai.
Circuit & Chisel's logo, two coral half-discs facing each other Confirmed by the maker · 18 September 2026 Circuit & Chisel's Pareto 26.9, sold through Unbiased

On 18 September the Union Alpha account on X posted "You found us", and named the stealth model on OpenRouter, OpenCode and Cloudflare: Pareto 26.9 from The Unbiased Co, Circuit & Chisel's blended system, sold through its Unbiased platform. The reveal described what had been answering all along: several models, open and frontier, work the same task, a harness checks the work, and a stronger model is called in where the task needs one.

The account said the stealth listing existed to gather real usage ahead of a 26.10 release the following month, that demand reached billions of tokens a minute within hours, and that AWS helped triple capacity overnight. It put the paid Pareto 26.9 card at:

  • $2.50per million input tokens
  • $7.50per million output tokens
  • $0.25per million cached input
Called before the reveal: the pricing arithmetic, the description trail and the three fingerprints below are the case as we published it while the model was anonymous, with every reading left in view in the order we took it.
  • 16 SeptemberListed anonymously. OpenRouter carries stealth/union-alpha: free, 262K of context, text and images. Cloudflare documents it as a blended model, with five worked cost examples.
  • 17 SeptemberThree tokenizers surface. Early fingerprints had split between Qwen and Llama 3, with GLM reported on images. Differential measurement finds all three behind the one endpoint: Llama 3 and Qwen answer on text, the GLM family with images and tools.
  • 17 SeptemberWe name Pareto. At 04:51 UTC we publish the attribution on X: Circuit & Chisel's Pareto, from the pricing arithmetic and Cloudflare's own first description.
  • 17 SeptemberThe paper trail thins. Cloudflare rewrites its catalogue entry, and Circuit & Chisel's release-note repository stops resolving.
  • Before the revealAn independent trace agrees. Palmer matches Union Alpha's completion IDs to Unbiased's platform and reports the same vLLM fleet answering both.
  • 18 SeptemberConfirmed. The Union Alpha account posts "You found us" and names Pareto 26.9.

We called it a day before the reveal

The attribution went up while the model was anonymous: on this page, and in a post on X at 04:51 UTC on 17 September, the earliest public post we can verify connecting Union Alpha to Pareto and Circuit & Chisel. The reveal came the next day.

Our public call · 17 September 2026, 04:51 UTC Union Alpha, called as Circuit & Chisel's Pareto, a day before the reveal View the post ↗

The call rested on three legs, each one set out below. Cloudflare’s day-one entry described a blended system in nearly the same words Circuit & Chisel uses for Pareto. Five worked cost examples in the same entry solve to Pareto’s then-published rates, to the cent. And the endpoint carried three tokenizer fingerprints, which is what several models behind one door look like from the outside. The reveal confirmed the identity and the architecture both.

Three tokenizers behind one endpoint

Measuring the endpoint’s tokenizer returned three different answers depending on what the request carried, which is a separate finding and is set out further down. Once the system had a name, the three answers read as its architecture showing through: a blend carries several tokenizers because it carries several models.

Cloudflare’s developer documentation gave the first reason to look at a blend. From 16 September its catalogue entry described Union Alpha as “a blended AI model that engages multiple language models in parallel for each request, synthesizes one answer, and supports text and vision inputs through a single API response”, and tagged it Blended Model.

That description was removed the next day. On 17 September, commit 333aed8, the entry was rewritten to “a multimodal model designed for research, coding, and agentic workflows”. Three separate fields carrying the same claim went in that one commit: the description, the Blended Model tag, and an Architecture: Blended model metadata entry. Its pull request summary was left as an empty template, so the record does not say why.

Both revisions are frozen in our research folder, so the quotes stay checkable. The measurements below were taken before either edit and do not depend on a word of Cloudflare’s wording.

Cloudflare's developer documentation page for stealth slash union-alpha, tagged Third-party, carrying the replacement description that calls it a multimodal model designed for research, coding and agentic workflows, above a TypeScript usage example
The same page after the edit. The Third-party tag stays, and the blended-model description is gone. Captured 17 September 2026 from developers.cloudflare.com.
Table headed Union Alpha has three fingerprints, showing that plain text requests return counts matching the Llama 3 family or the Qwen family while requests carrying an image or a tool declaration return counts matching the GLM family every time
The same endpoint, measured under different request shapes.

What we sent, and what came back

Only one number is read here, usage.prompt_tokens, with the reply capped at a single token. That field measures the input. Nothing the model generates can reach it.

What the request carried Which vocabulary answered Steady?
Plain text Llama 3 family, or Qwen family Flips, roughly half and half
A system prompt Llama 3 family, or Qwen family Flips
A prior assistant turn Llama 3 family, or Qwen family Flips
Code, or a very long context Llama 3 family, or Qwen family Flips
An image GLM family Same answer every time
A tool declaration GLM family Same answer every time

Attach an image and the endpoint stops varying: fourteen identical readings out of fourteen, and again under tool declarations. Send plain text and it alternates between two other vocabularies.

How you can tell a real tokenizer from a bad number

A tokenizer is exactly proportional to what you send it. So send a string, then send three copies of it, and subtract. Everything fixed cancels, including the chat template, the image and the tool schema, and what survives is the cost of the content alone.

Thirty digits are the sharpest single test on the text paths. The Llama 3 vocabulary bundles digits in threes, so thirty cost it ten. The Qwen vocabulary splits them, so the same thirty cost thirty. One endpoint returns both, from requests identical down to the byte.

Digits cannot separate Llama 3 from GLM, because both charge ten, which is why the image path needs strings chosen to split them. Our first reading of that path used digits and landed in the Llama 3 group. Re-running with eight strings that do separate the two, chosen from emoji, Thai, Hindi and combining marks, the image path resolves to the GLM family on seven of eight, while plain text under the identical strings resolves to Llama 3 on eight and Qwen on three, with no GLM at all.

Scored against the catalogue

Across the full fifty strings, the two text signatures score as follows against 288 measured models:

Reading Best match Score Median of 288
Text path A Llama 3 family 47 of 50 12
Text path B Qwen family 46 of 50 14

43 of the 50 strings returned both, and not one returned a reading that fitted neither. Eight models tie at the top of path A across three different companies, and Qwen 3.5, 3.6, 3.7 and Kwaipilot all tie near the top of path B. So each path names a vocabulary family, never a checkpoint and never a company.

Qwen first, then Llama 3, then GLM

Our first reading, on 16 September, measured Qwen. We re-measured and read Llama 3 on 17 September. An independent toolkit, majiayu000/stealthprint, read the same endpoint as Llama 3 as well.

Both readings were of real tokenizers, because the endpoint was returning both. That pair is what pointed at a system behind the address. A third team reported GLM behaviour on image requests, and once we re-ran that path with strings chosen to separate GLM from Llama 3 we reproduced it exactly.

The lesson worth keeping is about the question rather than any lab: a measurement can be sound and still answer a question that was posed too narrowly. The reveal graded every reading at once: the endpoint held several tokenizers, one per component model, so each early reading had caught a real component. The question with an answer was the one the differential test asked: how many, and which families.

A deleted price tag pointed at one company

Cloudflare documented Union Alpha in pull request 33478 on 16 September 2026. Its first two revisions carry five worked examples with real costs on them. A later commit in the same pull request, “Mark Union Alpha launch pricing free”, replaces every cost with zero. That reads as an ordinary launch-pricing correction. The earlier figures are still in the history, and they are arithmetic.

Example Input tokens Output tokens Cost as first published
System Guidance 19 46 $0.00031125
Coding Example 14 39 $0.00026125
Follow-up Conversation 43 96 $0.00065375
Creative Writing 12 39 $0.00025875

Two of those give two equations in two unknowns. Solving them, assuming nothing, returns $1.25 per million input tokens and $6.25 per million output. Those rates then reproduce the rest. A fifth example, 29 input tokens of which 28 were cached and five output, cost $0.0000367. With the first two rates known, that forces a third: $0.15 per million cached input.

When we ran this arithmetic, Unbiased published Pareto at exactly those three numbers: $1.25 input, $6.25 output, $0.15 cached input. The paid 26.9 card announced at the reveal reads higher, $2.50 in, $7.50 out and $0.25 cached, so the match dates the documents it was run on: the examples were priced the way Pareto was priced when they were written.

Unbiased's published rate card for Pareto, a three-column panel headed The rate card reading one dollar twenty-five per million input tokens, fifteen cents per million cached input tokens and six dollars twenty-five per million output tokens
Pareto's published rates, checked again on 17 September 2026 at unbiased.ai/pricing. The three figures Cloudflare's arithmetic returns are these three figures.

Pareto is built by Circuit & Chisel, which sells it through Unbiased, and its own pages describe it as a model that runs several models against each other on every request and keeps the best answer. That is close to the wording Cloudflare used for Union Alpha on 16 September and removed on 17 September.

Circuit and Chisel's home page reading The new frontier AI lab, Frontier intelligence for 75 per cent less, and We build Unbiased, the platform for Pareto, our model that blends frontier and open models, above a ticker noting a private beta with a few dozen customers
Circuit & Chisel states the relationship itself: Pareto is the model, Unbiased is the platform it is sold through, and it blends frontier and open models. Captured 17 September 2026 from circuitandchisel.com.

When we published, this connected two documents, and that was all it connected. The Union Alpha page could have been adapted from Pareto’s documentation, carrying the example costs and the architecture wording across together, with a third party serving the endpoint. The reveal closed the distance: the endpoint was Pareto itself, so the examples described the system they were priced for.

A release-note repository came down

On 17 September a Circuit & Chisel repository, circuitandchisel/unbiased-releases, published as “release notes for customer-facing changes across Unbiased products”, stopped resolving. It had held a pull request titled “Add routing release notes for gpu-router-v2026.09.13.3”, opened by a release-note bot and closed by a human with the comment “Closing for WAY too much fucking information about how Pareto works”.

The removal is real and specific. The sibling repository circuitandchisel/pareto-evals is still public, so this is not a site-wide fault, and no Wayback Machine snapshot exists of either the pull request or the repository root.

A third-party account reported that a routing release note in that repository names union-alpha as a per-account public name for Pareto, in a change dated 14 September. That would be a direct documentary link, and far stronger than the pricing arithmetic above. We cannot verify it.

We went looking again on 17 September. The repository answers 404 and so does every pull request under it. Search engines still carry titles from it, which is how we know what was there: pull requests 5, 6 and 14, the last of them “Add routing release notes for gpu-router-v2026.09.13.3”. The one the report points to, number 16, appears in no index we can reach, and the Wayback Machine holds no snapshot of it. So we record the report, and we record that we could not open the document behind it.

The closing comment fits two readings equally well: a bot published an internal alias and someone caught it, or the repository was pulled for exactly the reason the comment gives, that it revealed too much about how Pareto routes, with no Union Alpha connection in it. It complains about Pareto’s routing and says nothing about Union Alpha.

The reveal settled the identity a day later without the diff. The report stays logged here as history: if that pull request ever resurfaces, it would show whether the alias sat on paper four days before the account said so.

Palmer traced it to the same back end

A researcher posting as Palmer reached the same name along a different route, and published it before the reveal as his own finding. His method read the plumbing rather than the paperwork. Union Alpha’s completion IDs follow one pattern, chatcmpl-mu followed by 14 lowercase alphanumeric characters, and he reported matching the pattern, exactly, on The Unbiased Co’s own platform, then watching a request he sent to Union Alpha answered by the same vLLM fleet that answers Unbiased. He closed the thread with “CASE CLOSED”, and afterwards credited the earlier work here: “Thanks for the help sidestepping the red herring @YFarmX @cheatyyyy”.

The two observations carry different weight. An ID format travels between systems that share software, so on its own it is a lead. The fleet observation is the strong one: two names, one back end answering both. Both are his reports of his own testing rather than something we can rerun, and the reveal the same day put the answer on the record in the operator’s own words.

Two more things we checked

Cloudflare’s archived examples publish their exact requests alongside the token counts the endpoint returned, so they can be replayed. Six attempts each, against Union Alpha and two unrelated models on the same budget:

Model Examples reproduced Readings matching
Union Alpha 4 of 5 20 of 29
Nemotron 3 Super 0 of 5 0 of 28
Nemotron 3.5 Lightning 0 of 5 0 of 30

The one Union Alpha miss is the example whose archived response records 28 of its 29 tokens as cached, so the published figure belongs to a cache state we cannot summon on demand. It stays scored as a miss.

The archived responses also carry generation ids whose first eight characters, read as base-36 milliseconds, decode to a 16.479-second run on 15 September, in the order the file lists them, about 15.1 hours before OpenRouter’s catalogue entry was created. That is consistent with an automated batch made shortly before listing. It dates the completions and says nothing about who produced them, or when the cost beside them was calculated.

How far the Union Alpha readings go

Every figure here comes from usage.prompt_tokens with a one-token cap, measured differentially so the envelope cancels, and scored against a frozen corpus of 288 models measured on 3 September 2026. The raw readings, the per-envelope histograms and the working are committed alongside the page.

Read precisely, that gave three statements before the reveal, and the reveal graded all three.

  • A shared vocabulary places a family. Families have guests: Perplexity’s Sonar and the Hermes line sit in the Llama-3 group, Kwaipilot sits in the Qwen group, and both were built by companies that did not create those vocabularies. The reveal bore this out: the three vocabularies were three real components inside one system.
  • Cloudflare’s own file called the provider third-party, and left its provider and terms fields empty. The reveal put a name in the space: The Unbiased Co.
  • Input accounting tells you what counted the prompt. It follows the text as far as the tokenizer and no further. Under a blend that reading was exact: the models that counted the prompts sat inside a system that also called stronger ones to answer them.

Fingerprinting a stealth endpoint now means fingerprinting a system. The token counts identify the components, and the documents, the prices and the completion IDs identify the operator. Union Alpha took two days from listing to reveal with all of those in play, and the working on this page is the part anyone can rerun.

What the September releases turned out to be

Forty-two models joined the OpenRouter catalogue between 16 September and 5 October 2026, and on 5 October we measured the 27 of them that can be measured: the rest are serving variants of a model already on record, routers, or “pro” endpoints whose reasoning tokens move the baseline between calls. The same run picked up the popular older models the budget had skipped until now, GPT-4o, o3, GPT-4 Turbo, Claude Sonnet 4 and 4.5, and Claude Opus 4.1, 4.5 and 4.6 among them. Where each one landed:

Model Measures with What that says
GPT-6 Luna, GPT-6 Sol, GPT-6 Astra GPT-4o, o3, GPT-5, GPT-5.5 and GPT-5.5 Pro on all 50 strings One OpenAI vocabulary from GPT-4o in May 2024 to GPT-6. GPT-6.1 Sol agrees on the 49 strings it returned.
Claude Sonnet 5.5 Claude Opus 5, Opus 5.5 and Fable 5.1 on all 49 strings they share The Claude 5 vocabulary. It agrees with Claude Sonnet 4.5 on 15 strings of 50.
Claude Sonnet 4, Sonnet 4.5, Opus 4.5 Claude 3 Haiku on all 50 strings; Opus 4.1 and 4.6 on 49 Every Claude from the 3 line to the 4 line shares one vocabulary, and every Claude 5 shares another.
Grok 4.7 Grok 4.5 and Grok 4.6 on all 50 xAI kept its vocabulary.
GLM 5.3 Prime, GLM 5.3 FlashX GLM 5.3 and Ox Alpha on all 50 Same vocabulary as the rest of the GLM 5 line.
Qwen 3.8 Max (0902), Qwen 3.8 Max Prime, Qwen 3.6 Plus, Qwen 3.6 Max preview, Qwen 3.5 Plus The Qwen 3.5 and 3.6 vocabulary, 22 models Qwen 3.8 27B and Qwen 3.8 Omni Flash sit instead with Qwen 3.6 27B and 35B: the 3.8 generation ships on two vocabularies, one for the large models and one for the 27B line.
DeepSeek V4.1 Flash DeepSeek V4 Pro and the rest of the DeepSeek line on all 50 Same vocabulary. DeepSeek R1, measured cleanly for the first time, agrees with DeepSeek Chat on 47 of 50, one template apart.
MiMo V2.6 Flash, V2.6 Pro, V2.6 Pro UltraSpeed MiMo V2.5 and the Qwen 2.5 vocabulary, 41 models Xiaomi’s new generation keeps the Qwen 2.5 vocabulary.
Aion 3.5, Aion 3.5 Mini Each other on all 50, and nothing else on more than 6 strings A vocabulary this corpus had never seen. Both listings say “built on the GLM family”; Aion 3.0, which says the same, does measure with GLM.
Pareto 26.10 preview Llama 3.2 3B on 48 of 49 Where Union Alpha showed three tokenizers behind one endpoint, the 26.10 preview answers from a Llama 3 tokenizer.
Solar Mini 4 Solar Pro 3 and Solar Pro 4 on all 50 Upstage’s own vocabulary, with an identical API contract.
Command A Plus Cohere’s North Mini Code on all 50 Cohere’s vocabulary; Command R+ 08-2024 agrees with Command R7B on 49 of 49.
Schematron V2 Small, Schematron V2 Turbo Llama 3.1, 3.2 and 3.3 on all 50; IBM Granite 4.2 8B on all 50 Two sizes of one product line, built on two different bases.
Perceptron Mk1.5, Apodex 1.1 Mini, Nex-AGI N2.5 Mini and Pro, Ternary Bonsai 2 27B Qwen vocabularies: Mk1.5 with Perceptron Mk1 and Morph V3, Apodex and N2.5 Pro with Qwen 3.6 27B, N2.5 Mini and Bonsai with Qwen 3.5 and 3.6 Five more small-lab models on Alibaba’s vocabularies, which now carry 54 measured models.
GPT-4 Turbo, GPT-3.5 Turbo 16K GPT-3.5 Turbo, Granite 4.0 Micro and Phi-4 on all 50 OpenAI’s earlier cl100k vocabulary, the one its open tiktoken file documents.
Gemini 2.5 Flash Gemini 3.7 Flash and the Gemini line on all 50 Measured cleanly on the third attempt; caching had corrupted the first two.

Claude 5 counts tokens differently from Claude 4

The Anthropic rows are the clearest lineage split on the page. Claude 3 Haiku, Claude Sonnet 4, Sonnet 4.5, Opus 4.5 and Haiku 4.5 return identical counts on all 50 strings, and Opus 4.1 and Opus 4.6 match them on the 49 they return. Claude Opus 5, Opus 5.5, Fable 5.1 and now Sonnet 5.5 return identical counts to each other and agree with that older group on 15 strings of 50. Two vocabularies, one boundary, and the boundary is the version number: whatever else changed between the 4 and 5 lines, the tokenizer did.

OpenAI has kept one vocabulary from GPT-4o to GPT-6

The OpenAI rows say the opposite. GPT-4o, its two dated 2024 snapshots, o3, GPT-5 mini, GPT-5, GPT-5.5, GPT-5.5 Pro, GPT-6 Luna, GPT-6 Sol and GPT-6 Astra all share one exact signature, now the joint-largest cluster on the page at 41 models. GPT-4 Turbo and GPT-3.5 Turbo sit on the earlier cl100k vocabulary with IBM’s Granite 4.0 Micro and Microsoft’s Phi-4. Six “pro” endpoints (GPT-5.6 Luna, Sol and Terra Pro, GPT-6 Luna, Sol and Astra Pro, and GPT-6.1 Sol Pro) stay unmeasurable: each reasons before it answers, the reasoning moves the token baseline between calls, and the harness discards a run whose baseline moves rather than guess.

We measured 66 models twice, and 61 fingerprints held

A fingerprint is a reading taken on a date, and a reading is only worth selling if the next one can be checked against it. On 5 October 2026 we sent the fifty strings a second time to the 66 most-used models already on record, from GPT-5.5 and Claude Opus 5.5 down to Gemma 4 and Llama 4 Scout, and read each answer against the reading from August or September.

  • 66models measured a second time on 5 October 2026
  • 61agreed with their earlier reading on every string clean in both runs
  • 5moved, and each one is explained on its record page
  • 2of the five were a second provider counting the same vocabulary differently

Sixty-one held. GPT-5.5, GPT-5.4, Claude Opus 5.5, Gemini 3.8 Flash, DeepSeek V4 Pro, GLM 5.3, Kimi K3, Grok 4.5, MiniMax M3 and Llama 4 Maverick all returned the same count on every string they returned in August or September. Five moved, and the five are more useful than the sixty-one, because each one shows a different way a reading can move while the model stays where it was.

Model What moved What it was
GLM 5.1 9 of 49 strings, each exactly 7 tokens lower A second provider. Friendli, which served the September reading, reproduced it on all 50 strings when we pinned the calls to it. AtlasCloud now also serves the model and counts nine strings seven tokens lower.
DeepSeek V3.2 1 of 50 strings, 3 tokens higher One provider’s count of one string. DigitalOcean reproduced the August reading on every string it answered; AtlasCloud, which served the August reading, now counts the mathematical-alphabet string three tokens higher.
Qwen 3.8 27B 5 of 50 strings, 4 or 5 tokens higher The August reading, which was spread across eight providers. Four providers measured one at a time in October agree with each other, and the corrected record matches Qwen 3.6 27B and 35B on all 50 strings.
MiMo V2.5 Pro all 50 strings, each exactly 2 tokens higher The August baseline. The 4 September note on this record predicted exactly this: a 253-token hidden system prompt had pulled every row two tokens low, and the direct reading matches MiMo V2.5 on all 50.
DeepSeek Chat 2 of 50 strings, each 1 token lower A serving template. DeepInfra served the October reading where StreamLake served the August one, and the two templates glue the string to a neighbouring token on two rows.

The provider is part of the reading

GLM 5.1 is the one to remember. Pinned to Friendli, the model answers today exactly as it did on 3 September. Pinned to AtlasCloud, the same model with the same vocabulary reports a fixed overhead of 12 tokens against Friendli’s 5, and nine of the fifty strings come back seven tokens cheaper: Chinese, Japanese, Hebrew, Greek, combining marks, CJK extension B, Python, JSON and a URL. The vocabulary held. The bill depends on the route.

So every record now carries the provider that served each reading, the sweep can pin a model to one provider on demand, and a reading that moves is re-taken on the original provider before the record calls it a change. A move that the original provider reproduces is logged as provider accounting, which is what two of the five were.

What the re-check changed in the dataset

Two placements improved. Qwen 3.8 27B, which August’s composite reading had left with a signature of its own, now sits with the Qwen 3.6 vocabulary. MiMo V2.5 Pro, which the 4 September correction had folded onto its sibling by arithmetic, now carries a direct reading that matches its sibling on all 50 strings. One listing-level label is also fixed: Aion 3.0’s own listing says it is built on GLM, and its fingerprint sits in the GLM group, so the group is now named Z.ai throughout rather than being left unlabelled because one guest sat in it. GLM 5, 5.1, 5.2, 5.3 and Ox Alpha all read “Z.ai” on their pages now.

Every re-measurement is on the model’s record page under “Measured twice”, with the strings that moved, the count on each side and the provider on each side, and a fingerprint that moved appears in the change history beside the price and context changes it already tracked.

What the sweep covered

  • 465models in the catalogue on 5 October 2026
  • 333measured since 20 August, of which 310 gave a complete signature
  • 68distinct fingerprints among those signatures
  • 55tokenizer families, once template differences are folded in

Of the 465 models listed on 5 October, 313 carry a measurement; the other 20 of the 333 were measured before OpenRouter withdrew them, and their records stay. The 152 listed models without one are all accounted for. Eighty-three are variants of a model already measured, 26 route through an aggregator that hides which model answers, and 17 cost more to measure than the budget allowed. The remaining 26 split three ways: 14 returned readings corrupted by caching or by overhead that moved during the run (every OpenAI “pro” endpoint is among them, because reasoning tokens move the baseline between calls), six are unavailable to this account (Meta’s Muse Spark line asks for an age confirmation the account has never given), and six errored at the provider.

Eighty-nine unlabelled models now have a family

OpenRouter lists a tokenizer field for each model. For 138 of the 467 models on record, that field reads “Other”. Those are the interesting ones, and they are exactly the cases measurement can speak to.

Eighty-nine of those 138 now carry a family at high confidence or better. Eighty-one of them were measured directly. The other eight were never sent to the sweep at all: they are variants of a model that was measured, a free tier or a batch endpoint of the same weights, and they inherit that model’s family.

  • 138models whose stated tokenizer is "Other"
  • 81measured directly and placed at high confidence
  • 8more placed as variants of a measured model

Confidence comes from three signals. An exact fingerprint match with a known family is one. Identical published settings, all seven of them, is a second and independent one. The third covers a fingerprint sitting one token away from a group whose members all name the same family. Two different signals agreeing gets a record marked very high; one signal alone gets high.

Labs that name their tokenizer get it right

Among the 212 measured models that named a tokenizer rather than saying “Other”, the fingerprint sits with the family the label names in 200 cases, and contradicts it in none. Nine carry a label and sit alone or beside unlabelled peers, so there is nothing to compare against, and three sit in a group whose other members carry mixed labels.

This is the least dramatic result on the page and one of the most useful. It says the field is reliable where it is filled in, so a reader can take a stated tokenizer at face value. It also acts as a check on our own method: a measurement that disagreed with hundreds of correctly declared models would be measuring something other than what it claims to.

The one placement that reads as a disagreement is worth naming, because it is the same model that surfaces in the controls below. DeepSeek’s R1 distillation of Llama 70B carries a Llama 3 tokenizer tag and measures with the DeepSeek family. The dataset records that as partial rather than as a contradiction: it is DeepSeek’s own cross-family model, built with DeepSeek tooling on a Llama base, so both labels describe something real about it. That is an editorial judgement made in August 2026 and written into the build, rather than something the arithmetic decided on its own.

The families, and who is in them

Fifty-five labs have at least one measured model. These are the thirteen with the most, counted by models we measured rather than by models listed. Below this point four labs tie on six apiece, so thirteen is where the ranking stops being a ranking.

  • QwenAlibaba (Qwen)54
  • OpenAIOpenAI43
  • GoogleGoogle27
  • Mistral AIMistral AI19
  • ZZ.ai17
  • AnthropicAnthropic16
  • DeepSeekDeepSeek15
  • MetaMeta9
  • MiniMaxMiniMax8
  • NVIDIANVIDIA9
  • Moonshot AIMoonshot AI7
  • TencentTencent7
  • iAInclusionAI7

The placements worth a second look

Most of the eighty land where you would expect: a lab’s unlabelled model measures with that lab’s other models. A handful cross a boundary, and those are the ones worth reading carefully. Each row below is a model whose stated tokenizer is “Other”, and the family its fingerprint puts it in.

Model, as listed Measures with Confidence
NVIDIA Nemotron 3.5 Lightning Mistral High
Thinking Machines Inkling Small GPT High
Dots Studio Dots3-Note Preview Qwen High
Meta Muse Glimmer 30B Llama 4 High

None of these is an accusation. A model can carry one company’s name and another company’s tokenizer for entirely ordinary reasons: adapting an openly published base is common practice across the industry, and a fine-tune inherits the tokenizer of whatever it started from. The value here is that the relationship is stated as a measurement anyone can repeat, rather than inferred from a naming convention.

How far the evidence goes

This is the part most reports of this kind leave out, so it gets its own section.

The sweeps produced 310 complete signatures and found 68 distinct ones among them. That ratio is the whole story of what the method can do. Sixty-eight buckets for 310 models means the average bucket holds between four and five models, five buckets hold more than twenty, and the two largest, a Qwen group and the OpenAI group that runs from GPT-4o to GPT-6, hold 41 each. The measurement tells you which bucket a model sits in. Picking one model out of 41 is a different question, and this evidence does not answer it.

The honest summary

This method identifies a family. Naming an individual model needs something else on top: a distinctive API contract, a release date, a capability only one model has. The Ox Alpha case worked because all three were available at once.

Two further limits are worth stating. The measurement reads the tokenizer, which sits in front of the model, so a lab that swapped the weights behind an unchanged tokenizer would look identical here. And the counts come through a billing pipeline, so a provider that caches aggressively or adds variable overhead can corrupt a reading; fourteen models were marked unmeasurable for exactly that reason rather than being guessed at.

The checks behind all this

The three controls, and what each one rules out

A positive control asks whether models known to be related do group together. Twelve models declare the Llama 3 tokenizer. Eleven of them land in one family. The twelfth, DeepSeek's R1 distillation of Llama 70B, measures with the DeepSeek family instead, which is the correct answer: it carries a Llama tag but was rebuilt with DeepSeek tooling. A control that produced no outliers at all would be the suspicious result.

A negative control asks whether unrelated models are pushed apart. If the method were secretly measuring something generic, everything would look alike. Across three pairs chosen to be unrelated, agreement ran between 18 and 25 of the 50 strings, against the 50 of 50 that a genuine match produces.

A reproducibility check asks whether the same model answers the same way twice. GLM-5.3, measured on 21 August, 3 September and 5 October through different sessions, returned 50 identical counts out of 50 each time. The October re-check above extends the same test to 66 models.

All three pass. They are rerun every time the dataset is rebuilt, and a failure blocks the build.

How a single measurement is taken

Every request carries fixed scaffolding: a system prompt, role markers, formatting the provider adds. That baseline is measured first, with an empty payload, and subtracted. What remains is the cost of the string itself.

The baseline is then measured a second time after the fifty strings have run. If it has drifted, the whole run is discarded rather than corrected, because a moving baseline means the arithmetic underneath is unreliable.

Requests are pinned to the first provider that answers, since two providers serving one model can wrap it differently. Where routing moves anyway, a fresh baseline is taken for the new provider before its counts are used, and every observation records which providers served it.

A count below one token, or above the string's own byte length, is impossible and is marked corrupt. Eight or more corrupt rows retires the model as unmeasurable instead of admitting a partial signature to the dataset.

Why some models share a fingerprint but not a template

A serving template that glues the user's text against an adjacent character merges exactly one token, on exactly the rows where that join is possible. The result is two signatures that differ by one token on a handful of rows and agree everywhere else.

Treated naively, that splits one real family into two. The pipeline instead recognises the pattern, links the two signatures and records the link. Fifteen such links exist across the dataset, which is why 68 distinct signatures resolve into 55 families.

Where this came from

The method was built to answer one question in August 2026: a model called Ox Alpha appeared on OpenRouter with no owner attached to it. Its fingerprint matched Z.ai’s GLM-5.3 on all fifty strings, and its API contract matched exactly one model in a catalogue of 422. We first published the Ox Alpha fingerprint on the site on 20 August 2026, then expanded the write-up on 21 August. On 26 August Z.ai confirmed the model was theirs and named it GLM-5.3-Flash.

That case is written up in full in our Ox Alpha investigation and the confirmation that followed. This page is what happened when the same measurement was pointed at the whole catalogue rather than one model.

Union Alpha made it two for two. Ox Alpha took six days from first fingerprint to Z.ai’s confirmation; Union Alpha took two from listing to the maker’s reveal, with our attribution on the record in between. Each case also stretched the method: Ox Alpha showed a fingerprint plus an API contract can name a checkpoint, and Union Alpha showed three fingerprints on one endpoint can name a system.

Nex-N2.5: the Qwen attribution, checked against released files

Research note · 8 September 2026

A dated supplement: it measures released files, separately from the endpoint catalogue above, and leaves the historical sweep totals unchanged.

Nex-N2.5 Mini’s released tokenizer is an exact continuation of the earlier Nex-N2 Mini tokenizer. Its ordinary vocabulary and merge rules match the Qwen3.5 and Qwen3.6 controls we checked. This is a confirmation of the published foundation and a reproducible comparison of files.

Three distinct kinds of evidence

The first signal is a declaration. OpenRouter’s catalogue labels both Nex-N2.5 Mini and Pro with architecture.tokenizer: Qwen3. That field supplies a provider-facing family label. On its own it gives us the catalogue’s description of the endpoint.

The second signal is the maker’s released configuration. Nex’s Mini repository names Qwen3_5MoeForConditionalGeneration and model_type: qwen3_5_moe. This explicitly identifies the architecture used by the released package. It is stronger and more specific evidence than guessing a foundation from generated answers. The Pro repository, at the time of this check, carries a model card saying its weights are coming soon. Its evidence remains separate from Mini’s released files.

The third signal is our local measurement. We downloaded the public tokenizer files and encoded the same 50 diagnostic strings used in the catalogue study, with automatically added special tokens disabled. The comparisons were:

Control tokenizer Exact token-count matches with Nex-N2.5 Mini
Qwen3.5-35B-A3B 50 of 50
Qwen3.6-35B-A3B 50 of 50
Qwen3-30B-A3B 25 of 50

The older Qwen3 control is useful because it tests how much information the broad catalogue label carries. Its counts diverge on half the probes. The later pair both reproduce Mini’s full signature, which supports attribution to their shared tokenizer family. Since both controls match, these counts alone leave Qwen3.5 and Qwen3.6 indistinguishable.

Comparing the actual vocabulary

We also inspected the tokenizer structure. All 248,044 base-vocabulary entries have the same token IDs as the Qwen3.5 and Qwen3.6 controls. All 247,587 ordered merge rules match after normalising their JSON representation: Nex stores a merge as a two-element array, while these Qwen files store a space-separated pair. Comparing the raw JSON would mistake that formatting difference for a change in tokenisation.

Nex’s file adds seven tokens beyond the controls’ added-token lists, taking its total vocabulary to 248,077, against 248,070 in those Qwen controls. Their labels concern audio or text-to-speech handling. Such labels identify reserved tokens in a file; an operational audio capability requires separate model and endpoint evidence.

The decisive novelty check was the previous Nex release. Nex-N2 Mini and Nex-N2.5 Mini supplied byte-identical tokenizer files, including those seven additions. Both downloaded files have SHA-256:

87a7830d63fcf43bf241c3c5242e96e62dd3fdc29224ca26fed8ea333db72de4

The extra tokens are therefore inherited from the earlier released tokenizer. Their presence supplies continuity evidence.

How far this attribution reaches

Our supported conclusion is that the released Mini package uses a Qwen3.5-family architecture and the same tokenizer foundation as the checked Qwen3.5/Qwen3.6 releases, continuing Nex-N2 Mini’s tokenizer unchanged. The configuration and measured files support that statement through different kinds of evidence.

Tokeniser continuity can coexist with substantial changes to model weights, training, tool use and performance. Those properties require their own measurements. The local comparison identifies the released tokenisation machinery; identifying the exact checkpoint behind a hosted service requires evidence from that service.

We attempted the hosted fingerprint run as well. Mini returned rate-limit responses at the baseline stage. The harness then deferred Pro under its shared free-tier allowance rule. Neither endpoint produced a completed 50-string measurement in this attempt. The results above are labelled local-file measurements throughout, and add zero models to the historical endpoint totals.

Primary files: Nex-N2.5 Mini configuration, Mini tokenizer, previous Nex-N2 Mini tokenizer, Qwen3.5 control, Qwen3.6 control, older Qwen3 control, Pro model card, OpenRouter catalogue.

The downloadable measurement receipt records every compared file’s SHA-256, the 50-count vectors, structural comparisons and the previous-release check. The primary links above follow their repositories’ current files; a matching hash identifies the exact bytes used in this dated comparison. The diagnostic text remains in the research notebook.

Jev brings its own tokenizer

Research note · 18 September 2026

A dated supplement, measured through a route the sweep cannot reach. It adds one model to the catalogue study's method and leaves the historical sweep totals unchanged.

TypeSafe AI left stealth on 15 September 2026, and its first System One model reached OpenRouter’s catalogue at 00:01 UTC on 18 September. typesafe/jev-1.13 takes a piece of state plus typed questions and returns calibrated probabilities, at $0.042 per million input tokens with output free. We measured it the same day it listed.

Measuring it needed a different route

Jev answers at POST /api/alpha/decisions, the route OpenRouter’s own Go SDK binds in decisions.go, and its text->decisions modality places it outside the /api/v1/models chat catalogue that lists the other 446 models. The endpoints API describes it in full, so the record is public; the chat completions route answers 500 for it under every request shape.

That is a gap in the method worth naming. A sweep that plans from the chat catalogue will keep missing models of this class as more of them appear. The measurement here ports the sweep’s method onto the decisions route without changing it: the same fifty strings, the same single-character baseline taken before and after, the same rule that marks a row corrupt when its marginal falls below one token or runs past the string’s own byte length.

The run was clean twice over

Two full passes, fifty clean rows each, and a fixed overhead of 273 tokens that held on every check. The two passes agree on all fifty counts. Union Alpha’s two vocabularies announced themselves as disagreements between repeated readings of byte-identical requests, so an endpoint of that kind would have shown itself here. This one returns one vocabulary, from one provider, on one dated checkpoint, for about a tenth of a penny.

The closest match reaches 48%

Scored against every record in the dataset the comparison could score, 286 of them, counting exact agreements over the strings both sides hold clean:

Model Agreement Family
Tencent HY-MT2 7B 24 of 50 Tencent
Tencent Hunyuan A13B Instruct 24 of 50 Tencent
Poolside Laguna XS 2.1 21 of 50 Unplaced
Thinking Machines Inkling Small 20 of 50 GPT
OpenAI o4-mini 20 of 50 GPT

Thirty-three models tie at 20 of 50, spanning OpenAI’s line, Thinking Machines and Meituan, so the lower half of that table is a sort order rather than a ranking. The median across all 286 compared models is 14 of 50.

Jev carries a new signature

A genuinely shared vocabulary reads above 90% in this corpus: Union Alpha matched the Llama 3 family on 47 of 50, and the reproducibility control returned 50 of 50. Jev’s best match reaches 24 of 50, twenty points clear of the median and far below that bar, with the top of the table spread across three different labs. The constant-offset scan that rescued earlier readings peaks at the same 24 and at an offset of zero, so a reporting constant explains none of it.

The dataset’s own clustering reaches the same verdict without being asked: ingesting Jev took the distinct signature count from 64 to 65 and the tokenizer groups from 50 to 51, so the pipeline opened a new group for it rather than folding it into one that existed.

So Jev’s tokenizer is a signature this corpus has never held: its own vocabulary, measured here for the first time. That fits a lab that built a decision model from its own stack rather than adapting a published base. The limit is the one this page states everywhere: a fingerprint places a vocabulary, and a model added to the corpus tomorrow could still turn out to share it.

What this page reports

The figures here are a snapshot: the OpenRouter catalogue as it stood on 5 October 2026, built into a dataset the same day. The catalogue moves. Forty-two models on it were new since 16 September, and nineteen had been withdrawn in the same three weeks.

Anyone with an API key can run the same measurement and check the arithmetic. What we publish here is the aggregate: the counts, the families, the controls and the limits. The fifty strings themselves stay in our notebook for now.