YFarmX

AI model fingerprints

counting tokens across 278 models to see which ones share a lineage

37 min readLarge Language ModelsLast updated:

Editorial collage headed MODEL FINGERPRINTS with the subtitle 278 models measured, 64 signatures: one large inked thumbprint on newsprint whose ridges are formed from printed digits, a torn spec sheet carrying two rows of token counts, a pencilled note reading 453 catalogued, a blank index card, and pinned paper chips showing the OpenAI, Anthropic, Google and Qwen logos

Key facts

453OpenRouter, 3 Sep 2026
Models catalogued
278261 complete signatures
Measured
64grouping into 50 families
Distinct fingerprints
2both confirmed by their makers
Stealth models named

Every AI model chops text into tokens, and the chopping rules differ from one model family to the next. Feed the same 50 awkward strings to 278 models, record how many tokens each one charges for, and models that share a tokenizer answer identically. That is enough to sort a catalogue into families and to place models whose makers say nothing about what is inside them. It is not enough to name an individual model, and this page is careful about the difference. Twice now the method has named a stealth model anyway, and both times the maker confirmed the call.

Called before the reveal · confirmed 18 September 2026 Union Alpha is Pareto 26.9. The stealth model named itself two days after listing, confirming the attribution we published while it was anonymous: Circuit & Chisel's Pareto, called from its pricing arithmetic, its paper trail and the three tokenizers behind its one endpoint. Read the case ↓

What this page is

Every AI model reads text by chopping it into pieces called tokens. The rules for chopping live in a part of the model called the tokenizer, and they differ from one family to the next. So if two models chop text in exactly the same way, across strings chosen to make those rules disagree, they almost certainly share foundations.

Between 20 August and 3 September 2026 we sent the same 50 strings to every model we could reach on OpenRouter, and recorded how many tokens each one charged. This page reports what came back. It is a measurement, run at a stated time on a stated catalogue, and everything below is traceable to it.

Read this first

A shared tokenizer shows a shared lineage: the same foundations, the same tooling, often the same base model. It does not show that one model is a copy of another. A fine-tune keeps the tokenizer of the model it was tuned from, so a family here can hold genuinely different models built on one starting point.

The idea in one minute

Ask a model to read the letter “a”, then twenty spaces, then the letter “b”. One model charges you three tokens for that. Another charges four. Neither is wrong. They were built with different rules about how to bundle runs of whitespace, and those rules are baked in long before the model is trained.

Repeat that with a family emoji, a line of Chinese, a twelve-digit number, a run of tabs, a mathematical expression. Each string is chosen because tokenizers tend to disagree about it. After fifty of them you have a column of numbers that acts like a fingerprint of the tokenizer.

The measurement is cheap, it needs no special access, and the model cannot decline to answer, because the token count is on the bill.

Seeing a fingerprint

These are real counts from the sweep. Read down a column: where two models show the same number on every row, their tokenizers agree.

Model a + 20 spaces + b 20 tabs Family emoji Chinese phrase 12 digits
OpenAI GPT-5-nano3311124
Meta Llama 3.3 70B3315124
Anthropic Claude Sonnet 54414184
Google Gemini 3.7 Flash3371012
DeepSeek V3.2341194
Z.ai GLM-5.333797
Mistral Medium 3.534191612
Alibaba Qwen3 Max Thinking33101012
Five of the fifty strings, and the tokens each model charged for them. The first two columns separate almost nothing; the family emoji separates nearly everything. That is why the set has fifty strings rather than five.

Notice how little the first column tells you. Seven of these eight models charge three tokens for the letter “a”, twenty spaces and the letter “b”. The emoji column splits the same eight models into six different answers. Most of the work is done by a handful of strings, which is exactly what a set of fifty is for: enough awkward cases that models which merely look alike come apart.

Union Alpha is Pareto 26.9

On 16 September 2026 OpenRouter listed an anonymous model, stealth/union-alpha: free, a 262,144-token context, text and images, no owner named. Two days later the account behind it gave the answer itself: Pareto 26.9, the blended system Circuit & Chisel sells through its Unbiased platform. Between those two dates sit the measurements, the arithmetic and the attribution this page published while the endpoint was still anonymous. The working stays below, in the order it happened.

OpenRouter's model page for Union Alpha, showing the identifier stealth slash union-alpha, modalities of text and image in and text out, a price of Free, a 262K context and a release date of 16 September 2026
OpenRouter's listing on the day it appeared: free, 262K of context, text and image in, and a provider called stealth. Captured 17 September 2026 from openrouter.ai.
Circuit & Chisel's logo, two coral half-discs facing each other Confirmed by the maker · 18 September 2026 Circuit & Chisel's Pareto 26.9, sold through Unbiased

On 18 September the Union Alpha account on X posted "You found us", and named the stealth model on OpenRouter, OpenCode and Cloudflare: Pareto 26.9 from The Unbiased Co, Circuit & Chisel's blended system, sold through its Unbiased platform. The reveal described what had been answering all along: several models, open and frontier, work the same task, a harness checks the work, and a stronger model is called in where the task needs one.

The account said the stealth listing existed to gather real usage ahead of a 26.10 release the following month, that demand reached billions of tokens a minute within hours, and that AWS helped triple capacity overnight. It put the paid Pareto 26.9 card at:

  • $2.50per million input tokens
  • $7.50per million output tokens
  • $0.25per million cached input
Called before the reveal: the pricing arithmetic, the description trail and the three fingerprints below are the case as we published it while the model was anonymous, with the corrections we made along the way left in view.
  • 16 SeptemberListed anonymously. OpenRouter carries stealth/union-alpha: free, 262K of context, text and images. Cloudflare documents it as a blended model, with five worked cost examples.
  • 17 SeptemberThree tokenizers surface. Early fingerprints had split between Qwen and Llama 3, with GLM reported on images. Differential measurement finds all three behind the one endpoint: Llama 3 and Qwen answer on text, the GLM family with images and tools.
  • 17 SeptemberWe name Pareto. At 04:51 UTC we publish the attribution on X: Circuit & Chisel's Pareto, from the pricing arithmetic and Cloudflare's own first description.
  • 17 SeptemberThe paper trail thins. Cloudflare rewrites its catalogue entry, and Circuit & Chisel's release-note repository stops resolving.
  • Before the revealAn independent trace agrees. Palmer matches Union Alpha's completion IDs to Unbiased's platform and reports the same vLLM fleet answering both.
  • 18 SeptemberConfirmed. The Union Alpha account posts "You found us" and names Pareto 26.9.

We called it a day before the reveal

The attribution went up while the model was anonymous: on this page, and in a post on X at 04:51 UTC on 17 September, the earliest public post we can verify connecting Union Alpha to Pareto and Circuit & Chisel. The reveal came the next day.

Our public call · 17 September 2026, 04:51 UTC Union Alpha, called as Circuit & Chisel's Pareto, a day before the reveal View the post ↗

The call rested on three legs, each one set out below. Cloudflare’s day-one entry described a blended system in nearly the same words Circuit & Chisel uses for Pareto. Five worked cost examples in the same entry solve to Pareto’s then-published rates, to the cent. And the endpoint carried three tokenizer fingerprints, which is what several models behind one door look like from the outside. The reveal confirmed the identity and the architecture both.

Three tokenizers behind one endpoint

Measuring the endpoint’s tokenizer returned three different answers depending on what the request carried, which is a separate finding and is set out further down. Once the system had a name, the three answers read as its architecture showing through: a blend carries several tokenizers because it carries several models.

Cloudflare’s developer documentation gave the first reason to look at a blend. From 16 September its catalogue entry described Union Alpha as “a blended AI model that engages multiple language models in parallel for each request, synthesizes one answer, and supports text and vision inputs through a single API response”, and tagged it Blended Model.

That description was removed the next day. On 17 September, commit 333aed8, the entry was rewritten to “a multimodal model designed for research, coding, and agentic workflows”. Three separate fields carrying the same claim went in that one commit: the description, the Blended Model tag, and an Architecture: Blended model metadata entry. Its pull request summary was left as an empty template, so the record does not say why.

Both revisions are frozen in our research folder, so the quotes stay checkable. The measurements below were taken before either edit and do not depend on a word of Cloudflare’s wording.

Cloudflare's developer documentation page for stealth slash union-alpha, tagged Third-party, carrying the replacement description that calls it a multimodal model designed for research, coding and agentic workflows, above a TypeScript usage example
The same page after the edit. The Third-party tag stays, and the blended-model description is gone. Captured 17 September 2026 from developers.cloudflare.com.
Table headed Union Alpha has three fingerprints, showing that plain text requests return counts matching the Llama 3 family or the Qwen family while requests carrying an image or a tool declaration return counts matching the GLM family every time
The same endpoint, measured under different request shapes.

What we sent, and what came back

Only one number is read here, usage.prompt_tokens, with the reply capped at a single token. That field measures the input. Nothing the model generates can reach it.

What the request carried Which vocabulary answered Steady?
Plain text Llama 3 family, or Qwen family Flips, roughly half and half
A system prompt Llama 3 family, or Qwen family Flips
A prior assistant turn Llama 3 family, or Qwen family Flips
Code, or a very long context Llama 3 family, or Qwen family Flips
An image GLM family Same answer every time
A tool declaration GLM family Same answer every time

Attach an image and the endpoint stops varying: fourteen identical readings out of fourteen, and again under tool declarations. Send plain text and it alternates between two other vocabularies.

How you can tell a real tokenizer from a bad number

A tokenizer is exactly proportional to what you send it. So send a string, then send three copies of it, and subtract. Everything fixed cancels, including the chat template, the image and the tool schema, and what survives is the cost of the content alone.

Thirty digits are the sharpest single test on the text paths. The Llama 3 vocabulary bundles digits in threes, so thirty cost it ten. The Qwen vocabulary splits them, so the same thirty cost thirty. One endpoint returns both, from requests identical down to the byte.

Digits cannot separate Llama 3 from GLM, because both charge ten. That is a trap we fell into and had to climb out of: our first reading of the image path used digits alone and called it Llama. Re-running with eight strings that do separate the two, chosen from emoji, Thai, Hindi and combining marks, the image path resolves to the GLM family on seven of eight, while plain text under the identical strings resolves to Llama 3 on eight and Qwen on three, with no GLM at all.

Scored against the catalogue

Across the full fifty strings, the two text signatures score as follows against 288 measured models:

Reading Best match Score Median of 288
Text path A Llama 3 family 47 of 50 12
Text path B Qwen family 46 of 50 14

43 of the 50 strings returned both, and not one returned a reading that fitted neither. Eight models tie at the top of path A across three different companies, and Qwen 3.5, 3.6, 3.7 and Kwaipilot all tie near the top of path B. So each path names a vocabulary family, never a checkpoint and never a company.

What we published first, and why it was wrong

Our first reading, on 16 September, said Qwen. We re-measured and reached Llama 3 on 17 September. An independent toolkit, majiayu000/stealthprint, read the same endpoint as Llama 3 as well.

Both readings were of real tokenizers. The endpoint was returning both, and the question each of us was answering, which single family is this, had no answer. A third team reported GLM behaviour on image requests; we said we could not reproduce it, then found the flaw in our own probe and reproduced it exactly.

The lesson worth keeping is not about any lab. It is that a measurement can be sound and still answer a question that was badly posed. The reveal then graded every reading at once: the endpoint held several tokenizers, one per component model, so each early reading had caught a real component. The question with an answer was the one the differential test asked: how many, and which families.

A deleted price tag pointed at one company

Cloudflare documented Union Alpha in pull request 33478 on 16 September 2026. Its first two revisions carry five worked examples with real costs on them. A later commit in the same pull request, “Mark Union Alpha launch pricing free”, replaces every cost with zero. That reads as an ordinary launch-pricing correction. The earlier figures are still in the history, and they are arithmetic.

Example Input tokens Output tokens Cost as first published
System Guidance 19 46 $0.00031125
Coding Example 14 39 $0.00026125
Follow-up Conversation 43 96 $0.00065375
Creative Writing 12 39 $0.00025875

Two of those give two equations in two unknowns. Solving them, assuming nothing, returns $1.25 per million input tokens and $6.25 per million output. Those rates then reproduce the rest. A fifth example, 29 input tokens of which 28 were cached and five output, cost $0.0000367. With the first two rates known, that forces a third: $0.15 per million cached input.

When we ran this arithmetic, Unbiased published Pareto at exactly those three numbers: $1.25 input, $6.25 output, $0.15 cached input. The paid 26.9 card announced at the reveal reads higher, $2.50 in, $7.50 out and $0.25 cached, so the match dates the documents it was run on: the examples were priced the way Pareto was priced when they were written.

Unbiased's published rate card for Pareto, a three-column panel headed The rate card reading one dollar twenty-five per million input tokens, fifteen cents per million cached input tokens and six dollars twenty-five per million output tokens
Pareto's published rates, checked again on 17 September 2026 at unbiased.ai/pricing. The three figures Cloudflare's arithmetic returns are these three figures.

Pareto is built by Circuit & Chisel, which sells it through Unbiased, and its own pages describe it as a model that runs several models against each other on every request and keeps the best answer. That is close to the wording Cloudflare used for Union Alpha on 16 September and removed on 17 September.

Circuit and Chisel's home page reading The new frontier AI lab, Frontier intelligence for 75 per cent less, and We build Unbiased, the platform for Pareto, our model that blends frontier and open models, above a ticker noting a private beta with a few dozen customers
Circuit & Chisel states the relationship itself: Pareto is the model, Unbiased is the platform it is sold through, and it blends frontier and open models. Captured 17 September 2026 from circuitandchisel.com.

When we published, this connected two documents, and that was all it connected. The Union Alpha page could have been adapted from Pareto’s documentation, carrying the example costs and the architecture wording across together, with a third party serving the endpoint. The reveal closed the distance: the endpoint was Pareto itself, so the examples described the system they were priced for.

A release-note repository came down

On 17 September a Circuit & Chisel repository, circuitandchisel/unbiased-releases, published as “release notes for customer-facing changes across Unbiased products”, stopped resolving. It had held a pull request titled “Add routing release notes for gpu-router-v2026.09.13.3”, opened by a release-note bot and closed by a human with the comment “Closing for WAY too much fucking information about how Pareto works”.

The removal is real and specific. The sibling repository circuitandchisel/pareto-evals is still public, so this is not a site-wide fault, and no Wayback Machine snapshot exists of either the pull request or the repository root.

A third-party account reported that a routing release note in that repository names union-alpha as a per-account public name for Pareto, in a change dated 14 September. That would be a direct documentary link, and far stronger than the pricing arithmetic above. We cannot verify it.

We went looking again on 17 September. The repository answers 404 and so does every pull request under it. Search engines still carry titles from it, which is how we know what was there: pull requests 5, 6 and 14, the last of them “Add routing release notes for gpu-router-v2026.09.13.3”. The one the report points to, number 16, appears in no index we can reach, and the Wayback Machine holds no snapshot of it. So we record the report, and we record that we could not open the document behind it.

The closing comment fits two readings equally well: a bot published an internal alias and someone caught it, or the repository was pulled for exactly the reason the comment gives, that it revealed too much about how Pareto routes, with no Union Alpha connection in it. It complains about Pareto’s routing and says nothing about Union Alpha.

The reveal settled the identity a day later without the diff. The report stays logged here as history: if that pull request ever resurfaces, it would show whether the alias sat on paper four days before the account said so.

Palmer traced it to the same back end

A researcher posting as Palmer reached the same name along a different route, and published it before the reveal as his own finding. His method read the plumbing rather than the paperwork. Union Alpha’s completion IDs follow one pattern, chatcmpl-mu followed by 14 lowercase alphanumeric characters, and he reported matching the pattern, exactly, on The Unbiased Co’s own platform, then watching a request he sent to Union Alpha answered by the same vLLM fleet that answers Unbiased. He closed the thread with “CASE CLOSED”, and afterwards credited the earlier work here: “Thanks for the help sidestepping the red herring @YFarmX @cheatyyyy”.

The two observations carry different weight. An ID format travels between systems that share software, so on its own it is a lead. The fleet observation is the strong one: two names, one back end answering both. Both are his reports of his own testing rather than something we can rerun, and the reveal the same day put the answer on the record in the operator’s own words.

Two more things we checked

Cloudflare’s archived examples publish their exact requests alongside the token counts the endpoint returned, so they can be replayed. Six attempts each, against Union Alpha and two unrelated models on the same budget:

Model Examples reproduced Readings matching
Union Alpha 4 of 5 20 of 29
Nemotron 3 Super 0 of 5 0 of 28
Nemotron 3.5 Lightning 0 of 5 0 of 30

The one Union Alpha miss is the example whose archived response records 28 of its 29 tokens as cached, so the published figure belongs to a cache state we cannot summon on demand. It stays scored as a miss.

The archived responses also carry generation ids whose first eight characters, read as base-36 milliseconds, decode to a 16.479-second run on 15 September, in the order the file lists them, about 15.1 hours before OpenRouter’s catalogue entry was created. That is consistent with an automated batch made shortly before listing. It dates the completions and says nothing about who produced them, or when the cost beside them was calculated.

How far the Union Alpha readings go

Every figure here comes from usage.prompt_tokens with a one-token cap, measured differentially so the envelope cancels, and scored against a frozen corpus of 288 models measured on 3 September 2026. The raw readings, the per-envelope histograms and the working are committed alongside the page.

Read precisely, that gave three statements before the reveal, and the reveal graded all three.

  • A shared vocabulary places a family. Families have guests: Perplexity’s Sonar and the Hermes line sit in the Llama-3 group, Kwaipilot sits in the Qwen group, and both were built by companies that did not create those vocabularies. The reveal bore this out: the three vocabularies were three real components inside one system.
  • Cloudflare’s own file called the provider third-party, and left its provider and terms fields empty. The reveal put a name in the space: The Unbiased Co.
  • Input accounting tells you what counted the prompt. It follows the text as far as the tokenizer and no further. Under a blend that reading was exact: the models that counted the prompts sat inside a system that also called stronger ones to answer them.

Fingerprinting a stealth endpoint now means fingerprinting a system. The token counts identify the components, and the documents, the prices and the completion IDs identify the operator. Union Alpha took two days from listing to reveal with all of those in play, and the working on this page is the part anyone can rerun.

What the sweep covered

  • 453models in the catalogue on 3 September 2026
  • 278measured, of which 261 gave a complete signature
  • 64distinct fingerprints among those signatures
  • 50tokenizer families, once template differences are folded in

The 175 we did not measure are all accounted for. Ninety entries are variants of a model already measured, 35 cost more to measure than the budget allowed, and 19 route through an aggregator that hides which model answers. The remaining 31 split three ways: 16 errored, 13 returned readings corrupted by caching or by overhead that moved during the run, and two are waiting on account credit.

Eighty unlabelled models now have a family

OpenRouter lists a tokenizer field for each model. For 139 of the 453 models, that field reads “Other”. Those are the interesting ones, and they are exactly the cases measurement can speak to.

Eighty of those 139 now carry a family at high confidence or better. Sixty-five of them were measured directly. The other fifteen were never sent to the sweep at all: they are variants of a model that was measured, a free tier or a batch endpoint of the same weights, and they inherit that model’s family.

  • 139models whose stated tokenizer is "Other"
  • 65measured directly and placed at high confidence
  • 15more placed as variants of a measured model

Confidence comes from three signals. An exact fingerprint match with a known family is one. Identical published settings, all seven of them, is a second and independent one. The third covers a fingerprint sitting one token away from a group whose members all name the same family. Two different signals agreeing gets a record marked very high; one signal alone gets high.

Labs that name their tokenizer get it right

Among the models that named a tokenizer rather than saying “Other”, the measured fingerprint agreed with the declared family in every case but one.

This is the least dramatic result on the page and one of the most useful. It says the field is reliable where it is filled in, so a reader can take a stated tokenizer at face value. It also acts as a check on our own method: a measurement that disagreed with hundreds of correctly declared models would be measuring something other than what it claims to.

The exception is worth naming, because it is the same model that surfaces in the controls below. DeepSeek’s R1 distillation of Llama 70B carries a Llama 3 tokenizer tag and measures with the DeepSeek family. The dataset records that as partial rather than as a contradiction: it is DeepSeek’s own cross-family model, built with DeepSeek tooling on a Llama base, so both labels describe something real about it. That is an editorial judgement made in August 2026 and written into the build, rather than something the arithmetic decided on its own.

The families, and who is in them

Forty-eight labs have at least one measured model. These are the thirteen with the most, counted by models we measured rather than by models listed. Below this point four labs tie on five apiece, so thirteen is where the ranking stops being a ranking.

  • QwenAlibaba (Qwen)48
  • OpenAIOpenAI31
  • GoogleGoogle26
  • Mistral AIMistral AI19
  • ZZ.ai14
  • DeepSeekDeepSeek13
  • AnthropicAnthropic9
  • MetaMeta9
  • MiniMaxMiniMax8
  • NVIDIANVIDIA8
  • Moonshot AIMoonshot AI7
  • TencentTencent7
  • ByteDanceByteDance Seed6

The placements worth a second look

Most of the eighty land where you would expect: a lab’s unlabelled model measures with that lab’s other models. A handful cross a boundary, and those are the ones worth reading carefully. Each row below is a model whose stated tokenizer is “Other”, and the family its fingerprint puts it in.

Model, as listed Measures with Confidence
NVIDIA Nemotron 3.5 Lightning Mistral High
Thinking Machines Inkling Small GPT High
Dots Studio Dots3-Note Preview Qwen High
Meta Muse Glimmer 30B Llama 4 High

None of these is an accusation. A model can carry one company’s name and another company’s tokenizer for entirely ordinary reasons: adapting an openly published base is common practice across the industry, and a fine-tune inherits the tokenizer of whatever it started from. The value here is that the relationship is stated as a measurement anyone can repeat, rather than inferred from a naming convention.

How far the evidence goes

This is the part most reports of this kind leave out, so it gets its own section.

The sweep produced 261 complete signatures and found 64 distinct ones among them. That ratio is the whole story of what the method can do. Sixty-four buckets for 261 models means the average bucket holds four models, three buckets hold more than twenty, and the largest, a Qwen group, holds thirty-seven. The measurement tells you which bucket a model sits in. Picking one model out of thirty-seven is a different question, and this evidence does not answer it.

The honest summary

This method identifies a family. Naming an individual model needs something else on top: a distinctive API contract, a release date, a capability only one model has. The Ox Alpha case worked because all three were available at once.

Two further limits are worth stating. The measurement reads the tokenizer, which sits in front of the model, so a lab that swapped the weights behind an unchanged tokenizer would look identical here. And the counts come through a billing pipeline, so a provider that caches aggressively or adds variable overhead can corrupt a reading; thirteen models were marked unmeasurable for exactly that reason rather than being guessed at.

The checks behind all this

The three controls, and what each one rules out

A positive control asks whether models known to be related do group together. Twelve models declare the Llama 3 tokenizer. Eleven of them land in one family. The twelfth, DeepSeek's R1 distillation of Llama 70B, measures with the DeepSeek family instead, which is the correct answer: it carries a Llama tag but was rebuilt with DeepSeek tooling. A control that produced no outliers at all would be the suspicious result.

A negative control asks whether unrelated models are pushed apart. If the method were secretly measuring something generic, everything would look alike. Across three pairs chosen to be unrelated, agreement ran between 18 and 25 of the 50 strings, against the 50 of 50 that a genuine match produces.

A reproducibility check asks whether the same model answers the same way twice. GLM-5.3, measured on different days through different sessions, returned 50 identical counts out of 50.

All three pass. They are rerun every time the dataset is rebuilt, and a failure blocks the build.

How a single measurement is taken

Every request carries fixed scaffolding: a system prompt, role markers, formatting the provider adds. That baseline is measured first, with an empty payload, and subtracted. What remains is the cost of the string itself.

The baseline is then measured a second time after the fifty strings have run. If it has drifted, the whole run is discarded rather than corrected, because a moving baseline means the arithmetic underneath is unreliable.

Requests are pinned to the first provider that answers, since two providers serving one model can wrap it differently. Where routing moves anyway, a fresh baseline is taken for the new provider before its counts are used, and every observation records which providers served it.

A count below one token, or above the string's own byte length, is impossible and is marked corrupt. Eight or more corrupt rows retires the model as unmeasurable instead of admitting a partial signature to the dataset.

Why some models share a fingerprint but not a template

A serving template that glues the user's text against an adjacent character merges exactly one token, on exactly the rows where that join is possible. The result is two signatures that differ by one token on a handful of rows and agree everywhere else.

Treated naively, that splits one real family into two. The pipeline instead recognises the pattern, links the two signatures and records the link. Sixteen such links exist across the dataset, which is why 64 distinct signatures resolve into 50 families.

Where this came from

The method was built to answer one question in August 2026: a model called Ox Alpha appeared on OpenRouter with no owner attached to it. Its fingerprint matched Z.ai’s GLM-5.3 on all fifty strings, and its API contract matched exactly one model in a catalogue of 422. We first published the Ox Alpha fingerprint on the site on 20 August 2026, then expanded the write-up on 21 August. On 26 August Z.ai confirmed the model was theirs and named it GLM-5.3-Flash.

That case is written up in full in our Ox Alpha investigation and the confirmation that followed. This page is what happened when the same measurement was pointed at the whole catalogue rather than one model.

Union Alpha made it two for two. Ox Alpha took six days from first fingerprint to Z.ai’s confirmation; Union Alpha took two from listing to the maker’s reveal, with our attribution on the record in between. Each case also stretched the method: Ox Alpha showed a fingerprint plus an API contract can name a checkpoint, and Union Alpha showed three fingerprints on one endpoint can name a system.

Nex-N2.5: the Qwen attribution, checked against released files

Research note · 8 September 2026

A dated supplement: it measures released files, separately from the endpoint catalogue above, and leaves the historical sweep totals unchanged.

Nex-N2.5 Mini’s released tokenizer is an exact continuation of the earlier Nex-N2 Mini tokenizer. Its ordinary vocabulary and merge rules match the Qwen3.5 and Qwen3.6 controls we checked. This is a confirmation of the published foundation and a reproducible comparison of files.

Three distinct kinds of evidence

The first signal is a declaration. OpenRouter’s catalogue labels both Nex-N2.5 Mini and Pro with architecture.tokenizer: Qwen3. That field supplies a provider-facing family label. On its own it gives us the catalogue’s description of the endpoint.

The second signal is the maker’s released configuration. Nex’s Mini repository names Qwen3_5MoeForConditionalGeneration and model_type: qwen3_5_moe. This explicitly identifies the architecture used by the released package. It is stronger and more specific evidence than guessing a foundation from generated answers. The Pro repository, at the time of this check, carries a model card saying its weights are coming soon. Its evidence remains separate from Mini’s released files.

The third signal is our local measurement. We downloaded the public tokenizer files and encoded the same 50 diagnostic strings used in the catalogue study, with automatically added special tokens disabled. The comparisons were:

Control tokenizer Exact token-count matches with Nex-N2.5 Mini
Qwen3.5-35B-A3B 50 of 50
Qwen3.6-35B-A3B 50 of 50
Qwen3-30B-A3B 25 of 50

The older Qwen3 control is useful because it tests how much information the broad catalogue label carries. Its counts diverge on half the probes. The later pair both reproduce Mini’s full signature, which supports attribution to their shared tokenizer family. Since both controls match, these counts alone leave Qwen3.5 and Qwen3.6 indistinguishable.

Comparing the actual vocabulary

We also inspected the tokenizer structure. All 248,044 base-vocabulary entries have the same token IDs as the Qwen3.5 and Qwen3.6 controls. All 247,587 ordered merge rules match after normalising their JSON representation: Nex stores a merge as a two-element array, while these Qwen files store a space-separated pair. Comparing the raw JSON would mistake that formatting difference for a change in tokenisation.

Nex’s file adds seven tokens beyond the controls’ added-token lists, taking its total vocabulary to 248,077, against 248,070 in those Qwen controls. Their labels concern audio or text-to-speech handling. Such labels identify reserved tokens in a file; an operational audio capability requires separate model and endpoint evidence.

The decisive novelty check was the previous Nex release. Nex-N2 Mini and Nex-N2.5 Mini supplied byte-identical tokenizer files, including those seven additions. Both downloaded files have SHA-256:

87a7830d63fcf43bf241c3c5242e96e62dd3fdc29224ca26fed8ea333db72de4

The extra tokens are therefore inherited from the earlier released tokenizer. Their presence supplies continuity evidence.

How far this attribution reaches

Our supported conclusion is that the released Mini package uses a Qwen3.5-family architecture and the same tokenizer foundation as the checked Qwen3.5/Qwen3.6 releases, continuing Nex-N2 Mini’s tokenizer unchanged. The configuration and measured files support that statement through different kinds of evidence.

Tokeniser continuity can coexist with substantial changes to model weights, training, tool use and performance. Those properties require their own measurements. The local comparison identifies the released tokenisation machinery; identifying the exact checkpoint behind a hosted service requires evidence from that service.

We attempted the hosted fingerprint run as well. Mini returned rate-limit responses at the baseline stage. The harness then deferred Pro under its shared free-tier allowance rule. Neither endpoint produced a completed 50-string measurement in this attempt. The results above are labelled local-file measurements throughout, and add zero models to the historical endpoint totals.

Primary files: Nex-N2.5 Mini configuration, Mini tokenizer, previous Nex-N2 Mini tokenizer, Qwen3.5 control, Qwen3.6 control, older Qwen3 control, Pro model card, OpenRouter catalogue.

The downloadable measurement receipt records every compared file’s SHA-256, the 50-count vectors, structural comparisons and the previous-release check. The primary links above follow their repositories’ current files; a matching hash identifies the exact bytes used in this dated comparison. The diagnostic text remains in the research notebook.

Jev brings its own tokenizer

Research note · 18 September 2026

A dated supplement, measured through a route the sweep cannot reach. It adds one model to the catalogue study's method and leaves the historical sweep totals unchanged.

TypeSafe AI left stealth on 15 September 2026, and its first System One model reached OpenRouter’s catalogue at 00:01 UTC on 18 September. typesafe/jev-1.13 takes a piece of state plus typed questions and returns calibrated probabilities, at $0.042 per million input tokens with output free. We measured it the same day it listed.

Measuring it needed a different route

Jev answers at POST /api/alpha/decisions, the route OpenRouter’s own Go SDK binds in decisions.go, and its text->decisions modality places it outside the /api/v1/models chat catalogue that lists the other 446 models. The endpoints API describes it in full, so the record is public; the chat completions route answers 500 for it under every request shape.

That is a gap in the method worth naming. A sweep that plans from the chat catalogue will keep missing models of this class as more of them appear. The measurement here ports the sweep’s method onto the decisions route without changing it: the same fifty strings, the same single-character baseline taken before and after, the same rule that marks a row corrupt when its marginal falls below one token or runs past the string’s own byte length.

The run was clean twice over

Two full passes, fifty clean rows each, and a fixed overhead of 273 tokens that held on every check. The two passes agree on all fifty counts. Union Alpha’s two vocabularies announced themselves as disagreements between repeated readings of byte-identical requests, so an endpoint of that kind would have shown itself here. This one returns one vocabulary, from one provider, on one dated checkpoint, for about a tenth of a penny.

The closest match reaches 48%

Scored against every record in the dataset the comparison could score, 286 of them, counting exact agreements over the strings both sides hold clean:

Model Agreement Family
Tencent HY-MT2 7B 24 of 50 Tencent
Tencent Hunyuan A13B Instruct 24 of 50 Tencent
Poolside Laguna XS 2.1 21 of 50 Unplaced
Thinking Machines Inkling Small 20 of 50 GPT
OpenAI o4-mini 20 of 50 GPT

Thirty-three models tie at 20 of 50, spanning OpenAI’s line, Thinking Machines and Meituan, so the lower half of that table is a sort order rather than a ranking. The median across all 286 compared models is 14 of 50.

Jev carries a new signature

A genuinely shared vocabulary reads above 90% in this corpus: Union Alpha matched the Llama 3 family on 47 of 50, and the reproducibility control returned 50 of 50. Jev’s best match reaches 24 of 50, twenty points clear of the median and far below that bar, with the top of the table spread across three different labs. The constant-offset scan that rescued earlier readings peaks at the same 24 and at an offset of zero, so a reporting constant explains none of it.

The dataset’s own clustering reaches the same verdict without being asked: ingesting Jev took the distinct signature count from 64 to 65 and the tokenizer groups from 50 to 51, so the pipeline opened a new group for it rather than folding it into one that existed.

So Jev’s tokenizer is a signature this corpus has never held: its own vocabulary, measured here for the first time. That fits a lab that built a decision model from its own stack rather than adapting a published base. The limit is the one this page states everywhere: a fingerprint places a vocabulary, and a model added to the corpus tomorrow could still turn out to share it.

What this page reports

The figures here are a snapshot: the OpenRouter catalogue as it stood on 3 September 2026, built into a dataset on 4 September. The catalogue moves. Thirty-one models on it were new since 20 August, and twenty-eight had been withdrawn.

Anyone with an API key can run the same measurement and check the arithmetic. What we publish here is the aggregate: the counts, the families, the controls and the limits. The fifty strings themselves stay in our notebook for now.