Does the declared label hold?
Does the tokenizer family the catalogue declares for this model agree with what the measurement found?
noul → probability 0 to 1
Model identity · decision layer
YFarmX measures what a model's tokenizer actually does. Jev turns each of those evidence records into a decision a router can act on: whether the declared label holds, whether the model reads as somebody else's under a new name, and how sure it is of both.
Every record in the identity catalogue carries a measurement: a prompt overhead, a tokenizer fingerprint, and the models that share it. Reading one tells you what was found. Acting on it means answering a narrower question, 266 times over, and the answer needs a confidence attached rather than a label.
A chat model would write a paragraph per record that somebody then has to parse. Jev answers the typed question directly and returns the distribution behind it, so a verdict arrives as a number this page can render and a router can threshold on.
One call per model carries all five. Each answer comes back with the full distribution, so a split verdict stays visible instead of collapsing to its winner.
Does the tokenizer family the catalogue declares for this model agree with what the measurement found?
noul → probability 0 to 1
Does this model read as an existing model from another lab, served under a new name?
noul → probability 0 to 1
A router is deciding how to treat this model in its catalogue. What should it do with this record?
choice → trust · verify · flag · hold
How strong is the evidence behind the family this record infers?
score → thin · suggestive · solid · conclusive
A customer pays per token on this endpoint. How exposed are they by any gap between what this model is declared to be and what it was measured to be?
score → none · slight · notable · severe
for the whole catalogue
266 evidence briefs come to 349,383 input tokens. At the published rate of $0.042 per million input tokens, with output tokens free, one full sweep returns 1,330 calibrated answers for less than two pence, in 16 seconds.
That is what makes re-running it on every catalogue refresh reasonable: the verdicts stay current with the measurements instead of ageing between reviews.
Swept 18 September 2026 against ~typesafe/jev-latest. 265 models answered, at a cost of $0.01467. Ranked by the action Jev picked, so anything it wants flagged or held sits at the top, showing the first 40 of 265.
One record shows why the judgement needs a probability rather than a rule.
Union Alpha is served from a stealth endpoint and declares its tokenizer as "Other". Measured against the same 47 strings as every other model in the catalogue, it agrees with Meta's Llama 3 family on 46 of them.
Its six nearest neighbours are four Meta models and two community fine-tunes of Llama. Sao10K's Euryale carries Llama's tokenizer because it is built on Llama, which is an honest lineage. A stealth endpoint with the same reading is a question, and the evidence alone settles neither.
| Declared | Other |
|---|---|
| Measured family | Llama3 |
| Grade | moderate |
| Overhead | 17 tokens |
| 46/47 | Meta: Llama 3.2 3B Instruct (Meta) |
| 46/47 | Meta: Llama 3.1 8B Instruct (Meta) |
| 46/47 | Meta: Llama 3.1 70B Instruct (Meta) |
| 46/47 | Meta: Llama 3.3 70B Instruct (Meta) |
| 46/47 | Sao10K: Llama 3.1 Euryale 70B v2.2 (Sao10k) |
| 46/47 | Sao10K: Llama 3.3 Euryale 70B (Sao10k) |
Reads as a rebadge: 80%. Declared label holds: 58%.
Flag and verify came back level. A single label would print one of them and lose the fact that Jev split evenly between showing buyers a discrepancy and asking for another measurement.
These are the records where a rule gives the wrong answer either way. Each one matches a model from a different lab on at least 95% of the test strings, which is the honest signature of an open-base fine-tune and equally the signature of a rebadge. Grouped into the 51 relationships between labs that produce them, closest first:
| Closest example | Lab | Matches across the line | Strings | Models |
|---|---|---|---|---|
| NVIDIA: Nemotron 3.5 Lightning | NVIDIA | Mistral: Mistral Nemo (Mistral AI) | 50/50 | 74 |
| Qwen: Qwen3 Max Thinking | Alibaba (Qwen) | ByteDance: UI-TARS 7B (Bytedance) | 50/50 | 35 |
| Dots Studio: Dots3-Note Preview (free) | Dots Studio | Qwen: Qwen3 30B A3B Instruct 2507 (Alibaba (Qwen)) | 50/50 | 35 |
| Xiaomi: MiMo-V2.5 | Xiaomi | Qwen: Qwen3 30B A3B Instruct 2507 (Alibaba (Qwen)) | 50/50 | 35 |
| DeepSeek: DeepSeek V4 Flash Vision Exp | DeepSeek | StepFun: Step 3.5 Flash (StepFun) | 50/50 | 34 |
| Qwen: Qwen3.8 Flash | Alibaba (Qwen) | Nex AGI: Nex-N2-Mini (Nex Agi) | 50/50 | 20 |
| Sao10K: Llama 3.3 Euryale 70B | Sao10k | Meta: Llama 3.2 3B Instruct (Meta) | 50/50 | 16 |
| Nous: Hermes 4 405B | Nous Research | Meta: Llama 3.2 1B Instruct (Meta) | 50/50 | 13 |
| Perplexity: Sonar Pro | Perplexity | Meta: Llama 3.2 3B Instruct (Meta) | 50/50 | 12 |
| Thinking Machines: Inkling Small | Thinking Machines | OpenAI: GPT-5 Nano (OpenAI) | 50/50 | 12 |
| Morph: Morph V3 Large | Morph | Qwen: Qwen3.5 Plus 2026-04-20 (Alibaba (Qwen)) | 27/27 | 7 |
| Z.ai: GLM 5.3 Flash | Z.ai | Ox Alpha (Undisclosed (stealth)) | 50/50 | 7 |
| Perplexity: Sonar Pro | Perplexity | Sao10K: Llama 3.1 Euryale 70B v2.2 (Sao10k) | 50/50 | 6 |
| TheDrummer: Cydonia 24B V4.1 | Thedrummer | Mistral: Mistral Small 3 (Mistral AI) | 49/49 | 6 |
| Qwen: Qwen3.6 35B A3B | Alibaba (Qwen) | Kwaipilot: KAT-Coder-Air V2.5 (Kwaipilot) | 50/50 | 4 |
| IBM: Granite 4.0 Micro | IBM | OpenAI: GPT-3.5 Turbo (OpenAI) | 50/50 | 4 |
| Perceptron: Perceptron Mk1 | Perceptron | Qwen: Qwen3.5 Plus 2026-04-20 (Alibaba (Qwen)) | 27/27 | 4 |
| TheDrummer: Cydonia 24B V4.1 | Thedrummer | Venice: Uncensored (Cognitivecomputations) | 49/49 | 4 |
| Magnum v4 72B | Anthracite Org | Qwen: Qwen3 30B A3B Instruct 2507 (Alibaba (Qwen)) | 50/50 | 3 |
| Venice: Uncensored | Cognitivecomputations | Mistral: Mistral Small 3 (Mistral AI) | 50/50 | 3 |
| IBM: Granite 4.0 Micro | IBM | Microsoft: Phi 4 (Microsoft) | 50/50 | 3 |
| Microsoft: Phi 4 | Microsoft | OpenAI: GPT-3.5 Turbo (OpenAI) | 50/50 | 3 |
| Relace: Relace Search | Relace | Qwen: Qwen3 30B A3B Instruct 2507 (Alibaba (Qwen)) | 50/50 | 3 |
| Dots Studio: Dots3-Note Preview (free) | Dots Studio | Xiaomi: MiMo-V2.5 (Xiaomi) | 50/50 | 2 |
| Dots Studio: Dots3-Note Preview (free) | Dots Studio | ByteDance: UI-TARS 7B (Bytedance) | 50/50 | 2 |
The same five questions, asked about a single model, as a command that needs no checkout. It fetches the records it names from this site, works out how far two fingerprints agree, and puts the result to Jev. One call, about two thousandths of a penny.
npx yfarmx-model-guard stealth/union-alpha \
--against meta-llama/llama-3.3-70b-instructThe exit code follows the verdict, so it gates a deploy: 0 trust, 1 verify, 2 flag, 3 hold. Add --json for machine-readable output, or --explain to see what each question asks.
The script reads the committed catalogue, builds one evidence brief per model, and writes the answers back as JSON. It invents nothing: a model it could not ask lands in an errors list rather than a verdict.
# a trial run over ten models
node scripts/model-identity/jev-verdicts.mjs --limit 10
# the full catalogue
node scripts/model-identity/jev-verdicts.mjs
# build the payloads and print the cost, calling nothing
node scripts/model-identity/jev-verdicts.mjs --dry-runBuilt on TypeSafe's Jev through the OpenRouter Decisions API, over the YFarmX Model Identity catalogue measured 16 September 2026.