YFarmX

Model identity · decision layer

The Jev desk

YFarmX measures what a model's tokenizer actually does. Jev turns each of those evidence records into a decision a router can act on: whether the declared label holds, whether the model reads as somebody else's under a new name, and how sure it is of both.

266measured models asked
1,330typed decisions per sweep
126share a fingerprint across lab lines
$0.0147input cost billed, output free
trust (221) verify (21) flag (19) hold (4)One stripe per model answered, 265 of 266. Height shows how strong Jev judged the evidence.

Evidence is not a decision

Every record in the identity catalogue carries a measurement: a prompt overhead, a tokenizer fingerprint, and the models that share it. Reading one tells you what was found. Acting on it means answering a narrower question, 266 times over, and the answer needs a confidence attached rather than a label.

A chat model would write a paragraph per record that somebody then has to parse. Jev answers the typed question directly and returns the distribution behind it, so a verdict arrives as a number this page can render and a router can threshold on.

Five questions, asked of every measured model

One call per model carries all five. Each answer comes back with the full distribution, so a split verdict stays visible instead of collapsing to its winner.

noul

Does the declared label hold?

Does the tokenizer family the catalogue declares for this model agree with what the measurement found?

noul → probability 0 to 1

noul

Does it read as a rebadge?

Does this model read as an existing model from another lab, served under a new name?

noul → probability 0 to 1

choice

What should a router do?

A router is deciding how to treat this model in its catalogue. What should it do with this record?

choice → trust · verify · flag · hold

score

How strong is the evidence?

How strong is the evidence behind the family this record infers?

score → thin · suggestive · solid · conclusive

score

How exposed is a paying customer?

A customer pays per token on this endpoint. How exposed are they by any gap between what this model is declared to be and what it was measured to be?

score → none · slight · notable · severe

$0.0147

for the whole catalogue

266 evidence briefs come to 349,383 input tokens. At the published rate of $0.042 per million input tokens, with output tokens free, one full sweep returns 1,330 calibrated answers for less than two pence, in 16 seconds.

That is what makes re-running it on every catalogue refresh reasonable: the verdicts stay current with the measurements instead of ageing between reviews.

The verdicts

Swept 18 September 2026 against ~typesafe/jev-latest. 265 models answered, at a cost of $0.01467. Ranked by the action Jev picked, so anything it wants flagged or held sits at the top, showing the first 40 of 265.

ModelLabActionLabel holdsReads as rebadge
LiquidAI: LFM2.5-2.6B (free)Liquid AIflag39%74%
AionLabs: Aion-RP 1.0 (8B)Aion Labsflag40%73%
TheDrummer: UnslopNemo 12BThedrummerflag63%73%
WizardLM-2 8x22BMicrosoftflag34%71%
Amazon: Nova 2 LiteAmazonflag27%69%
Mistral: Mistral Small 3Mistral AIflag72%66%
Tencent: Hy4 previewTencentflag37%64%
Dots Studio: Dots3-Note Preview (free)Dots Studioflag81%63%
NVIDIA: Nemotron 3.5 Content Safety (free)NVIDIAflag38%59%
Cohere: North Mini Code (free)Cohereflag52%58%
Baidu: ERNIE 4.5 VL 424B A47B Baiduflag40%57%
Microsoft: Phi 4Microsoftflag62%57%
IBM: Granite 4.2 8BIBMflag42%56%
IBM: Granite 4.0 MicroIBMflag69%51%
Z.ai: GLM 4.7 FlashZ.aiflag54%46%
Google: Lyria 3 Clip PreviewGoogleflag47%27%
Google: Gemma 2 27BGoogleflag51%7%
Anthropic: Claude Opus 4.7Anthropicflag64%5%
Anthropic: Claude Sonnet 4.6Anthropicflag77%5%
Meituan: LongCat 2.0Meituanhold43%66%
Arcee AI: Trinity Large ThinkingArcee AIhold42%65%
Poolside: Laguna XS 2.1Poolsidehold42%61%
Reka EdgeRekaaihold45%50%
Union AlphaUndisclosed (stealth)verify58%80%
TheDrummer: Cydonia 24B V4.1Thedrummerverify51%77%
Qwen: Qwen3.5-122B-A10BAlibaba (Qwen)verify64%42%
Meta: Llama Guard 4 12BMetaverify47%25%
AionLabs: Aion-2.0Aion Labsverify75%21%
AionLabs: Aion-3.0-MiniAion Labsverify69%19%
Tencent: Hy-MT2-7BTencentverify55%18%
Cohere: Command R7B (12-2024)Cohereverify83%18%
Tencent: Hunyuan A13B InstructTencentverify54%16%
SpaceXAI: Grok 4.5xAIverify75%12%
Cohere: Command R (08-2024)Cohereverify75%12%
SpaceXAI: Grok 4.6xAIverify72%9%
DeepSeek: R1 Distill Llama 70BDeepSeekverify47%8%
Google: Nano Banana 2 (Gemini 3.1 Flash Image)Googleverify76%7%
Anthropic: Claude Fable 5Anthropicverify77%6%
Anthropic: Claude Fable 5.1Anthropicverify79%5%
Anthropic: Claude Opus 5Anthropicverify75%5%

Union Alpha reads as Llama 3

One record shows why the judgement needs a probability rather than a rule.

What the measurement found

Union Alpha is served from a stealth endpoint and declares its tokenizer as "Other". Measured against the same 47 strings as every other model in the catalogue, it agrees with Meta's Llama 3 family on 46 of them.

Its six nearest neighbours are four Meta models and two community fine-tunes of Llama. Sao10K's Euryale carries Llama's tokenizer because it is built on Llama, which is an honest lineage. A stealth endpoint with the same reading is a question, and the evidence alone settles neither.

Read the full evidence record

DeclaredOther
Measured familyLlama3
Grademoderate
Overhead17 tokens
46/47Meta: Llama 3.2 3B Instruct (Meta)
46/47Meta: Llama 3.1 8B Instruct (Meta)
46/47Meta: Llama 3.1 70B Instruct (Meta)
46/47Meta: Llama 3.3 70B Instruct (Meta)
46/47Sao10K: Llama 3.1 Euryale 70B v2.2 (Sao10k)
46/47Sao10K: Llama 3.3 Euryale 70B (Sao10k)

What Jev said about it

Reads as a rebadge: 80%. Declared label holds: 58%.

flag 46%verify 46%trust 7%hold 1%

Flag and verify came back level. A single label would print one of them and lose the fact that Jev split evenly between showing buyers a discrepancy and asking for another measurement.

126 models share a fingerprint across lab lines

These are the records where a rule gives the wrong answer either way. Each one matches a model from a different lab on at least 95% of the test strings, which is the honest signature of an open-base fine-tune and equally the signature of a rebadge. Grouped into the 51 relationships between labs that produce them, closest first:

Closest exampleLabMatches across the lineStringsModels
NVIDIA: Nemotron 3.5 LightningNVIDIAMistral: Mistral Nemo (Mistral AI)50/5074
Qwen: Qwen3 Max ThinkingAlibaba (Qwen)ByteDance: UI-TARS 7B (Bytedance)50/5035
Dots Studio: Dots3-Note Preview (free)Dots StudioQwen: Qwen3 30B A3B Instruct 2507 (Alibaba (Qwen))50/5035
Xiaomi: MiMo-V2.5XiaomiQwen: Qwen3 30B A3B Instruct 2507 (Alibaba (Qwen))50/5035
DeepSeek: DeepSeek V4 Flash Vision ExpDeepSeekStepFun: Step 3.5 Flash (StepFun)50/5034
Qwen: Qwen3.8 FlashAlibaba (Qwen)Nex AGI: Nex-N2-Mini (Nex Agi)50/5020
Sao10K: Llama 3.3 Euryale 70BSao10kMeta: Llama 3.2 3B Instruct (Meta)50/5016
Nous: Hermes 4 405BNous ResearchMeta: Llama 3.2 1B Instruct (Meta)50/5013
Perplexity: Sonar ProPerplexityMeta: Llama 3.2 3B Instruct (Meta)50/5012
Thinking Machines: Inkling SmallThinking MachinesOpenAI: GPT-5 Nano (OpenAI)50/5012
Morph: Morph V3 LargeMorphQwen: Qwen3.5 Plus 2026-04-20 (Alibaba (Qwen))27/277
Z.ai: GLM 5.3 FlashZ.aiOx Alpha (Undisclosed (stealth))50/507
Perplexity: Sonar ProPerplexitySao10K: Llama 3.1 Euryale 70B v2.2 (Sao10k)50/506
TheDrummer: Cydonia 24B V4.1ThedrummerMistral: Mistral Small 3 (Mistral AI)49/496
Qwen: Qwen3.6 35B A3BAlibaba (Qwen)Kwaipilot: KAT-Coder-Air V2.5 (Kwaipilot)50/504
IBM: Granite 4.0 MicroIBMOpenAI: GPT-3.5 Turbo (OpenAI)50/504
Perceptron: Perceptron Mk1PerceptronQwen: Qwen3.5 Plus 2026-04-20 (Alibaba (Qwen))27/274
TheDrummer: Cydonia 24B V4.1ThedrummerVenice: Uncensored (Cognitivecomputations)49/494
Magnum v4 72BAnthracite OrgQwen: Qwen3 30B A3B Instruct 2507 (Alibaba (Qwen))50/503
Venice: UncensoredCognitivecomputationsMistral: Mistral Small 3 (Mistral AI)50/503
IBM: Granite 4.0 MicroIBMMicrosoft: Phi 4 (Microsoft)50/503
Microsoft: Phi 4MicrosoftOpenAI: GPT-3.5 Turbo (OpenAI)50/503
Relace: Relace SearchRelaceQwen: Qwen3 30B A3B Instruct 2507 (Alibaba (Qwen))50/503
Dots Studio: Dots3-Note Preview (free)Dots StudioXiaomi: MiMo-V2.5 (Xiaomi)50/502
Dots Studio: Dots3-Note Preview (free)Dots StudioByteDance: UI-TARS 7B (Bytedance)50/502

Check one endpoint from your terminal

The same five questions, asked about a single model, as a command that needs no checkout. It fetches the records it names from this site, works out how far two fingerprints agree, and puts the result to Jev. One call, about two thousandths of a penny.

npx yfarmx-model-guard stealth/union-alpha \ --against meta-llama/llama-3.3-70b-instruct

The exit code follows the verdict, so it gates a deploy: 0 trust, 1 verify, 2 flag, 3 hold. Add --json for machine-readable output, or --explain to see what each question asks.

Run the sweep yourself

The script reads the committed catalogue, builds one evidence brief per model, and writes the answers back as JSON. It invents nothing: a model it could not ask lands in an errors list rather than a verdict.

# a trial run over ten models node scripts/model-identity/jev-verdicts.mjs --limit 10 # the full catalogue node scripts/model-identity/jev-verdicts.mjs # build the payloads and print the cost, calling nothing node scripts/model-identity/jev-verdicts.mjs --dry-run

Built on TypeSafe's Jev through the OpenRouter Decisions API, over the YFarmX Model Identity catalogue measured 16 September 2026.