Google: Gemma 4 31B

google/gemma-4-31b-it · GoogleDeclared and measured agree
Tokenizer lineage (measured)
Gemma
Evidence confidence
High
Why high: an exact fingerprint shared with 1 model that declares the same family, without a second signal of a different kind yet.
Declared tokenizer tag
Gemma
Entered catalogue
2 April 2026
Last tested
21 August 2026
Measurement
Measured
Evidence
  • exact tokenizer signature shared with 1 model

Identity statement

Google: Gemma 4 31B is Google's model. Its measured tokenizer sits with the Gemma group, which is evidence that it builds on the Gemma tokenizer. Tokenizer reuse is normal engineering, open vocabularies travel between labs, and this says nothing against Google's authorship of the model itself.

Evidence stack

Signals of different kinds, weighed together
SignalResultReference
Tokenizer signatureExact match with 1 other modeltk_ba03436b
API surfaceNo other entry shares this exact contractapi_6e963360
Context and output262,144 context · 262,144 max outputdeclared
Reasoning contractOptionaldeclared
Serving providers observedDeepInframeasured 21 August 2026

Shares this fingerprint

Exact signature first; template-boundary shifts of the same signature beneath

Closest measured models

Agreement across the mutually clean test strings

Click a row to open the full comparison.

The measured fingerprint

Marginal prompt-token cost of each test string, grouped by script
All 50 rows

English and whitespace

en-prose14
en-long16
spaces-203
spaces-604
tabs-203
newlines-203
mixed-ws9

Digits

digits-99
digits-1212
digits-3030
digits-sep12
float-long22

CJK

zh-common10
zh-long12
zh-rare25
ja-kana7
ja-kanji9
ko11

Other scripts

ru12
ar10
he13
hi8
th10
el13

Emoji

emoji-basic5
emoji-skin10
emoji-zwj-family7
emoji-zwj-x321
emoji-flags8
emoji-prof11

Rare Unicode

math26
boxdraw17
combining10
cjk-ext-b16
surrogates27
zalgo22
rtl-mix5

Code

py-code28
py-indent18
json28
html17
regex44
camel6
snake11

Repetition and encodings

rare-word-x525
repeat-tok21
base6434
hex16
url23
uuid36

An x marks a row excluded as corrupt (caching or provider interference detected during measurement). Overhead subtracted: 13 prompt tokens.

Declared record

What the catalogue claims about this model

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

Modalities
text+image+video->text
Declared tokenizer
Gemma
Prompt price
$0.10 / M tokens
Completion price
$0.34 / M tokens

Supported parameters

frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p

defaults: {"temperature":1,"top_p":0.95,"top_k":64,"frequency_penalty":null,"presence_penalty":null,"repetition_penalty":null}

Catalogue entryWeights on Hugging Face

History

Every observation, kept as taken
  • 2 April 2026Enters the OpenRouter catalogue declared as Gemma.
  • 21 August 2026Fingerprinted in the YFarmX catalogue sweep · 50 of 50 strings measured clean.
  • 21 August 2026YFarmX assessment: consistent with the Gemma family, high confidence.

Cite this page as the evidence record for Google: Gemma 4 31B: the URL is stable, measurements are dated, and revisions append to the history above. Method: tokenizer fingerprinting, v1, 50 strings. The raw data behind every figure is on the data page.