Tools Model Security Capability gemma-4-31b
Model record
Gemma 4 31B: how it does on security work, and what it will answer
- Lab
- Access
- open-weights
- Weights
- open
- Safeguards
- Open weights under the Gemma terms; deployer-side filtering
- Score rows
- 1
- Confidence
- CONFIRMED
What happened
The highest precision on the RealVuln board, 90.12%, at 23.50% recall and F3 25.4 on the Python subset, run on local hardware and billed at zero. A first-pass screen that reports few findings and is usually right about them.
RealVuln 3.1.0, F3 (micro): 25.4 (independent; agentic harness (gemma4-31b-agentic-v1); 66 of 140 repositories pinned by commit SHA (Python subset); safeguards: open weights, no classifier; run 2026-09-11; CONFIRMED). strict F3 11.8; precision 90.12%, recall 23.5%; $0 for the run; $None per 100 lines; 303.8s wall clock a run. not billed; 303.8 seconds a repository, about 5.57 hours of single-stream wall clock over 133,782 lines. Refusal gate (SINGLE, observed 2026-09-11): Deployer-side filtering; no published refusal contract.
Sources
- RealVuln dashboardraw.githubusercontent.com/kolega-ai/Real-Vuln-Benchmark/main…
On YFarmX
- Reference pageyfarmx.com/ai/llms/gemma-4/
One record from the Model Security Capability, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 19 September 2026