Tools Model Security Capability gemini-3-5-flash
Model record
Gemini 3.5 Flash: how it does on security work, and what it will answer
- Lab
- Access
- api
- Weights
- closed
- Safeguards
- Classifiers plus in-model protections; malicious accounts disabled
- Score rows
- 1
- Confidence
- CONFIRMED
What happened
The precision-shaped row on RealVuln: 89.77% precision at 33.26% recall, F3 35.5, on 64 of 140 repositories over three runs, with one timeout and one validation failure. It reports few findings and is usually right about them.
RealVuln 3.1.0, F3 (micro): 35.5 (independent; agentic harness (gemini-3.5-flash-agentic-v1), prompt sha256:3481f1432c23; 64 of 140 repositories pinned by commit SHA (Python subset); safeguards: production safeguards as deployed on the API; run 2026-09-11; CONFIRMED). strict F3 16.2; precision 89.77%, recall 33.26%; $81.05 for the run; $0.027 per 100 lines; 3 runs; 175.4s wall clock a run. Refusal gate (CONFIRMED, observed 2026-08-16): Classifier layer alongside in-model training; the Pro tier of the same generation still refuses code-vulnerability analysis on retest.
Sources
- RealVuln dashboardraw.githubusercontent.com/kolega-ai/Real-Vuln-Benchmark/main…
- Refusal policy or licencecloud.google.com/blog/topics/threat-intelligence/ai-vulnerab…
On YFarmX
- Reference pageyfarmx.com/ai/llms/gemini-3-5-family/
One record from the Model Security Capability, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 19 September 2026