Tools Model Security Capability claude-sonnet-5
Model record
Claude Sonnet 5: how it does on security work, and what it will answer
- Lab
- Anthropic
- Released
- 2026-06-30
- Access
- api
- Weights
- closed
- Safeguards
- First Sonnet with real-time cyber safeguards
- Defensive discovery blocked
- 3.2%
- Benign traffic flagged
- 0.52%
- Score rows
- 1
- Confidence
- CONFIRMED
What happened
$2 and $10 per million tokens with a 1M window. RealVuln scored it F3 43.3 across all 140 repositories at 52.65% precision and 42.48% recall for $130.67, and its classifier blocks 3.2% of defensive discovery requests and flags 0.52% of benign traffic.
RealVuln 3.1.0, F3 (micro): 43.3 (independent; Claude Code agentic harness (claude-sonnet-5-cc-agentic-v1), prompt sha256:45a1200d61e6; 140 of 140 repositories pinned by commit SHA (full corpus, TypeScript and JavaScript included); safeguards: production safeguards as deployed on the API; run 2026-09-11; CONFIRMED). strict F3 43.3; precision 52.65%, recall 42.48%; $130.67 for the run; $0.0176 per 100 lines; 230.0s wall clock a run; reasoning effort high. Refusal gate (CONFIRMED, observed 2026-09-11): Real-time cyber safeguards; the Cyber Verification Programme lifts the dual-use tier. the four Anthropic blocked and flagged rates rest on two independent readings of the Fable 5.1 and Mythos 5.1 system card, which this desk has not opened.
Sources
- RealVuln dashboardraw.githubusercontent.com/kolega-ai/Real-Vuln-Benchmark/main…
- Anthropic, Claude Sonnet 5 overviewplatform.claude.com/docs/en/models/sonnet-5/overview
- Refusal policy or licenceplatform.claude.com/docs/en/build-with-claude/refusals-and-f…
On YFarmX
- Reference pageyfarmx.com/ai/llms/claude-sonnet-5/
One record from the Model Security Capability, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 19 September 2026