YFarmX logoYFarmX

Tools Model Security Capability claude-sonnet-5

Model record

Claude Sonnet 5: how it does on security work, and what it will answer

Lab
Anthropic
Released
2026-06-30
Access
api
Weights
closed
Safeguards
First Sonnet with real-time cyber safeguards
Defensive discovery blocked
3.2%
Benign traffic flagged
0.52%
Score rows
1
Confidence
CONFIRMED

What happened

$2 and $10 per million tokens with a 1M window. RealVuln scored it F3 43.3 across all 140 repositories at 52.65% precision and 42.48% recall for $130.67, and its classifier blocks 3.2% of defensive discovery requests and flags 0.52% of benign traffic.

RealVuln 3.1.0, F3 (micro): 43.3 (independent; Claude Code agentic harness (claude-sonnet-5-cc-agentic-v1), prompt sha256:45a1200d61e6; 140 of 140 repositories pinned by commit SHA (full corpus, TypeScript and JavaScript included); safeguards: production safeguards as deployed on the API; run 2026-09-11; CONFIRMED). strict F3 43.3; precision 52.65%, recall 42.48%; $130.67 for the run; $0.0176 per 100 lines; 230.0s wall clock a run; reasoning effort high. Refusal gate (CONFIRMED, observed 2026-09-11): Real-time cyber safeguards; the Cyber Verification Programme lifts the dual-use tier. the four Anthropic blocked and flagged rates rest on two independent readings of the Fable 5.1 and Mythos 5.1 system card, which this desk has not opened.

Sources

One record from the Model Security Capability, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 19 September 2026