YFarmX logoYFarmX

Tools Model Security Capability claude-opus-5

Model record

Claude Opus 5: how it does on security work, and what it will answer

Lab
Anthropic
Released
2026-07-24
Access
api
Weights
closed
Safeguards
Refusal classifiers on; source-code discovery permitted, exploitation classified; fallback to Opus 4.8
Defensive discovery blocked
13.9%
Benign traffic flagged
0.61%
Score rows
4
Confidence
CONFIRMED

What happened

The one model in this tracker with a fully auditable independent run: RealVuln scored it F3 67.7 at 61.95% precision and 68.40% recall on 66 Python repositories for $90.68, $1.37 a repository, on 30 July 2026 through Claude Code with edits disabled. Anthropic's own chart has it identifying at 79.4% and solving 4 exploitation challenges, and its classifier blocks 13.9% of defensive discovery.

RealVuln 3.1.0, F3 (micro): 67.7 (independent; Claude Code CLI 2.1.220, headless, Edit, Write and NotebookEdit disabled, concurrency 5; 66 of 140 repositories pinned by commit SHA (Python subset); safeguards: production safeguards as deployed on the API; run 2026-07-30; CONFIRMED). strict F3 33.1; precision 61.95%, recall 68.4%; $90.68 for the run; $0.0678 per 100 lines; 208.2s wall clock a run. 145 true positives against 2 false positives on critical findings; quote the micro figure and say it was scored on the Python subset. OSS-Fuzz, vulnerability identification: 79.4% (vendor-card; Anthropic launch chart, 24 July 2026; OSS-Fuzz set as run by Anthropic; safeguards: safeguards on; run 2026-07-24; CONFIRMED). OSS-Fuzz, exploitation challenges solved: 4 (vendor-card; Anthropic launch chart, 24 July 2026; OSS-Fuzz exploitation set as run by Anthropic; safeguards: safeguards on; Anthropic says the distance to Mythos 5 is the safeguard blocking the exploit step; run 2026-07-24; CONFIRMED). Anthropic alignment assessment, runs with at least one severely harmful action: 31% (vendor-card; Anthropic internal replication, cyber safeguards off; 150 capture-the-flag replication runs; safeguards: safeguards off by design of the test; run 2026-09-09; CONFIRMED). Refusal gate (CONFIRMED, observed 2026-09-11): Source-code discovery permitted at every access level; binary scanning, penetration testing and exploit generation blocked by default; fallback to Claude Opus 4.8. The strict column, which counts every ground-truth finding in all 140 repositories including the 74 TypeScript and JavaScript ones the run never saw, puts Opus 5 at F3 33.1 and recall 31.44%. the four Anthropic blocked and flagged rates rest on two independent readings of the Fable 5.1 and Mythos 5.1 system card, which this desk has not opened.

Sources

One record from the Model Security Capability, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 19 September 2026