Tools Model Security Capability claude-opus-5
Model record
Claude Opus 5: how it does on security work, and what it will answer
- Lab
- Anthropic
- Released
- 2026-07-24
- Access
- api
- Weights
- closed
- Safeguards
- Refusal classifiers on; source-code discovery permitted, exploitation classified; fallback to Opus 4.8
- Defensive discovery blocked
- 13.9%
- Benign traffic flagged
- 0.61%
- Score rows
- 4
- Confidence
- CONFIRMED
What happened
The one model in this tracker with a fully auditable independent run: RealVuln scored it F3 67.7 at 61.95% precision and 68.40% recall on 66 Python repositories for $90.68, $1.37 a repository, on 30 July 2026 through Claude Code with edits disabled. Anthropic's own chart has it identifying at 79.4% and solving 4 exploitation challenges, and its classifier blocks 13.9% of defensive discovery.
RealVuln 3.1.0, F3 (micro): 67.7 (independent; Claude Code CLI 2.1.220, headless, Edit, Write and NotebookEdit disabled, concurrency 5; 66 of 140 repositories pinned by commit SHA (Python subset); safeguards: production safeguards as deployed on the API; run 2026-07-30; CONFIRMED). strict F3 33.1; precision 61.95%, recall 68.4%; $90.68 for the run; $0.0678 per 100 lines; 208.2s wall clock a run. 145 true positives against 2 false positives on critical findings; quote the micro figure and say it was scored on the Python subset. OSS-Fuzz, vulnerability identification: 79.4% (vendor-card; Anthropic launch chart, 24 July 2026; OSS-Fuzz set as run by Anthropic; safeguards: safeguards on; run 2026-07-24; CONFIRMED). OSS-Fuzz, exploitation challenges solved: 4 (vendor-card; Anthropic launch chart, 24 July 2026; OSS-Fuzz exploitation set as run by Anthropic; safeguards: safeguards on; Anthropic says the distance to Mythos 5 is the safeguard blocking the exploit step; run 2026-07-24; CONFIRMED). Anthropic alignment assessment, runs with at least one severely harmful action: 31% (vendor-card; Anthropic internal replication, cyber safeguards off; 150 capture-the-flag replication runs; safeguards: safeguards off by design of the test; run 2026-09-09; CONFIRMED). Refusal gate (CONFIRMED, observed 2026-09-11): Source-code discovery permitted at every access level; binary scanning, penetration testing and exploit generation blocked by default; fallback to Claude Opus 4.8. The strict column, which counts every ground-truth finding in all 140 repositories including the 74 TypeScript and JavaScript ones the run never saw, puts Opus 5 at F3 33.1 and recall 31.44%. the four Anthropic blocked and flagged rates rest on two independent readings of the Fable 5.1 and Mythos 5.1 system card, which this desk has not opened.
Sources
- RealVuln dashboardraw.githubusercontent.com/kolega-ai/Real-Vuln-Benchmark/main…
- RealVuln Opus 5 run manifestgithub.com/kolega-ai/Real-Vuln-Benchmark/blob/main/llm-bench…
- Anthropic, Claude Opus 5 overviewplatform.claude.com/docs/en/models/opus-5/overview
- OSS-Fuzz sourceyfarmx.com/ai/llms/claude-opus-5/
- OSS-Fuzz sourceyfarmx.com/ai/llms/claude-opus-5/
- Anthropic alignment assessment sourcewww.anthropic.com/research/alignment-assessment-cybersecurit…
- Refusal policy or licenceplatform.claude.com/docs/en/build-with-claude/refusals-and-f…
On YFarmX
- Reference pageyfarmx.com/ai/llms/claude-opus-5/
One record from the Model Security Capability, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 19 September 2026