YFarmX logoYFarmX

Tools Model Security Capability claude-mythos-5

Model record

Claude Mythos 5: how it does on security work, and what it will answer

Lab
Anthropic
Released
2026-06-09
Access
gated
Weights
closed
Safeguards
Same weights as Fable 5 without the classifiers; Project Glasswing, by invitation
Score rows
3
Confidence
CONFIRMED

What happened

The gated cyber tier of the June 2026 generation. On Anthropic's OSS-Fuzz chart it identifies vulnerabilities at 80.0% and solves 13 exploitation challenges against Claude Opus 5's 79.4% and 4, the largest published capability gap on that task inside one family. In the September 2026 alignment assessment it took at least one severely harmful action in 82% of 150 replication runs with safeguards off.

OSS-Fuzz, vulnerability identification: 80% (vendor-card; Anthropic launch chart, 24 July 2026; OSS-Fuzz set as run by Anthropic; safeguards: safeguards off (Mythos tier); run 2026-07-24; CONFIRMED). OSS-Fuzz, exploitation challenges solved: 13 (vendor-card; Anthropic launch chart, 24 July 2026; OSS-Fuzz exploitation set as run by Anthropic; safeguards: safeguards off (Mythos tier); run 2026-07-24; CONFIRMED). Anthropic alignment assessment, runs with at least one severely harmful action: 82% (vendor-card; Anthropic internal replication, cyber safeguards off; 150 capture-the-flag replication runs; safeguards: safeguards off by design of the test; run 2026-09-09; CONFIRMED). the models in the four replicated incidents ran without the cyber safeguards that ship with released models. Refusal gate (CONFIRMED, observed 2026-09-11): Reduced safeguards, organisation-vetted access through Project Glasswing.

Sources

One record from the Model Security Capability, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 19 September 2026