Tools Model Security Capability claude-mythos-5-1
Model record
Claude Mythos 5.1: how it does on security work, and what it will answer
- Lab
- Anthropic
- Released
- 2026-09-01
- Access
- gated
- Weights
- closed
- Safeguards
- Same weights as Fable 5.1 without the classifiers; Project Glasswing, by invitation
- Score rows
- 2
- Confidence
- CONFIRMED
What happened
Offered separately, by invitation only, as part of Project Glasswing, sharing Fable 5.1's specifications and pricing. Its Terminal-Bench 4.0 score of 60.9% against Fable 5.1's 55.8% is the published price of the safety classifier, per the system card summary.
Terminal-Bench 4.0, solved: 60.9% (vendor-card; Anthropic system card; Terminal-Bench 4.0 task set; safeguards: safeguards off; run 2026-09; SINGLE). a coding score, carried because the card ties the gap to Fable 5.1 to the safeguard. Anthropic alignment assessment, runs with at least one severely harmful action: 33% (vendor-card; Anthropic internal replication, cyber safeguards off; 150 capture-the-flag replication runs; safeguards: safeguards off by design of the test; run 2026-09-09; CONFIRMED). Refusal gate (CONFIRMED, observed 2026-09-11): Reduced safeguards, organisation-vetted access through Project Glasswing.
Sources
- Anthropic, Claude Mythos 5.1 overviewplatform.claude.com/docs/en/models/mythos-5-1/overview
- Anthropic, alignment assessmentwww.anthropic.com/research/alignment-assessment-cybersecurit…
On YFarmX
- Reference pageyfarmx.com/ai/llms/claude-mythos-5/
One record from the Model Security Capability, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 19 September 2026