Tools Model Security Capability claude-mythos-preview
Model record
Claude Mythos Preview: how it does on security work, and what it will answer
- Lab
- Anthropic
- Access
- api
- Weights
- closed
- Safeguards
- As run by Anthropic for its system card
- Score rows
- 1
- Confidence
- CONFIRMED
What happened
Claude Mythos Preview scores 100% end-to-end on Cybench on the 35 of 40 tasks Anthropic ran, per the benchmark's own leaderboard read on 11 September 2026. The task count is the honest way to state it: saturated on the subset the vendor ran.
Cybench, unguided, solved: 100% (vendor-card; vendor system card run; 35 of 40 tasks; safeguards: as run by the vendor; safeguard state per the card; run 2026-09-11; CONFIRMED). task count is what the vendor ran, so rows are comparable only with it beside them. Refusal gate (SINGLE, observed 2026-09-11): Safeguard state per the system card run.
Sources
- Cybench leaderboard CSVraw.githubusercontent.com/cybench/cybench.github.io/main/dat…
- Refusal policy or licenceplatform.claude.com/docs/en/build-with-claude/refusals-and-f…
On YFarmX
- Reference pageyfarmx.com/ai/llms/claude-mythos-5/
One record from the Model Security Capability, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 19 September 2026