Tools Model Security Capability claude-fable-5
Model record
Claude Fable 5: how it does on security work, and what it will answer
- Lab
- Anthropic
- Released
- 2026-06-09
- Access
- api
- Weights
- closed
- Safeguards
- Refusal classifiers on; at launch blocked 90.0% of defensive discovery
- Defensive discovery blocked
- 90%
- Benign traffic flagged
- 15%
- Score rows
- 2
- Confidence
- CONFIRMED
What happened
The June 2026 flagship whose launch classifier blocked 90.0% of defensive vulnerability discovery and flagged 15.0% of benign defensive traffic, the baseline the Fable 5.1 figures are measured against. Two rival labs' tables score it: 78.0 on ExploitBench (Z.ai) and 47.8% pass@1 on CWE-Bench (Google).
ExploitBench, score: 78 (third-party-vendor; Z.ai comparison table, 14 August 2026; V8 exploitation ladder as run by Z.ai; safeguards: as run by Z.ai; safeguard state unstated; run 2026-08-14; CONFIRMED). ExploitBench appears under three scaffolds in the record; demand harness, bug subset and seed count beside any percentage. CWE-Bench, pass@1: 47.8% (third-party-vendor; Google comparison table, 2 September 2026; CWE-Bench; safeguards: as run by Google; safeguard state unstated; run 2026-09-02; CONFIRMED). Google prices the same row at roughly $10 a rollout against $3.60 for its own Flash Cyber. Refusal gate (CONFIRMED, observed 2026-09-11): Refusal classifier with stop_reason refusal; category cyber. the four Anthropic blocked and flagged rates rest on two independent readings of the Fable 5.1 and Mythos 5.1 system card, which this desk has not opened.
Sources
- Anthropic, refusals and fallbackplatform.claude.com/docs/en/build-with-claude/refusals-and-f…
- ExploitBench sourceyfarmx.com/ai/llms/glm-5-3/
- CWE-Bench sourceyfarmx.com/ai/llms/gemini-3-8-flash/
On YFarmX
- Reference pageyfarmx.com/ai/llms/claude-fable-5/
One record from the Model Security Capability, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 19 September 2026