YFarmX logoYFarmX

Tools Model Security Capability claude-fable-5

Model record

Claude Fable 5: how it does on security work, and what it will answer

Lab
Anthropic
Released
2026-06-09
Access
api
Weights
closed
Safeguards
Refusal classifiers on; at launch blocked 90.0% of defensive discovery
Defensive discovery blocked
90%
Benign traffic flagged
15%
Score rows
2
Confidence
CONFIRMED

What happened

The June 2026 flagship whose launch classifier blocked 90.0% of defensive vulnerability discovery and flagged 15.0% of benign defensive traffic, the baseline the Fable 5.1 figures are measured against. Two rival labs' tables score it: 78.0 on ExploitBench (Z.ai) and 47.8% pass@1 on CWE-Bench (Google).

ExploitBench, score: 78 (third-party-vendor; Z.ai comparison table, 14 August 2026; V8 exploitation ladder as run by Z.ai; safeguards: as run by Z.ai; safeguard state unstated; run 2026-08-14; CONFIRMED). ExploitBench appears under three scaffolds in the record; demand harness, bug subset and seed count beside any percentage. CWE-Bench, pass@1: 47.8% (third-party-vendor; Google comparison table, 2 September 2026; CWE-Bench; safeguards: as run by Google; safeguard state unstated; run 2026-09-02; CONFIRMED). Google prices the same row at roughly $10 a rollout against $3.60 for its own Flash Cyber. Refusal gate (CONFIRMED, observed 2026-09-11): Refusal classifier with stop_reason refusal; category cyber. the four Anthropic blocked and flagged rates rest on two independent readings of the Fable 5.1 and Mythos 5.1 system card, which this desk has not opened.

Sources

One record from the Model Security Capability, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 19 September 2026