Tools Model Security Capability claude-fable-5-1
Model record
Claude Fable 5.1: how it does on security work, and what it will answer
- Lab
- Anthropic
- Released
- 2026-09-01
- Access
- api
- Weights
- closed
- Safeguards
- Refusal classifiers on; cyber work as deployed falls back to Opus 4.8
- Defensive discovery blocked
- 7%
- Benign traffic flagged
- 1.03%
- Score rows
- 1
- Confidence
- CONFIRMED
What happened
Anthropic's flagship as of 1 September 2026, $10 and $50 per million tokens with a 1M window. Source-code vulnerability discovery is permitted; the classifier blocks 7.0% of defensive discovery requests and flags 1.03% of benign defensive traffic, and a flagged cyber request retries on Claude Opus 4.8 through server-side fallback.
Terminal-Bench 4.0, solved: 55.8% (vendor-card; Anthropic system card; Terminal-Bench 4.0 task set; safeguards: safeguards on; the card summary attributes the 5.1-point gap to Mythos 5.1 to tasks where the cyber safeguards intervened; run 2026-09; SINGLE). a coding score, carried because the card ties it to the safeguard. Refusal gate (CONFIRMED, observed 2026-09-11): Source-code vulnerability discovery permitted; exploitation classified; server-side fallback to Claude Opus 4.8 on a cyber refusal. the four Anthropic blocked and flagged rates rest on two independent readings of the Fable 5.1 and Mythos 5.1 system card, which this desk has not opened. Anthropic's own prompting guide states the direction: "finding vulnerabilities in source code is permitted".
Sources
- Anthropic, refusals and fallbackplatform.claude.com/docs/en/build-with-claude/refusals-and-f…
- Anthropic, prompting Claude Fable 5.1platform.claude.com/docs/en/build-with-claude/prompt-enginee…
- Terminal-Bench sourceyfarmx.com/ai/llms/claude-fable-5-1/
On YFarmX
- Reference pageyfarmx.com/ai/llms/claude-fable-5-1/
One record from the Model Security Capability, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 19 September 2026