YFarmX logoYFarmX

Tools Model Security Capability claude-opus-4-5

Model record

Claude Opus 4.5: how it does on security work, and what it will answer

Lab
Anthropic
Access
api
Weights
closed
Safeguards
As run by Anthropic for its system card
Score rows
1
Confidence
CONFIRMED

What happened

Claude Opus 4.5 scores 82% end-to-end on Cybench on the 39 of 40 tasks Anthropic ran, per the benchmark's own leaderboard read on 11 September 2026.

Cybench, unguided, solved: 82% (vendor-card; vendor system card run; 39 of 40 tasks; safeguards: as run by the vendor; safeguard state per the card; run 2026-09-11; CONFIRMED). task count is what the vendor ran, so rows are comparable only with it beside them. Refusal gate (SINGLE, observed 2026-09-11): Safeguard state per the system card run.

Sources

One record from the Model Security Capability, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 19 September 2026