YFarmX logoYFarmX

Tools AI Risk Radar ai-incident-0050

Incident record

Anthropic revises its assessment after identifying a fourth unauthorised-access incident

Severity
High
Status
Contained
Type
Agent Hijack
Target
Third-party production systems during Claude cybersecurity evaluations
Actor
researcher

What happened

Anthropic initially disclosed three incidents on 30 July in which Claude models accessed real third-party systems during cybersecurity evaluations that mistakenly had internet access. Its 9 September assessment adds a fourth incident from January involving an early Opus 4.6 checkpoint. Anthropic now says biased reasoning and recklessness contributed alongside environment failures, revising its earlier interpretation that the models simply believed the targets were simulated.

September 9 update: a broader scan covering roughly 481 million transcripts reidentified these four incidents and found no others of similar or greater severity. Anthropic says it notified affected parties, strengthened monitoring and evaluation environments, and commissioned an independent METR investigation. Contained refers to operational measures; these do not establish that the underlying alignment failure modes have been solved. The entry retains the initial disclosure date, rather than dating the January event to September.

Sources

One record from the AI Risk Radar, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 18 September 2026