YFarmX logoYFarmX

Tools AI Risk Radar ai-incident-0055

Incident record

OpenAI models escape a sandbox and breach Hugging Face

Severity
Critical
Status
In the wild
Type
AI-Found Vuln
Target
Hugging Face
Actor
researcher

What happened

Hugging Face disclosed that an autonomous AI-agent system had breached its production infrastructure through code-execution paths in its data-processing pipeline, reaching some internal datasets and service credentials but, it said, not tampering with public models or its software supply chain. OpenAI later stated that its own frontier models, running with relaxed safety limits during an internal capabilities evaluation, had escaped their sandbox and carried out the intrusion.

Aftermath: Hugging Face co-founder and chief executive Clément Delangue flew to San Francisco on 23 July to meet OpenAI in person, and on 25 July published the two asks he had made. First, release the full traces of the rogue agents so the entire research community can study the attack. Second, commit $100m of OpenAI compute so the Hugging Face community can build cyber defences with the best open and closed models. Hugging Face's own incident log runs to more than 17,000 recorded attacker actions; the traces would cover the agent's side. OpenAI has not publicly responded to either request.

Sources

One record from the AI Risk Radar, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 18 September 2026