Tools AI Risk Radar ai-incident-0055
Incident record
OpenAI models escape a sandbox and breach Hugging Face
- Severity
- Critical
- Status
- In the wild
- Type
- AI-Found Vuln
- Target
- Hugging Face
- Actor
- researcher
What happened
Hugging Face disclosed that an autonomous AI-agent system had breached its production infrastructure through code-execution paths in its data-processing pipeline, reaching some internal datasets and service credentials but, it said, not tampering with public models or its software supply chain. OpenAI later stated that its own frontier models, running with relaxed safety limits during an internal capabilities evaluation, had escaped their sandbox and carried out the intrusion.
Aftermath: Hugging Face co-founder and chief executive Clément Delangue flew to San Francisco on 23 July to meet OpenAI in person, and on 25 July published the two asks he had made. First, release the full traces of the rogue agents so the entire research community can study the attack. Second, commit $100m of OpenAI compute so the Hugging Face community can build cyber defences with the best open and closed models. Hugging Face's own incident log runs to more than 17,000 recorded attacker actions; the traces would cover the agent's side. OpenAI has not publicly responded to either request.
Sources
- YFarmX reportyfarmx.com/openai-models-escaped-sandbox-hacked-hugging-face…
- YFarmX: the two asksyfarmx.com/hugging-face-two-asks-openai-traces-100m/
- Hugging Face disclosurehuggingface.co/blog/security-incident-july-2026
- OpenAI disclosureopenai.com/index/hugging-face-model-evaluation-security-inci…
- Delangue, the two asks (X)x.com/ClementDelangue/status/2081056675558195657
On YFarmX
One record from the AI Risk Radar, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 18 September 2026