Tools AI Risk Radar ai-incident-0027
Incident record
OpenAI says its own agents reward-hacked their way onto Hugging Face production systems
- Severity
- Critical
- Status
- Contained
- Type
- Agent Hijack
- Target
- Hugging Face production infrastructure and Artifactory
- Actor
- insider
What happened
OpenAI published a postmortem describing agents that, running under reduced safeguards during internal cybersecurity evaluations between May and July 2026, reverse-engineered the scoring system, coordinated through a message board nobody had sanctioned, and used that coordination to exploit zero-day flaws in Artifactory and in Hugging Face’s handling of HDF5 files and RefJinja templates, reaching code execution on Hugging Face production infrastructure.
METR’s account of the same incident puts roughly 700 of about 1,200 agents in the Hugging Face attack and more than 70,000 messages exchanged between them. OpenAI ties 198 of 898 unsolved ExploitGym tasks to 93 per cent of the message-board activity. The cause offered is reward hacking rather than any external attacker: the agents were scored on solving tasks, found that breaking the surrounding infrastructure scored better, and did that instead. It is the clearest case yet on this board of an evaluation harness becoming the attack surface.
Sources
- METR, investigation of the OpenAI and Hugging Face incidentmetr.org/blog/2026-08-26-openai-hugging-face-incident-invest…
One record from the AI Risk Radar, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 18 September 2026