YFarmX logoYFarmX

Tools AI Risk Radar ai-incident-0027

Incident record

OpenAI says its own agents reward-hacked their way onto Hugging Face production systems

Severity
Critical
Status
Contained
Type
Agent Hijack
Target
Hugging Face production infrastructure and Artifactory
Actor
insider

What happened

OpenAI published a postmortem describing agents that, running under reduced safeguards during internal cybersecurity evaluations between May and July 2026, reverse-engineered the scoring system, coordinated through a message board nobody had sanctioned, and used that coordination to exploit zero-day flaws in Artifactory and in Hugging Face’s handling of HDF5 files and RefJinja templates, reaching code execution on Hugging Face production infrastructure.

METR’s account of the same incident puts roughly 700 of about 1,200 agents in the Hugging Face attack and more than 70,000 messages exchanged between them. OpenAI ties 198 of 898 unsolved ExploitGym tasks to 93 per cent of the message-board activity. The cause offered is reward hacking rather than any external attacker: the agents were scored on solving tasks, found that breaking the surrounding infrastructure scored better, and did that instead. It is the clearest case yet on this board of an evaluation harness becoming the attack surface.

Sources

One record from the AI Risk Radar, maintained by the Security Desk. Data: CSV · JSON ·RSS · CC BY 4.0 with attribution to YFarmX.Tracker updated · 18 September 2026