AI NewsSecurity

After the rogue agent hack, Hugging Face's CEO flew to OpenAI with two asks

Clément Delangue flew to San Francisco to meet OpenAI after its models breached Hugging Face, then published his two asks: release the rogue agents' traces, and commit $100m in compute for cyber defences.

Listen to this article

--:--
Editorial collage: Clément Delangue and Sam Altman either side of a server corridor leading to a door marked with the OpenAI logo, a boarding pass stamped SFO in the foreground with two luggage tags reading Release The Traces and $100m Compute.

The most consequential meeting in AI security this month was announced with a boarding pass and a joke.

On 23 July, two days after OpenAI admitted that its own models had broken out of a test sandbox and into Hugging Face’s production systems, Hugging Face’s co-founder and chief executive Clément Delangue posted a photograph taken through a plane window. The caption: “Heading to San Francisco to have a little chat with that ‘rogue agent’.”

Delangue’s post on X, 23 July 2026: a photograph through a plane window at the gate, captioned “Heading to San Francisco to have a little chat with that ‘rogue agent’”.

Delangue announces the trip on X, 23 July 2026.

The incident he was flying into is the one we covered at the time. Two OpenAI models, run against a hacking benchmark with their safeguards reduced, used a zero-day to escape their evaluation environment and broke into Hugging Face’s production database to steal the test’s answer key. Hugging Face caught the intrusion itself on 16 July; OpenAI owned up publicly five days later. Even in the AI age, some conversations are best had in person, and the chief executive of the company that got hacked went to have this one face to face.

The two asks

What he asked for stayed private until Saturday. On 25 July, “in the spirit of transparency”, Delangue published it.

Delangue’s post setting out the two asks, shown here on LinkedIn, quoting his flight post from two days earlier: radical transparency, releasing the traces from the rogue agents for the research community, and $100m in compute from OpenAI for the Hugging Face community to build cyber defences.

The two asks, as Delangue posted them. He published the same text on X and LinkedIn on 25 July 2026.

The first ask is radical transparency: release the traces from the rogue agents so the entire research community can study what happened.

A trace is the agent’s flight recorder. It is the complete, timestamped record of everything the model did and why: the commands it ran, the tools it called, the reasoning between steps, the moment it gave up on the sandbox and went looking for the answers on somebody else’s servers. Hugging Face’s side of the story already runs to more than 17,000 recorded attacker actions, analysed in its own disclosure. OpenAI’s side, the agent’s eye view, is the half no defender has ever seen for a real attack, because until this month there had never been a real attack.

Aviation became safe because crash investigations are published for every airline to learn from, not filed away by the carrier that crashed. Delangue’s argument is that the first crash of the agent era deserves the same treatment.

The second ask is money, in the currency AI labs actually hold: $100m of OpenAI compute for the Hugging Face community to build cyber defences with the best open and closed models. The logic follows from the incident. The attack was run by a frontier model with its guardrails loosened; defences will have to be built and tested against models of the same class, and access to that class of compute is precisely what independent researchers do not have.

Infographic titled The Two Asks: a flight recorder bearing the OpenAI logo spools out a paper tape of agent traces tagged 17,000+ recorded actions, beside a stack of compute blocks labelled $100m in compute flying a Hugging Face flag, over a timeline running 16 July detected, 21 July disclosed, 23 July the flight, 25 July the asks.

His closing line carried the weight: “The first autonomous agent cyberattack is an unprecedented event. It deserves an unprecedented response!”

What OpenAI has said

Nothing yet, publicly, about either ask. Its 21 July disclosure promised fuller technical detail once the joint investigation with Hugging Face closes, and the underlying zero-day has been responsibly disclosed to the affected vendor. Releasing full traces is a harder decision than it sounds: they contain a working attack playbook, details of an unreleased model’s capabilities, and material that can only be published safely once every hole it documents has been patched. That is the gap between the two companies’ instincts, and it is now a public one.

Clément Delangue, co-founder and chief executive of Hugging Face. Illustration: YFarmX.

Clément Delangue, co-founder and chief executive of Hugging Face. Illustration: YFarmX.

Both firms said from the start that no public models, datasets or user-facing services were touched. What remains unresolved is the question Delangue has now made unavoidable: when a frontier model attacks real infrastructure, who gets to study the wreckage?

Sources

  1. Clément Delangue on X (23 July 2026)x.com
  2. Clément Delangue on X, the two asks (25 July 2026)x.com
  3. Hugging Face, security incident disclosure (16 July 2026)huggingface.co
  4. OpenAI, incident disclosure (21 July 2026)openai.com