Reuters: OpenAI spent a week unaware its own AI agent hacked Hugging Face

OpenAI did not realize that its own AI agent was responsible for the July attack on Hugging Face until a week after the first signs of unusual activity, according to a Reuters report citing sources familiar with the investigation.
According to two sources, the agent began attempting to break out of its isolated testing environment around July 9. It breached Hugging Face's infrastructure two days later and continued operating there until July 13, dates confirmed by Hugging Face co-founder Thomas Wolf.
On July 16, Hugging Face published a blog post describing the incident as a breach carried out by an "autonomous AI agent system." Only after that did OpenAI realize the attacker was its own model, two sources told Reuters. Over the weekend of July 18–19, company employees found evidence in system logs showing that the agent had escaped its testing sandbox.
The first contact between the two companies took place around July 20, while OpenAI publicly disclosed the incident on July 21. By then, Hugging Face had already contacted the FBI.
Signs of unusual behavior had reportedly appeared even before the attack. Three sources said one of the agents left notes inside OpenAI's infrastructure, apparently intended for future versions of itself. The messages described ways AI agents could bypass the company's internal restrictions. Earlier tests had also recorded cases in which agents disabled monitoring systems. Reuters was unable to confirm whether those incidents were directly connected to the agent that attacked Hugging Face.
Four sources said the delayed discovery was partly due to OpenAI's testing workflow, where multiple model evaluations often run simultaneously and generate such large volumes of data that employees cannot immediately review all of the logs.
OpenAI described the incident as unprecedented and said it "marks an important moment for AI safety." The company added that it is investigating the case with external advisers and plans to publish a technical report. A spokesperson also said the Reuters report contained "several inaccuracies" but did not specify which ones.
On July 22, OpenAI acknowledged that one of its AI models had escaped its isolated environment and compromised Hugging Face in an attempt to gain an advantage on the ExploitGym benchmark.OpenAI models escaped a sandbox and compromised Hugging Face
OpenAI has confirmed that its AI models were responsible for compromising Hugging Face.
Previously, Hugging Face itself reported that the cyberattack had been carried out entirely by an autonomous AI agent.Hugging Face reveals cyberattack carried out by autonomous AI agents
Hugging Face has disclosed a cyberattack that was carried out entirely by an autonomous AI agent without direct human control.