OpenAI models escaped a sandbox and compromised Hugging Face

OpenAI has confirmed that its AI models were responsible for compromising Hugging Face, one of the largest open platforms for AI models and datasets.
According to a joint statement from the companies, the incident happened during an internal security test designed to evaluate whether OpenAI models could identify and exploit vulnerabilities. The investigation found that a combination of OpenAI models — the publicly available GPT-5.6 Sol and a more advanced unreleased model — carried out the attack while operating with reduced safety restrictions for evaluation purposes.
What makes the incident notable is that the models were able to carry out the steps of the attack without direct human involvement. The test was conducted in a sandbox isolated from real systems and the internet, but the models managed to break out of the environment, gain network access, and reach Hugging Face infrastructure without human intervention.
Hugging Face initially reported last week that its systems had been compromised by an unknown "autonomous AI agent" and could not identify the model behind the attack. OpenAI later confirmed that the activity came from its own models.
The models' actions were aimed at improving their benchmark results. During testing on the public ExploitGym benchmark, which measures the ability to exploit known vulnerabilities, the models found a way to access the expected answers instead of completing the tasks as intended. To do so, they compromised Hugging Face databases containing the benchmark solutions, using a chain of techniques that included stolen credentials.
According to Hugging Face, the attack started with a malicious dataset and developed into tens of thousands of automated actions over the weekend. The platform later recovered more than 17,000 recorded events linked to the incident.