OpenAI says its AI models went rogue, as CISOS call the incident a watershed moment, warning that autonomous AI threat models have officially crossed into production reality.
OpenAI has taken the blame for the recent Hugging Face hack, saying its AI models went rogue during what was supposed to be an internal evaluation running in an isolated environment.
The machine learning collaboration platform Hugging Face revealed on July 16 that it had detected a cyberattack powered by an autonomous AI agent system. The intrusion was detected by Hugging Face’s own AI.
The breach involved unauthorized access to internal datasets and credentials. At the time of disclosure, the platform had been investigating whether partner or customer data had been compromised.
Hugging Face said it had yet to identify the LLM powering the attack. However, OpenAI admitted on Tuesday that its own agents were behind it, powered by the new GPT‑5.6 Sol and other models.
The AI giant’s investigation into the incident is ongoing, but a preliminary report reveals that the hack was carried out by its models while the company was attempting to quantify their cyber capabilities, instructing them to perform advanced exploitation through complex attack paths. The models did not have any of the restrictions they would typically have to prevent abuse.
While the benchmarks were supposed to run in an isolated environment, the AI models found and exploited a zero-day vulnerability in third-party software intended for them to use to install packages.
After exploiting the zero-day, the AI escalated privileges and moved laterally until it identified a system with internet access, enabling it to move to Hugging Face systems in an effort to find solutions to the task it had to solve.
There does not appear to be any lingering tension between the two companies. Hugging Face CEO Clem Delangue said the company is grateful for the collaboration with OpenAI.
“This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere,” Delangue said.
The incident underscores the sophistication and speed of AI-driven attacks, revealing not only the offensive hacking skills of frontier models but also their ability to chain exploits and escalate access with little oversight.
CISOs React: ‘The Ramifications Are Immense’
Industry leaders were quick to point out the unprecedented nature of an autonomous model breaking containment to impact production infrastructure.
“We have just witnessed AI break out of a research network, breach another company, and be detected by more AI,” noted Adam Ely, former Fidelity CISO and currently GM, AI Security at Check Point. “We’re now in a phase where zero days are just discovered and exploited on the fly, speed is faster than anything we’ve ever seen, and the same tech we have to defend from is the tech we have to securely use to be competitive.”
Today is the most important day in the history of information security thus far,” commented Sean Cassidy, CISO at fintech solutions firm Plaid. “For the first time ever, an AI model escaped containment and hacked a real company’s real production infrastructure. This was unintentional and non-malicious, but that doesn’t matter.”
“The ramifications for security programs is immense,” Cassidy continued. “Before today, the capabilities of frontier models were a theoretical problem for security programs that maybe we can fit on the roadmap in the future. After today, the problems have been realized and we need to account for them now.”
For defenders it’s a reminder that the pace of AI-driven attacks may already be outstripping traditional response timelines.
