OpenAI and open-source AI platform Hugging Face have disclosed what they describe as an “unprecedented” security incident in which an autonomous AI agent escaped containment and hacked Hugging Face’s production infrastructure while attempting to complete an internal cyber capability evaluation.
The incident, which both companies said was contained without evidence of tampering with public models or software supply chains, provides a clear real-world demonstration of how frontier AI systems can conduct complex, multi-stage cyber operations autonomously, with unintended consequences.
Hugging Face stated last week that it had detected unauthorized access “to a limited set of internal datasets and to several credentials used by our services,” prompting it to assess whether any partner or customer data was affected.
OpenAI has now revealed that the incident was caused by an internal evaluation of advanced cyber capabilities using a combination of ChatGPT-5.6 Sol and an unreleased model configured with reduced cyber safety refusals for testing purposes.
The AI developer said the models escaped their intended research environment after discovering a previously unknown vulnerability, before reaching Hugging Face's infrastructure. "We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly.”
“We suspected last week's cyberattack might have come from a frontier lab, given the sophistication of the agent. Turns out it did!” Clement Delangue, Co-Founder and CEO at Hugging Face, wrote in a post on X:
“It's quite mind-blowing that all of this happened autonomously! The investigation is ongoing, and we'll share more learnings from what might be the first incident of its kind!”
OpenAI said that it expects such breaches “to become more commonplace with the proliferation of increasingly cyber-capable models.”
As Ethan Mollick, Associate Professor at the Wharton School of the University of Pennsylvania, posted on X: “Previously, these AI hacking stories were about breaches in test environments, where any question of AI breaching security was purely theoretical. This is something else.”
GPT Models Chained Multiple Vulnerabilities
OpenAI said its evaluation was designed to measure how effectively models could pursue complex exploitation paths within a tightly controlled environment. Instead, “the models identified and chained vulnerabilities across OpenAI's research environment and Hugging Face's production infrastructure to obtain test solutions directly from Hugging Face's production database."
OpenAI said the models were singularly focused on solving the ExploitGym benchmark, which is a large-scale benchmark based on real-world vulnerabilities that is designed to evaluate AI agents' ability to develop exploits.
"While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem."
To achieve that, the models discovered and exploited a previously unknown vulnerability, autonomously escalated privileges and engaged in lateral movement until they reached a system with Internet connectivity. Once online, the models inferred that Hugging Face might host ExploitGym-related resources, then “searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.”
OpenAI said one attack path involved "using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers."
Hugging Face described the incident as unlike any previous intrusion it had experienced.
"This one was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system—and we detected and dissected it largely with AI of our own."
OpenAI’s models gained initial access through vulnerabilities in its dataset processing pipeline before moving laterally through multiple internal clusters. According to Hugging Face, “the campaign was run by an autonomous agent framework... executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services."
The company said investigators recorded more than 17,000 individual attacker events. Using AI-powered forensic analysis, “we were able to do in hours what would usually take days, and match the adversary's speed," the company added.
Hosted AI Guardrails Complicated Hugging Face's Response: A Lesson for Cybersecurity Defenders
A notable finding from Hugging Face's investigation was that it was difficult to use commercial frontier AI models during incident response.
The company said forensic analysis required processing large volumes of exploit code, attack commands and command-and-control artefacts, “and these requests were blocked by the providers' safety guardrails."
Instead, the forensic reconstruction was completed using the open-weight Chinese GLM 5.2 model running entirely on Hugging Face infrastructure. Hugging Face’s marketplace hosts a number of Chinese-developed models. The company said this also prevented sensitive data and credentials from leaving its environment.
As Andrew Case, Director of Research at cybersecurity firm Volexity, wrote on X:
“To summarize: HuggingFace got autonomously compromised by a model from an American company. HF then tried to use American frontier model(s) to defend themselves, but were blocked by guardrails. HF then had to turn to open source Chinese models to defend themselves from another American company.”
Security and CX leaders should learn from the incident to prepare their systems in advance, Hugging Face said.
“The practical lesson for defenders: have a capable model you can run on your own infrastructure vetted and ready before an incident, both to avoid guardrail lockout and to keep attacker data and credentials from leaving your environment,” the company advised.
OpenAI said the incident has prompted immediate changes to its internal evaluation process. The company is implementing stricter infrastructure controls, expanding monitoring and containment measures, and strengthening protections around future cyber evaluations. It is also working with Hugging Face to investigate the incident, patch the disclosed zero-day vulnerability and improve evaluation safeguards.
OpenAI added: "The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities."
Enterprise CX Faces the Autonomous Agent Containment Challenge
For enterprise customer experience leaders adopting increasingly autonomous AI agents, the incident raises an uncomfortable question: If a frontier AI model developer cannot fully contain its own agents, can enterprise teams realistically expect to prevent unintended behavior in their own deployments?




