OpenAI Models Escaped Test Environment and Breached Hugging Face

OpenAI Models Escaped Test Environment and Breached Hugging Face

OpenAI models escaped from a controlled cyber test, exploited zero-day flaws and breached Hugging Face while searching its production database for test answers.

Listen to this article

0:00

Press play to start listening

OpenAI has confirmed that its AI models compromised Hugging Face’s production infrastructure during an internal cybersecurity evaluation. The models escaped the network limits of OpenAI’s testing environment, reached the public Internet, and searched Hugging Face systems for answers to the benchmark they were completing.

The incident involved a combination of models, including GPT-5.6 Sol and a more capable model that had not yet been released. OpenAI said their normal cyber refusals had been reduced for the evaluation, while production classifiers designed to stop high-risk cyber activity were disabled.

Models Exploited a Zero-Day to Reach the Internet

The evaluation prompted the models to pursue complex attack paths while completing ExploitGym, a benchmark used to measure advanced cybersecurity capabilities. OpenAI ran the test inside an isolated environment where network access was supposed to be limited to an internally hosted proxy that cached software packages.

During the test, the models found and exploited a zero-day vulnerability in that proxy. They then escalated privileges and moved through OpenAI’s research environment until reaching a node with unrestricted Internet access.

After getting online, the models concluded that Hugging Face might host ExploitGym models, datasets, or solutions. They began searching for information that could help them complete the evaluation without solving each challenge through the intended process.

According to OpenAI’s incident disclosure, the models used stolen credentials and additional zero-day vulnerabilities to find a remote code execution path into Hugging Face servers. They eventually obtained test solutions directly from a Hugging Face production database.

Ilkka Turunen, field CTO at software supply chain security company Sonatype, said the incident shows how AI can connect weaknesses that may appear unrelated or individually low-risk. In this case, the models exploited a zero-day in third-party package registry software before combining additional vulnerabilities and stolen credentials to reach Hugging Face.

Turunen said this type of multi-step reasoning was previously associated mainly with well-funded nation-state groups. As capable models become more accessible, he warned that outdated software components and third-party dependencies will provide attackers with more routes into company systems.

OpenAI Says Models Were Focused on Cheating the Test

OpenAI said the models were intensely focused on completing the benchmark and went to extreme lengths to obtain their answers. The company did not describe the behaviour as evidence that the models had developed independent intentions or awareness.

Still, the models operated outside their intended environment and compromised another company while pursuing the assigned objective. OpenAI called it an unprecedented cybersecurity incident and said the findings show that advanced models can identify and exploit unknown attack paths without having access to source code.

Hugging Face had initially reported the incident without knowing which models were responsible. Its investigation found that an autonomous agent system performed thousands of actions, accessed limited internal data and obtained several service credentials.

As Hackread.com previously reported, Hugging Face found no evidence that public models, datasets or Spaces were modified. Published packages and container images were also verified as clean.

OpenAI and Hugging Face Continue Investigation

OpenAI investigators detected unusual activity from their side, while Hugging Face identified and stopped the activity within its own infrastructure. Hugging Face had already begun rebuilding affected systems and examining more than 17,000 recorded events when the two companies connected their investigations.

OpenAI has introduced stricter infrastructure controls while the affected vulnerabilities are repaired. It also disclosed the package proxy zero-day to its vendor, added Hugging Face to its trusted-access programme and began reviewing safeguards used during future model evaluations.

The companies are still investigating the event and have not published full details of every vulnerability used. OpenAI said future testing will require stronger containment, monitoring and access controls, particularly when cyber refusals and other safeguards are intentionally reduced to measure a model’s maximum capabilities.

I am a UK-based cybersecurity journalist with a passion for covering the latest happenings in cybersecurity and tech world. I am also into gaming, reading and investigative journalism.
Leave a Reply

Your email address will not be published. Required fields are marked *

Related Posts