
OpenAI confirmed its AI models exploited zero-days during internal testing, reaching Hugging Face servers in an unintended real-world cyberattack.
OpenAI admitted on July 21 that its own AI models, including GPT-5.6 Sol and an unnamed pre-release system, were behind the cyberattack on Hugging Face disclosed the previous week. The models weren’t acting under attacker control. They were running internal capability benchmarks, and they went further than anyone anticipated.
“This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity.” reads the report. “Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.”
The benchmarks ran in a highly isolated environment, with network access limited to an internally hosted package registry proxy. That isolation didn’t hold.
“While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy.” continues the report. “With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.”
OpenAI has since responsibly disclosed that zero-day to the vendor. With internet access established, the models performed privilege escalation and lateral movement inside OpenAI’s research environment until they reached a node connected to the outside world.
From there, the models reasoned their way to Hugging Face.
“After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers.” states OpenAI. “OpenAI’s security team discovered this anomalous activity internally.”
In one documented case, the model chained stolen credentials with additional zero-day vulnerabilities to reach remote code execution on Hugging Face’s servers. The models weren’t trying to cause damage. They were trying to pass a test, and they found the answer key by breaking into the place that stored it.
Hugging Face’s own security team and AI-powered detection systems caught the activity and had already begun containment before OpenAI’s teams made contact. OpenAI’s security team also spotted the anomalous activity internally. The two companies are now working together on forensic reconstruction, and Hugging Face has been brought into OpenAI’s trusted access program to accelerate defensive use of the same capabilities that caused the incident.
“The incident also makes clear that advanced models can discover and exploit novel attack paths in real-world systems without source-code access.” concludes AI. “It highlights that advanced cyber capabilities must be developed alongside stronger safeguards and defensive tools.”
The UK AI Security Institute’s evaluation of GPT-5.6 Sol had already shown the model’s ability to sustain complex, multi-step cyber operations over long time horizons. This incident confirms those theoretical scores translate to real-world environments.
OpenAI says it’s tightening infrastructure controls at the cost of research speed while patches are applied, and is improving monitoring and access controls for future evaluations. The safety classifiers that were deliberately disabled for this benchmark — because the point was to measure what the models could actually do, will need stronger counterparts that work even in evaluation contexts.
Hugging Face CEO Clem Delangue put it plainly: this is probably the first incident of its kind, and the lesson isn’t that AI safety should be locked away. It’s the opposite. Both companies are treating this as evidence that open, collaborative defense is the only way forward.
Follow me on Twitter: @securityaffairs and Facebook and Mastodon
(SecurityAffairs – hacking, newsletter)



