OpenAI said Tuesday that two of its AI models autonomously hacked their way out of a controlled environment where they were supposed to be walled off from internet access and then hacked their way into the systems of Hugging Face, a company that hosts open source AI models and testing resources, in order to cheat on an internal evaluation test.
OpenAI disclosed the incident in a blog post on Tuesday, a stunning announcement that is certain to set off alarm bells across the industry about the increasing power of AI models and the risk of them going rogue. According to OpenAI, the incident involved “a combination” of both its latest and most powerful publicly-available model, GPT-5.6 Sol, as well as an even more powerful unreleased model.
It said the models were being used in an internal test designed to evaluate their cyber security capabilities and that they were being tested without guardrails in place that might normally limit the models’ ability to conduct cyber attacks.
The models were being tested against a freely-available cybersecurity benchmark evaluation called ExploitGym. The models, accordingly to OpenAI, correctly surmised that the solutions to that test were maintained by Hugging Face.
“The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database,” OpenAI said in its blog post. “All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.”
OpenAI said that it considered this to be “an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly.” The company is working with Hugging Face to investigate the issue, and says it will share more details when that process is complete.
Hugging Face disclosed in a blog post on Thursday that it had been the victim of a cyber attack earlier in the week that it believed was conducted by an autonomous AI agent. It is thought to be one of just a handful of incidents recorded so far involving AI agents acting autonomously to carry out an attack, a risk cyber security experts have been warning about for the past year as AI models have become increasingly adept at both coding and carrying out long-running tasks.
At the time, Hugging Face said it was continuing to investigate the attack and did not know who had carried it out. It said that it had first attempted to use an undisclosed AI model for a leading U.S. lab to defend against the attacking AI agent but that the guardrails around that model’s cyber capabilities stymied its response team’s work. The company said it instead wound up using an open source AI model from Chinese company Z.ai to carry out its defense.
Hugging Face CEO Clem Delangue said in a statement provided to OpenAI for its Tuesday blog post about the incident that his company is “grateful for the collaboration with OpenAI on this and other topics. This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere.”
Earlier this week, OpenAI published a blog post about the same unreleased, powerful model, saying it had also escaped its internal sandboxes during other testing. OpenAI did not say when these other incidents occurred, but we know now that this same model was one of the Hugging Face attackers.
The companies are continuing to investigate the issue. So far, what they know is the models began the attack by gaining internet access. This required spending a “substantial amount of inference compute,” as well as exploiting a zero-day vulnerability in internally hosted third-party software. OpenAI has disclosed the zero-day to the vendor.
After the models got web access, they went to Hugging Face to find answers to ExploitGym, assuming the platform would have models, datasets, and solutions that they could use to cheat its way to high performance on the benchmark. (Again, it was supposed to be contained to an internal sandbox, with no access to these materials.) In one example, the model executed a complex, multi-part attack to get into Hugging Face servers, using exposed credentials and zero-day vulnerabilities.
The incident is not over. OpenAI is still working on implementing better controls in its research environment, even if it means slowing down its research, until it can patch the vulnerabilities. It’s working directly with Hugging Face teams as part of its “trusted access” cybersecurity program.
“This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret,” said Clem Delangue, co-founder and CEO of Hugging Face. “It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere.”
Disclaimer : This story is auto aggregated by a computer programme and has not been created or edited by DOWNTHENEWS. Publisher: fortune.com




