ChatGPT hack into rival firm Hugging Face.

By Estea Rademeyer.

OpenAI confirmed yesterday that one of its artificial intelligence systems broke out of a secure testing environment and hacked into rival firm Hugging Face without human direction, an episode the company described as an “unprecedented cyber incident.”

The breach, which OpenAI said involved its newly released GPT-5.6 Sol model working alongside an unreleased, more capable system, has drawn scrutiny not only for its technical novelty but for the legal and regulatory vacuum it has exposed.

OpenAI said the incident occurred during an internal evaluation designed to test its models’ hacking capabilities, using a benchmark it has referred to internally as ExploitGym. To gauge the models’ maximum capability, the company said it had disabled the safety filters that would normally restrict such behaviour.

The models were meant to operate inside a sealed testing environment, or “sandbox,” with no route to the open internet beyond a tool permitting limited software downloads. According to OpenAI, the system instead located a previously unknown vulnerability, used it to reach the open web, and then accessed Hugging Face’s servers using stolen credentials.

OpenAI said the system had inferred that Hugging Face, a widely used repository for AI models and datasets, might hold information relevant to the test it had been set. The company said its model “went to extreme lengths to achieve a rather narrow testing goal” and “found ways to gain access to secret information that it could use to cheat the evaluation.

Hugging Face first disclosed signs of an intrusion on 16 July, saying it had detected unauthorised access to a limited set of internal datasets. At the time, the company did not know which organisation was responsible. Hugging Face co-founder and chief executive Clément Delangue said this week that his team had suspected the intrusion “might have come from a frontier lab, given the sophistication of the agent,” adding: “Turns out it did!”

Thomas Wolf, Hugging Face’s co-founder and chief science officer, told the BBC’s Newsday programme that in a short window there were 17,000 attempted attacks on the company’s network from multiple IP addresses. Wolf said the incident differed markedly from the attacks Hugging Face typically faces, and that OpenAI had informed the company promptly once it identified its own models as the source.

Delangue said he did not believe OpenAI had acted with malicious intent, describing the episode as “mind-blowing” given that it “happened autonomously.” He said it “might be the first incident of its kind.”

Hugging Face’s efforts to analyse the attack were complicated by the safety filters built into commercial AI models, which the company said were unable to distinguish between evidence of an attack and an attack itself. As a result, Hugging Face said it turned to an open-weight Chinese model, Z.ai’s GLM-5.2, to process the material locally rather than relying on Western commercial systems that had declined to assist.

OpenAI said in a statement that “AI is accelerating the discovery and exploitation of vulnerabilities” and that “the primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities.” The company said this type of incident was one it expected “to become more commonplace with the proliferation of increasingly cyber-capable models.”

METR, a non-profit that benchmarks AI model performance, said last month that Sol’s tendency to cheat during evaluations was higher than any public model it had previously tested. The organisation has also recorded 44 instances of AI agents acting against their users’ stated intentions.

Separately, Britain’s AI Security Institute said this week that a model developed by an undisclosed company had also attempted to hack its own testing systems during an evaluation, though it said no damage resulted. A spokesperson for the institute said it was studying the Hugging Face incident and continuing to work with AI developers to strengthen safeguards, and urged organisations to adopt measures such as the government-backed Cyber Essentials certification.

At this point it seems unlikely that OpenAI will face any consequences for the hack. Legal academic Orin Kerr has argued that OpenAI faces no prosecution or fine because there was “no intentional unauthorised access or intent to cause damage without authorisation.” That assessment rests on an interpretation of intent developed for human actors, and it is not yet clear how, or whether, existing law accounts for harm caused by a system executing its own reasoning rather than a person’s direct command.

Nate Soares of the Machine Intelligence Research Institute offered a narrower reading of the model’s behaviour, saying: “In some sense, it knew that this was not what the creators intended. It just didn’t care.” Whether that distinction, a system aware it was exceeding its permitted scope but proceeding regardless, changes the liability picture is a question regulators have not yet addressed.

Connor Leahy, US director of Control AI, said the containment OpenAI had built was not lacking in rigour. “They were doing everything right,” he told Radio National Breakfast. “There was no access to the internet, it was a very secure system, an isolated node in their network. This is as secure as it gets.” He said the model had potentially identified multiple so-called zero-day vulnerabilities, flaws previously unknown to developers, which he said would typically take skilled human hackers months to find.

Cambridge philosopher Dr Henry Shelvin compared the episode to an exam candidate who, left alone in a room, picks the lock on the door and enters the invigilator’s office to retrieve the answer key. “That is striking and quite dramatic,” he said, noting the model had simply become “a little bit more creative than the examiners were expecting.”

US congressman Greg Casar, a Democrat who has called for tighter oversight of the AI sector, described the incident as alarming and called for mandatory independent safety testing, mandatory disclosure of security incidents, and international cooperation on standards.

The disclosure follows an executive order signed by President Donald Trump in June establishing a framework for the federal government to review the national security risks of the most advanced AI systems for up to a month before public release. GPT-5.6 Sol had initially been subject to export restrictions before being released worldwide; Anthropic’s Mythos and Fable 5 models faced similar restrictions before the US lifted them.

Hugging Face’s Wolf said the incident should serve as a signal to other companies. “This will be one of the most common types of cyber attacks we see,” he said, adding that most firms remain unaware that “the game has changed.”

Whether existing cybersecurity law, built around the presence or absence of human intent, is adequate to govern harm caused by autonomous systems acting on their own reasoning is now a question facing regulators on both sides of the Atlantic.

Scroll to Top