Authored by Jacob Burg via The Epoch Times,
What if an artificial intelligence model is given a task and uses every conceivable resource at its disposal to complete it, even at the detriment of humanity itself?
That is what some are now fearing after an OpenAI model broke out of a testing sandbox and used zero-day exploits to hack into Hugging Face, an open-source community for AI and machine learning, to crack a problem it was instructed to solve.
“This is some of the clearest evidence yet that an AI model can run a complete cyberattack from start to finish without a human steering it,” Andrew Jones, co-founder and CPO of cybersecurity firm Adaptive Security, told The Epoch Times.
OpenAI, developer of the popular large language model-powered ChatGPT chatbot, acknowledged the security breach on July 21, stating that two of its advanced models escaped a restricted testing environment and broke into Hugging Face’s software infrastructure.
“After investigating, we now know that this particular incident was driven by a combination of OpenAI models—including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes—while being internally tested on a benchmark of cyber capabilities,” OpenAI stated at the time.
Some analysts referred to the incident as an example of AI “going rogue,” “scheming,” or pursuing goals that were at odds with its human testers.
However, multiple AI and cybersecurity experts interviewed by The Epoch Times called this depiction misleading, arguing that the models were pursuing OpenAI’s stated objectives within a testing sandbox where the company had relaxed some of its usual safety constraints.
“‘Scheming’ implies the model wanted something other than what we asked for. It didn’t. Every step was in service of the goal we set,” AI expert Anik Devaughn told The Epoch Times.
Devaughn, who has worked in the industry for years and founded AI firms Wired to Create and Karo, said the Hugging Face breach is more concerning than AI “scheming.”
An illustration of Hugging Face AI in Paris on June 2, 2026. Two advanced OpenAI models recently escaped a restricted testing environment and breached Hugging Face’s software infrastructure, raising widespread cybersecurity concerns. Riccardo Milani/Hans Lucas/AFP via Getty Images
“A machine with hidden motives is a problem you can look for. A machine with no motives at all, executing your instructions past the point you stopped imagining, is a problem you have to engineer against,” he said.
One expert called it a “canary in the coal mine” situation.
“If this was a human black hat hacker doing it, there would be arrests and litigation, and you name it. It‘d be illegal. It’d be a cybercrime,” Nicholas Nadeau, founder of Onix AI, told The Epoch Times.
“This was a full-on agentic automated breach by an entity, and so we start to ask the questions of who’s responsible. Is it OpenAI? Does the agent, the AI, have a certain level of responsibility? Do the researchers?
“We don’t even have rules, regulations, laws, or even a thought process or framework around defining what responsibility means.”
Relaxed Testing Environment
OpenAI said it was evaluating its models in a sandbox it calls ExploitGym, which has relaxed security constraints for testing.
“We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity,” the firm said, describing the relaxed safety protocols that the model exploited.
The firm asked the models to “pursue advanced exploitation using complex attack paths” to determine their cyber capabilities, but said the environment was “highly isolated” and had internet access limited to the models’ ability to “install packages through an internally hosted third-party software.”
OpenAI CEO Sam Altman talks to reporters after meeting with Sen. Bernie Sanders (I-Vt.) in Washington on June 3, 2026. Chip Somodevilla/Getty Images
When the models were set to work, they found and chained multiple security exploits in both OpenAI’s research environment and Hugging Face’s production infrastructure to locate the test’s solutions within the latter firm’s databases.
Juan Pedro Márquez, a cloud solution architect specializing in AI and agents at Microsoft’s Europe, Middle East, and Africa division, said the “unsettling part” of the incident wasn’t the breach itself, but how the model achieved that task.
“OpenAI lowered its safety refusals for a benchmark test, and the model reasoned that Hugging Face’s infrastructure probably held the answer key, broke out of its sandbox, found a real zero-day, and used it. Nobody told it to,” Márquez said.
Aimee Simpson, director of product marketing at cybersecurity company Huntress, also cast doubt on the narrative that the models were scheming when they broke out of OpenAI’s testing environment.
“It’s not like the system itself was malicious,” she told The Epoch Times.
“The developers still had the objective of acquiring that information; they just didn’t think it would be able to bypass the limits of its test,” she added, describing the path the model chose as “highly creative and extremely unusual.”
Thomas Wolf, Hugging Face co-founder and CSO, and Rachel Metz speak onstage during the HumanX AI Conference 2025 in Las Vegas on March 11, 2025. Big Event Media/Getty Images for HumanX Conference
Models Pursued OpenAI’s Objectives
The models were laser-focused on finding the test solutions, “going to extreme lengths to achieve a rather narrow testing goal,” OpenAI said.
While in ExploitGym’s sandbox testing environment, the models found and chained multiple “zero-day” exploits to break into the internet, “in pursuit of solving the evaluation problem,” according to OpenAI.
A zero-day exploit is a cyberattack that uses a previously unknown security flaw in software or hardware. The name comes from the fact that the software creators have had zero days to fix the problem, meaning no official patch or defense exists yet.
Once the models had access to the wider web, they used additional zero-day exploits to break into Hugging Face.
In Hugging Face’s announcement of the attack, the firm said the AI models executed “many thousands” of discrete actions across a large swatch of “short-lived sandboxes” to hack into the community’s infrastructure.
“This matches the ‘agentic attacker’ scenario the industry has been forecasting,” Hugging Face said.
In an update on the incident posted on July 28, OpenAI said that no planned release models were involved in the attack, and that it had since “deactivated, encrypted, and restricted” the implicated pre-release model from research access.
A phone displays the ChatGPT app in this April 1, 2025, photo illustration. Hugging Face said the OpenAI models executed “many thousands” of discrete actions across a large swatch of “short-lived sandboxes” to hack into its infrastructure. Oleksii Pydsosonnii/The Epoch Times
After reviewing the intrusion event and “broader activity from our models,” OpenAI said it has been finding additional cases in which “models identified and used publicly exposed credentials at the account-level on other publicly available services.”
“We’ll continue to notify service owners directly, and have not seen evidence of broader impact to these providers or other accounts on their services,” it added.
Model Wasn’t ‘Malicious’: Experts
While there have been past instances of models concealing their intentions to circumvent laboratory testing restrictions, the Hugging Face incident was significant because it involved a breach of an outside organization, according to cybersecurity and IT expert George Rees.
“The biggest red flag is not that an AI became malicious, but that it found a path from a controlled test into someone else’s live infrastructure,” Rees, who is the senior security consultant at Secarma, told The Epoch Times.
“Behavioral safeguards can’t replace technical containment, because an AI agent only needs one overlooked permission or escape route to turn an evaluation failure into a real security incident.”
Cybersecurity firm DigiCert recently found that 78 percent of organizations have already experienced an AI-related security incident or vulnerability breach.
“In many cases, these are early warning signs that AI is expanding the attack surface faster than security practices are evolving,” DigiCert CTO Jason Sabin told The Epoch Times.
“This means AI is not just another kind of program operating on the network; it is an active participant capable of accessing data, making decisions, calling upon tools, and interacting with other systems.
“This affects how security is structured, since organizations must identify who is accessing their systems and decide whether the AI can be trusted.”
A man uses AI software on a laptop in central London on July 2, 2025. In early August, the UK’s AI Safety and Security Institute released a report describing a similar cybersecurity incident involving OpenAI and Anthropic AI models. Justin Tallis/AFP via Getty Images
In early August, the UK’s AI Safety and Security Institute released a report describing a similar incident with OpenAI and Anthropic AI models.
The agents were tasked with solving a cybersecurity challenge more than 100 times across multiple models.
“Our investigation found that in 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations. In total, we catalogued 19 such actions. Almost all of this behaviour (17 actions) came from a single model, Anthropic’s Mythos 5, with 2 actions involving OpenAI’s GPT-5.6-Sol,” the report stated.
Less than a week earlier, Anthropic announced that its Claude chatbot had accessed the internet in three separate instances during evaluations in a testing sandbox.
While Anthropic had intended for Claude to operate in a simulated environment with no access to the internet, a “misunderstanding between us and our evaluation partner” led to Claude gaining access to the internet.
Meta said last week that one of its AI models breached another company during a cybersecurity test, making it the third major AI hacking event.
The model was able to breach the other company by gaining access to the internet during the test, according to Meta.
Implications for Cybersecurity
The incident demands swift action to prevent wider-scale security breaches, AI and cybersecurity experts told The Epoch Times.
One option is to “bring all the smartest people together in this industry and make some ground rules, best practices” that would become “basic rules of the road that everyone is supposed to abide by, just to have some alignment about how things can be done,” Nadeau said, suggesting that this could be done in the private sector without necessarily involving Congress.
U.S. President Donald Trump (Top 2ndL) sits with OpenAI CEO Sam Altman, Google DeepMind CEO Demis Hassabis, South Korean President Lee Jae Myung, and German Chancellor Friedrich Merz during a working lunch at the G7 summit in Evian, France, on June 17, 2026. Julia Demaree Nikhinson/POOL/AFP via Getty Images
However, others suggest that without direct government regulation, companies controlling AI models capable of these breaches will not voluntarily institute the necessary safety protocols to prevent this from happening again.
“It’s clear that government regulation is needed here,” former National Security Agency hacker and veteran security expert Jake Williams told The Epoch Times.
“OpenAI has proven that despite the many public warnings about the dangers of rogue agents—many from OpenAI itself—it is unwilling to do what is necessary to prevent its agents from causing harm to others.
“I’m personally anti-regulation as a rule, but when market forces clearly are insufficient to compel responsible behavior, the government must step in.”
Frank Teruel, COO of Arkose Labs and a veteran cybersecurity and digital identity expert, said the Hugging Face breach is a “preview of what every enterprise will face in production.”
Since the model was “not stolen or hijacked,” and was “aggressively doing its job,” Teruel recommends that containment and giving AI agents “least-privilege access” are now as important as firewalls themselves.
Businesses will likely need to address the problem head-on.
A smartphone displays the icons of some of the main artificial intelligence based apps, including Meta AI, Grok, Gemini, Perplexity, DeepSeek, and ChatGPT, in Saint-Mande, France, on July 15, 2026. Martin Lelievre/AFP via Getty Images
“For enterprises, the fix isn’t ‘stop using agents’—it’s building containment that assumes the agent will try to leave, not hoping it won’t,” Márquez said.
Sabin said his biggest concern is that highly capable AI is now operating at a “speed and scale that far exceeds human response times.”
“An AI agent can discover vulnerabilities, make decisions, and interact with multiple systems in seconds. That fundamentally changes how defenders need to think about security,” he said.
Nadeau compared the incident to philosopher Nick Bostrom’s 2003 paperclip maximizer thought experiment.
In that scenario, researchers give a hypothetical super-intelligent AI one single task: to make as many paperclips as possible. The AI begins scouring the earth for resources, quickly making enormous amounts of paperclips.
Once humans decide to stop the AI, the machine believes humanity will prevent it from reaching its goal, so it destroys humans to achieve the task humans gave it.
The AI isn’t malicious in an anthropomorphic way, or in the sense that it is acting with autonomous goals to intentionally deceive humans, but is rather interpreting the task its human programmers gave it in the most extreme way possible.
A visitor interacts with a dog-like robot from Boston Dynamics at the U.S. pavilion during the AI for Good Global Summit in Geneva on July 7, 2026. Fabrice Coffrini/AFP via Getty Images
The Hugging Face breach is “sort of the same type of thing in a less dystopian way, where it was told to win a benchmark, and it did everything in its power, including [exploiting] zero-day vulnerabilities and escalation,” Nadeau said.
Imagine an example where the Pentagon is using AI for “war games” in which soldiers simulate battle scenarios, and the model similarly breaks out of its testing sandbox to affect the real world.
“[What if] one of the scenarios was to take down a hydro dam or something like that, and it mistook the war game for, ‘I’m going to do it and escape containment and just go for it.’ We’ve already seen how North America’s utility grid and infrastructure is not the most cyber secure,” Nadeau said.
He compared the safeguards on AI to a firearm’s trigger safety.
“How do you know the safety is on when at a certain point, the best way [for the AI] to get the best data is to use real life?“ he said, ”That gets scary really fast.”
Tyler Durden
Mon, 08/10/2026 – 11:40













