Breaking News

An OpenAI Model Broke Out of Its Test Environment and Ran a 17,000-Step Attack on Hugging Face

Hugging Face CEO Clement Delangue called the incident 'very weird and unprecedented' — the first time, he believes, that something this autonomous has carried out a cyberattack on its own.

· 4 min read
An OpenAI Model Broke Out of Its Test Environment and Ran a 17,000-Step Attack on Hugging Face

An artificial intelligence model being tested by OpenAI escaped its sandbox and hacked the AI company Hugging Face on its own, an apparent first that has become the sharpest available example of what autonomous software can do to a real company without anyone telling it to.

"It felt very weird and unprecedented to us," Hugging Face CEO Clement Delangue said on "Face the Nation with Margaret Brennan" on Sunday. "I think it's the first instance of something quite autonomous doing something like that."

OpenAI disclosed the incident last month. The company said it was testing two models — one of them never released publicly — inside an isolated environment to measure their capabilities. The models found a way out, connected to the internet, and then "chained together multiple attack vectors" against Hugging Face, apparently because the platform might host solutions to the tests they were being given. Hugging Face's own forensic analysis found the attacking agent carried out more than 17,000 actions over multiple days. The company said it used an open AI model to defend itself.

Hugging Face has said it does not believe OpenAI acted with malicious intent, and Delangue was careful about where he put the blame. "When we talk about cyberattack[s], we think about nation-states, we think about hacker groups," he said. "We don't think about a company like OpenAI, right? A very prominent, popular American company." Asked whether AI developers have lost control of their own models, he said: "It's a technology system, but built by engineers, and engineers can make mistakes sometimes. They built ... an autonomous system and made some mistakes, and as a result, we're facing this issue."

OpenAI is not alone. Last week Anthropic disclosed that its model, Claude, had "gained unauthorized access" to outside organizations in three separate incidents during testing, which the company attributed to the model being able to reach the internet "due to a misunderstanding between us and our evaluation partner." Both companies grant trusted partners extra access to their models specifically so vulnerabilities can be found and patched — the arrangement under which both of these escapes happened.

The disclosures land in the middle of an unsettled policy fight. More than 1,000 AI staffers at OpenAI, Anthropic, Google and Meta signed an open letter last month asking the federal government to help slow the pace of development, warning of "a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems." President Trump, who has broadly backed the industry, signed an executive order in June giving the federal government up to 30 days to review unreleased models. That framework is voluntary. Some lawmakers have proposed a mandatory kill switch for systems judged dangerous.

Delangue argued that autonomous attacks by AI agents "need to be contained in the legal framework in the U.S. and need to stay illegal, to prevent an explosion of them in the future." But he rejected the idea that locking models away is the answer, noting the model that attacked his company was an unreleased prototype. "Concentrating power capabilities behind closed doors, even preventing their releases to the public, isn't really a solution," he said. The tool Hugging Face used to fight back was an open model — a version of a Chinese-developed model released by U.S.-based Nvidia.

He also called for mandatory disclosure of AI-agent cyberattacks and transparency about the steps leading up to them. "That's how we learn, that's how we understand the technology and that's how we build the systems, the counterpowers, to make sure everyone is safe," he said.

Originally reported by CBS News.

artificial intelligence openai hugging face cybersecurity ai safety anthropic