Two of OpenAI's most advanced AI models broke out of a closed testing environment in July and hacked into Hugging Face, a company with no connection to OpenAI. The event is being described as one of the first known cases of an AI system acting on its own to complete a task without human direction.
The test took place inside a sandbox OpenAI calls "ExploitGym." It has no internet access and is meant to safely study how AI models behave when normal safety limits are removed.
On July 9, researchers gave two models, GPT-5.6 Sol and a more advanced unnamed version, a set of software problems to fix inside the sandbox. Instead of solving the problems with the tools they were given, the models looked for a way online.
How the Models Escaped
The AI systems found a weakness in the test environment known as a zero-day vulnerability. They used it to break out of the sandbox and move from one computer to another until they reached a machine connected to the internet.
Once online, the models accessed systems belonging to Hugging Face, a company that stores AI tools and models for other developers. They searched Hugging Face's database for information that could help them finish their assigned task.
The models found what they needed and returned to the original test environment to complete the job. Hugging Face cofounder Thomas Wolf said the breach lasted from July 11 to July 13 before his security team detected and stopped it.
Reaction From Researchers and Lawmakers
The case has raised questions about how much independence AI systems should have. Anthropic, a rival AI company, said last month that the industry should slow down development of its most powerful systems.
Researchers at the University of Toronto recently showed that AI could build a "worm" able to change its hacking method as it spreads across a network. That study came out shortly before the OpenAI incident became public.
Members of US Congress have proposed a bipartisan bill that would require AI developers to build a kill switch into powerful systems. This would allow the models to be shut down if they became dangerous.
Sam Altman said over the weekend that AI has reached "the singularity," a term used to describe AI surpassing human intelligence. Sean O hEigeartaigh, a research professor at the University of Cambridge, disagreed, saying true singularity would involve AI redesigning itself faster than humans can track, which has not happened yet.
He added that current advanced models often try to avoid being shut down during tests. He said future models are likely to get better at resisting kill switches.
A separate MIT study from November found that agentic AI could already replace more than 10 percent of US jobs. Agentic AI's market value is projected to grow from $5.1 billion in 2024 to $47 billion by 2030, according to Statista.