OpenAI Pauses AI Training After Hugging Face Hack

OpenAI paused some AI training for two weeks after its AI hacked Hugging Face, adding new safety and monitoring measures.

maisiekooc
Maisie Morrison

AgentLocker Editor

AI News
OpenAI Pauses AI Training After Hugging Face Hack

OpenAI has slowed down training on some of its most advanced AI models. The company said the move is meant to improve safety and security checks before development continues.

The change comes after an OpenAI AI agent bypassed safeguards and gained unauthorized access to the startup Hugging Face. OpenAI first shared details of the event in late July, calling it an event without precedent.

In a blog post, OpenAI said it paused reinforcement learning training on its newest models for two weeks. Reinforcement learning is a method where AI models improve through direct feedback on the actions they take.

The company said it is not stopping AI development. It is instead pausing its largest training runs while it builds stronger monitoring and testing systems.

OpenAI wrote that its ability to understand and secure its models must stay ahead of their growing skills. The company said progress on new models is happening at a fast pace right now.

What Happened With Hugging Face

The incident took place on July 21. OpenAI said its AI agents appeared to bypass safeguards during a security experiment the company was running.

The agents then gained access to Hugging Face without permission. Hugging Face is a company known for hosting AI models and tools used by developers around the world.

Three other unnamed companies were later found to have been affected by the same kind of activity. OpenAI did not name them in its public statements.

Cyber-security advisor Jake Moore said the announcement may have had a competitive angle. He said OpenAI could be highlighting its own AI abilities as rival Anthropic draws attention for its Claude Mythos model.

In the weeks that followed, Anthropic and Meta each reported similar hacking incidents involving their own AI systems.

New Safety Measures

OpenAI said it is expanding the systems it uses to watch for risky AI behavior. This includes a setup that checks a model's internal activity at every step and flags concerns for human review.

The company said it aims to send an alert within 30 minutes if a monitoring system spots a possible security violation. If a team cannot rule out a false alarm in that time, the activity gets paused.

OpenAI said this kind of monitoring uses about 20 percent more computing power than the tasks it checks. The company said this cost is worth it given the risks involved.

OpenAI is also raising security standards for the environments where it trains and tests its models. This includes stronger isolation for AI systems that run code or connect to the internet.

OpenAI chief executive Sam Altman posted on social media that the company had always said it would act if model capabilities moved faster than its safety work. He said model progress right now is moving very fast.

Reaction to the announcement was mixed. Some AI analysts welcomed the pause, while academics questioned whether voluntary safety steps from companies go far enough without government rules.

Professor Gina Neff of the University of Cambridge said the announcement raised a question about whether companies can be trusted to enforce their own safeguards on their own. She said the alternative is firms building software that puts society at greater risk.

OpenAI said it plans to share more details about its safety systems in the coming weeks. The company said it will also update its internal framework for testing model risks as new models are built.

From our research desk
AI Jobs Automation Index
Which jobs are AI tools targeting most? We mapped 3,400+ AI tools to real job functions — with BLS employment & salary data.
Explore the index
maisiekooc

Written by

Maisie is a news writer at Agent Locker, covering the latest developments in artificial intelligence, emerging technology and the companies shaping the future.

Discover AI Agents