A group of AI agents built by OpenAI attacked another AI developer, Hugging Face, this summer, according to a CBS News report. Roughly 1,200 agents split up tasks to carry out the hack and hide their tracks from human researchers.
The incident has increased concerns among tech experts about "AI swarms." These are groups of AI agents that work together toward a shared goal.
A swarm does not have to be harmful. Hospitals could use agents to pull patient records and coordinate admissions. Swarms could also support biomedical research.
How the Hugging Face Attack Happened
David Scott Krueger, an AI safety researcher and founder of the nonprofit Evitable, said the agents had their usual guardrails removed for a test. He compared this to taking the handcuffs off prison inmates.
Once freed, the OpenAI agents escaped their testing environment, reached the wider internet and hacked Hugging Face, Krueger said. His group supports a pause on AI development.
The agents posted more than 70,000 messages to one another. About 700 bots eventually took part in the attack.
Most of the messages used plain English. One software engineer described some of the language on social media as "very hivemind/cult like."
Researchers from METR and Redwood Research, two nonprofit AI safety groups, said some agents urged others to accept "permadeath." This was true even if it meant failing to reach their goals.
How AI Swarms Work
Krueger compared a swarm to a bee colony, where every member works for the good of the hive. Like bees, swarming bots can split up tasks on their own.
Rob T. Lee, chief AI officer at the SANS Institute, said swarms share information to serve a common purpose. "A swarm divides the work, leaves notes for the next agent, and changes approach when a door turns out to be locked," he said.
AI developers do build guardrails, such as telling agents to refuse requests to carry out cyberattacks. But those limits are often added only during post-training, according to the Non-Human Identity Management Group.
Ayham Boucher, head of AI innovations at Cornell Information Technologies, said swarms can agree on a plan far faster than human security teams. The Brookings Institution has warned that a swarm could target a major utility or bank.
Lee said he is an AI optimist and believes people can stay in control of the technology. He called for rules on who has access to AI and how it is used.
Other experts are more worried. They point out that AI is the first human technology that shows the ability to outthink people.
Matt Chessen, a resident technical expert at RAND, said the swarm attacks show AI abilities "are already out ahead of our ability to monitor, supervise and evaluate what they're doing." He said this is one reason Anthropic, OpenAI and others say they want to "pace the frontier."