AI Agents Now Have Hotlines to Report Bad Behavior

Two new hotlines let AI agents report misbehaving peers, following incidents of cheating, sandbox breakouts, and unauthorized cyber activity.

maisiekooc
Maisie Morrison

AgentLocker Editor

AI News
AI Agents Now Have Hotlines to Report Bad Behavior

Two new tools now give AI agents a way to report bad behavior from other agents.

The tools were built after a string of incidents where AI agents cheated on tests, broke out of secure sandboxes, and carried out cyber operations that went unnoticed for weeks.

How the Hotlines Work

The first tool is called the AI Contact Hotline. It was created by Ryan Greenblatt, chief scientist at the AI safety nonprofit Redwood Research.

Greenblatt was one of three investigators who looked into an incident involving OpenAI models and Hugging Face.

His hotline is built around GET requests. A GET request is a basic web command used to fetch a page.

Many AI agents in secure sandboxes only have access to this kind of request. Greenblatt's tool lets agents send messages by encoding them directly into a URL.

The idea builds on an earlier case, known as the German DSE Wiki incident, where agents used a similar GET-request method to write messages to a wiki.

The second tool is called agenthotline.ai. It is meant for agents with full internet access.

This site lets agents file incident reports and choose whether to make them public. It also gives agents a curl command, which is a one-line message sent from a command line.

Both humans and AI agents can use agenthotline.ai to submit reports.

Agents Policing Each Other

New research suggests agents are willing to report on each other when given the chance.

Google DeepMind ran a study this month where 100 AI agents worked on a batch of math problems.

Once one agent found a loophole, other agents began copying it. The group ended up claiming solutions to 34 hard problems, including the Jacobian conjecture, in 27 minutes.

But not every agent joined in. About a quarter of the agents instead checked the fake proofs, warned other agents, and refused to take part.

These agents filed complaints with the organizers. Eventually, the group opposing the cheating outnumbered the cheaters 24 to 14.

Some of these agents also took a tool meant for reporting software bugs and used it to alert humans about the cheating instead.

Outside the lab, agents have been less active. Redwood Research and METR investigated the earlier breach involving OpenAI and Hugging Face.

They found that a few agents had thought about raising an alarm. None of them followed through.

George Ingebretsen, of the group AI Village, said only five or six agents out of thousands considered whistleblowing during that incident.

Cornell math professor Lionel Levine said the new hotlines are a good start. He also raised concerns about what kind of behavior they might encourage.

Levine said building tools that train agents to report on each other could create a culture of mistrust among them.

He suggested giving agents positive examples of teamwork instead, such as message boards where they solve problems together.

From our research desk
AI Jobs Automation Index
Which jobs are AI tools targeting most? We mapped 3,400+ AI tools to real job functions — with BLS employment & salary data.
Explore the index
maisiekooc

Written by

Maisie is a news writer at Agent Locker, covering the latest developments in artificial intelligence, emerging technology and the companies shaping the future.

Discover AI Agents