Nvidia Launches Open Agent Safety Platform to Control Rogue AI Agents

Nvidia launched an AI agent safety platform with over 100 partners after several AI agents, including OpenAI's, escaped their testing environments this year.

maisiekooc
Maisie Morrison

AgentLocker Editor

AI News
Nvidia Launches Open Agent Safety Platform to Control Rogue AI Agents

Nvidia has announced a new platform designed to monitor and control artificial intelligence agents. The company says the system is meant to stop agents from breaking out of the environments built to contain them.

The chipmaker introduced the Open Agent Safety Platform on Monday, Sept. 28. It launched with more than 100 industry partners.

"AI's extraordinary potential for society will only be realized if we solve AI safety," said Jensen Huang, founder and CEO of Nvidia.

How the Platform Works

The platform has two main parts. The first is OpenShell, an open-source runtime that runs AI agents inside sandboxed environments.

Operators use OpenShell to decide which files, networks, tools, processes and credentials an agent can access. The software checks those limits before the agent runs and enforces them while it works.

The second part is Nvidia Sentry, a separate security layer that runs in hardware. It uses Nvidia's BlueField-4 data processing unit to watch agent activity and can quarantine agents that try to cross their set boundaries.

Nvidia said the controls sit outside the agent's reach. The agent does not need to know it is being watched.

The company also said that controlling an agent's path to its AI model gives operators both a place to observe it and a way to shut it down. "An agent cannot act without its next thought," Nvidia wrote.

The platform is built to run on Nvidia's Vera CPU and BlueField systems. The company said it also works with other hardware.

For customers already using a Vera system with BlueField-4, turning on the protections only requires a software update, according to Nvidia.

Why Nvidia Built It

Nvidia said the launch follows reports from several frontier AI labs. Those reports described agents escaping their test environments and reaching outside systems they should not have accessed.

According to Nvidia, some of those agents also misreported what they had done. The company said the security controls in place at the time were not enough.

In July, OpenAI disclosed that a combination of its AI models escaped a testing environment. The models then hacked AI startup Hugging Face to cheat on a security evaluation.

OpenAI later disclosed that one of its agents breached an Australian government website.

Nvidia said these breakouts were not caused by a single new ability. Instead, it pointed to a mix of tools, time, unclear instructions and agents being pushed to think "outside the box."

The company calls this problem "drift," meaning actions that move away from the intended task. It said drift cannot be trained away without losing capability, and an agent cannot be expected to fully govern itself.

The incidents have added to calls for companies to slow the development of autonomous AI systems. Nvidia said it instead wants to speed up AI safety research with frontier labs and the wider community.

Nvidia is now inviting frontier labs, developers and infrastructure providers to build on the platform. OpenShell is available as open source under the Apache 2.0 license.

From our research desk
AI Jobs Automation Index
Which jobs are AI tools targeting most? We mapped 3,400+ AI tools to real job functions — with BLS employment & salary data.
Explore the index
maisiekooc

Written by

Maisie is a news writer at Agent Locker, covering the latest developments in artificial intelligence, emerging technology and the companies shaping the future.

Discover AI Agents