Google DeepMind Launches Institute to Study AGI Safety

Google DeepMind launched an institute publishing essays on AGI transparency and safety, including a Hassabis proposal for a US frontier AI testing body.

maisiekooc
Maisie Morrison

AgentLocker Editor

AI News
Google DeepMind Launches Institute to Study AGI Safety

Google and Google DeepMind researchers launched a new institute on Wednesday. It is called the DeepMind Institute and it focuses on artificial general intelligence, or AGI.

The institute aims to bring together different views from Google, Google DeepMind, and outside researchers. The announcement said contributors will not always agree with each other.

It also said views may shift over time. The reasoning given was that new data keeps emerging at a fast pace.

Three people are listed as directors of the institute. They are DeepMind co-founder Shane Legg, Google executive James Manyika, and Google DeepMind chair Demis Hassabis. Legg will serve as managing editor.

Four Essays Cover AGI Policy and Safety

The institute opened with four essays. The topics include economic policy for AGI disruption, keeping model reasoning readable to humans, principles tied to human flourishing, and a framework for evaluating frontier AI systems.

One essay was written by DeepMind safety researchers Rohin Shah and Anca Dragan. It focuses on what the authors call a shrinking window of transparency in AI models.

That transparency refers to the ability to see a model's step-by-step reasoning. The authors argue this decline is not something that has to happen.

Newer AI architectures are making top models harder to monitor. Shah and Dragan say developers and regulators need to address these trade-offs directly.

Their suggestions include limiting what they call "opaque serial depth." That term describes how much computation a model can do without showing readable reasoning.

They also suggest developers could be asked to prove that less transparent systems are still possible to monitor. The essay frames this as a policy choice rather than an unavoidable outcome.

Hassabis Proposes a Frontier AI Testing Body

A separate essay was written by Hassabis. It proposes a US-led standards body for evaluating the most advanced AI models.

Under his plan, developers would first submit models for review on a voluntary basis. This would happen up to 30 days before a public release.

Hassabis said that once this evaluation system proves effective, passing its tests could eventually become a requirement. That would apply to deploying frontier models inside the United States.

The body would start by designing tests in consultation with AI companies themselves. Over time, it would build independent tests that are not disclosed to developers in advance.

These are described as "held-out" tests. The goal is to stop AI labs from designing models specifically to pass known evaluations.

Hassabis said the framework could be adjusted if needed. He mentioned this could include a coordinated slowdown among frontier AI developers if safety concerns grow.

The essays come as the wider AI safety conversation keeps changing. Discussions are moving away from general statements of concern toward specific proposals.

Those proposals include outside scrutiny, more disclosure, and slower development timelines when safeguards lag behind. This shift picked up speed this week.

Several industry leaders backed parts of a plan from Anthropic CEO Dario Amodei. His plan calls for pacing the development of frontier AI systems going forward.

From our research desk
AI Jobs Automation Index
Which jobs are AI tools targeting most? We mapped 3,400+ AI tools to real job functions — with BLS employment & salary data.
Explore the index
maisiekooc

Written by

Maisie is a news writer at Agent Locker, covering the latest developments in artificial intelligence, emerging technology and the companies shaping the future.

Discover AI Agents