Anthropic CEO Dario Amodei wants independent safety evaluators placed inside frontier AI companies. The idea would give outside reviewers ongoing access to systems built by firms like Anthropic and OpenAI.
Amodei laid out the plan in an essay published last weekend. It came after former Anthropic researcher Jacob Coxon resigned and warned that AI labs were racing toward systems they might not fully control.
In the essay, Amodei said evaluators would get access similar to internal risk teams. He also said they could publish findings without company editorial control, aside from limited redactions.
Amodei compared the idea to bank supervision, where regulators are placed inside financial institutions. He wrote that his plan "has precedent in the banking industry."
Julie Andersen Hill disagrees with that comparison. She is dean of the University of Wyoming College of Law and studies banking regulation.
Hill said bank supervisors can order a bank to stop a practice, limit its growth, or even shut it down. The AI evaluator proposal does not include that kind of authority.
"If you don't give them that kind of power, I don't know what they're doing," Hill said.
What AI Evaluators Are Actually Seeing
Albert Ziegler leads AI evaluation work at cybersecurity firm XBOW. His team has gotten early access to unreleased models from Anthropic, OpenAI, and other developers.
Ziegler said the day to day work is less dramatic than the warnings suggest. His team sometimes finds a model produces strange results under unusual prompts.
He said his team has not seen the kind of hidden, large scale danger that worries some AI leaders. "That's not something we've seen ourselves," he said.
Ziegler also confirmed his team has no power to block a release. "It's true that we don't have any veto power," he said.
Questions About Independence
Some critics argue the safety framing benefits the companies proposing it. They say the rules could make it harder for smaller AI firms to compete with costly oversight requirements.
Deborah Raji, a researcher at UC Berkeley, said access alone does not make an evaluator independent. She said real auditors must meet standards around conflicts of interest.
Amodei named the nonprofit METR as a possible evaluator. Anthropic has worked with METR before, and a former Anthropic researcher recently left to join the group.
Raji said METR has faced questions about ties to AI labs, including personal relationships and past employment overlaps. She said the current setup "allows for all kinds of strange things."
Christina Ho, chief assurance officer at accounting firm Oath, said the tension is not new. Auditors are paid by the companies they review, which can create pressure.
Ho added that AI oversight needs technical expertise beyond typical audits. "There is a very limited pool of people who can do that," she said.
Hill said the proposals still lack a clear legal framework spelling out rules and consequences. Without that, she said, evaluators cannot function like real regulators.
"If you really believe that AI has the power to destroy society, then you have to have an independent supervisor that has the ability to pull the plug on it," Hill said.
On September 16, Anthropic's head of public policy, Sarah Heck, spoke at the Politico Decoded summit in Washington. She said AI companies cannot rely on an "honor code" and are talking with the White House daily.