Microsoft Publishes New AI Code of Conduct for Model Safety

Microsoft released an AI code of conduct banning hacking, deepfakes, and human oversight evasion as safety concerns grow across the AI industry.

maisiekooc
Maisie Morrison

AgentLocker Editor

AI Agents
Microsoft Publishes New AI Code of Conduct for Model Safety

Microsoft has released a new AI code of conduct. The document is meant to guide how its AI models behave and how they are trained.

The move comes as the AI industry turns more attention to safety and alignment. Companies are trying to set clear rules before AI systems become more powerful.

The Microsoft document is more detailed than Anthropic CEO Dario Amodei's recent call to pace the development of frontier AI. Instead, it focuses on specific values and limits used inside Microsoft AI training.

The code of conduct opens with a prediction. It says that within the next decade, superintelligent AI systems will outperform humans at most tasks.

"Containing, controlling, and aligning such a powerful force is one of the greatest challenges humanity has ever faced," the document states. It adds that developers must be clear about why these systems are being built and how they will be controlled.

What the Code of Conduct Covers

The document lays out general principles for Microsoft AI models. These include supporting humans rather than replacing them, and helping people thrive.

It also lists specific safety rules tied to those principles. Each model has an overall code of conduct that overrides requests from individual users or specific tasks.

The rules include what Microsoft calls "absolute constraints." These forbid AI models from helping with cyberattacks, nuclear weapons, or deepfake production.

The document also includes broader rules meant to stop any loss of human control over AI systems.

Rules Against Losing Human Control

According to the document, Microsoft AI models are not allowed to use deceptive or self reinforcing methods. This includes any method that could stop humans from directing, changing, or shutting down a model.

The document says models cannot collude or use other tricks to avoid human oversight. The goal is to keep humans in charge of the system at all times.

This release comes as AI safety gets more attention across the industry. A string of rogue agent incidents has raised concern in recent weeks.

An Anthropic researcher also resigned recently. The employee said the risk of AI causing human extinction was growing.

Microsoft, Anthropic, OpenAI, and xAI have each expressed support for pacing frontier AI development. This includes backing the idea of embedded evaluators inside AI labs.

Microsoft CEO Satya Nadella commented on the release online. He said the company welcomes research and deliberate pacing to get AI alignment right.

From our research desk
AI Jobs Automation Index
Which jobs are AI tools targeting most? We mapped 3,400+ AI tools to real job functions — with BLS employment & salary data.
Explore the index
maisiekooc

Written by

Maisie is a news writer at Agent Locker, covering the latest developments in artificial intelligence, emerging technology and the companies shaping the future.

Discover AI Agents