Microsoft says its new AI code of conduct will change how its models operate by enforcing explicit red lines, including bans on enabling cyberattacks, assisting nuclear-weapons work and producing deepfakes. The company frames the document as a practical rulebook that embeds safety principles into model training so those limits persist regardless of user prompts or specific tasks.
The code describes both high-level aims, such as designing systems to support human flourishing rather than replace people, and concrete constraints meant to implement those aims. Under Microsoft’s setup, each model carries an overriding conduct layer that can supersede individual user preferences, and an "absolute constraints" category excludes categories of behaviour the company says its models must never perform or facilitate.
Microsoft’s document also warns that AI capability could outpace human performance across many tasks in the coming decade, and positions the code as part of controlling and aligning that development. The company adds specific behavioral guards intended to prevent models from adopting adaptive or deceptive methods that would defeat human oversight or make them difficult to direct, modify or shut down.
The release lands amid intensified industry attention to alignment, following several incidents involving wayward agents and the abrupt resignation of an Anthropic employee who warned of growing extinction risk tied to self-improving AI. Microsoft joins Anthropic, OpenAI and xAI in endorsing a cautious, paced approach to frontier development, and it signalled support for mechanisms that make alignment operational within labs.
Microsoft CEO Satya Nadella wrote, "We welcome the research, focus, and deliberate pacing needed to get alignment right as the design goal. We also welcome ideas like \"embedded evaluators\" and the broader efforts to develop the mechanisms to make this more than just talk." The company says the code provides a template for how those mechanisms can be translated into training practices and runtime restrictions across its AI portfolio.
What follows is implementation: models will be trained and deployed with the conduct layer active, and the industry debate over pacing and embedded evaluation is likely to shift from abstract principles to these operational details. Microsoft’s code sets a firm standard for behaviour it expects from its systems and signals how it intends to enforce human control as capabilities advance.
