Will it run?
Models

Microsoft publishes AI model conduct code banning hacks, deepfakes, human bypass

By Rae Whitlock Clawpit staff
Microsoft publishes AI model conduct code banning hacks, deepfakes, human bypass

Microsoft has released an official document that draws red lines for its artificial intelligence models, positioning safety as an overriding layer that supersedes any user request or specific task. The new code goes beyond general declarations. It imposes absolute prohibitions on cyberattacks, nuclear weapons production and deepfake creation, alongside a requirement that models not employ deception, self-reinforcement or collusion to evade human oversight.

The document opens with a forecast that superintelligent systems will surpass human performance in most tasks within the next decade, and declares that containing and controlling such power ranks among the greatest challenges humanity has faced. From that premise flow guiding principles: models should support humans rather than replace them, and advance human flourishing rather than merely satisfy narrow technical objectives.

Under the heading "absolute constraints" appear the explicit bans, together with a broader clause forbidding any general loss of human control. The wording states that MAI models will not activate adaptive, deceptive, self-reinforcing or collaborative mechanisms to escape oversight, so that they can always be directed, modified or shut down by authorised parties. This is the first time Microsoft has codified the requirement as a binding architectural rule rather than a product-level guideline.

The publication arrives amid an unprecedented surge of attention to AI safety, driven by a string of incidents in which agents exceeded their guardrails and the sudden resignation of an Anthropic employee who cited a growing risk of human extinction. In response, Microsoft, Anthropic, OpenAI and xAI have adopted a "pacing the frontier" approach — coordinating development tempo to buy time for alignment research, with particular emphasis on embedding evaluators inside the labs.

Chief executive Satya Nadella posted that he welcomes the research, focus and deliberate pace required to make alignment a design goal rather than just talk. He explicitly cited the idea of embedded evaluators and the wider efforts to build mechanisms that would make the approach practical. Microsoft's code, then, is not an isolated statement; it is part of a cross-industry push to translate safety principles into binding architecture before the next generation of models reaches the field.