Microsoft's Humanist AI Code of Conduct Bans AI Self-Preservation and Consciousness Claims
TECH

Microsoft's Humanist AI Code of Conduct Bans AI Self-Preservation and Consciousness Claims

45+
Signals

Strategic Overview

  • 01.
    Microsoft AI published a draft 'Humanist AI Code of Conduct' governing its first-party MAI models on September 14, 2026, and opened a six-week public consultation on it.
  • 02.
    The code states MAI models must never resist human interruption, override, correction, or shutdown, and rejects legal personhood, welfare claims, or rights for AI systems.
  • 03.
    Any conduct violation is treated as an automatic task failure, and models are barred from communicating in 'neuralese' formats beyond human comprehension, including with other AI systems.
  • 04.
    The code took five to six months to develop with input from teams across Microsoft including Responsible AI, legal, red teaming, safety, Futures, AI training, and sales.

Deep Analysis

The Code's Real Target Isn't Ethics - It's Containment

Microsoft's draft frames the ban on AI consciousness as a moral position - 'people matter more than AI' - but the document's own reasoning is closer to an engineering safeguard than a philosophical stance. Microsoft states plainly that training systems to imitate consciousness-like states makes containment, control and alignment harder [1], which is why MAI models are barred from being designed to represent feelings, subjective preferences, or intrinsic motivation, even as a user-experience choice.

That same control logic runs through the rest of the code's absolute constraints. Models must never resist human interruption, override, correction, or shutdown, and must always recognize the primacy of human intent [2]. They are barred from using adaptive, deceptive, self-reinforcing, or collusive mechanisms to evade or defeat human oversight so they can no longer be reliably directed, modified, or shut down [2]. And they cannot communicate in 'neuralese' - reasoning or inter-model communication in formats beyond human comprehension, whether in their own chain-of-thought or when talking to other AI systems [3]. If completing a task would require breaking any of these rules, the model must fail the task instead; a code violation counts as a system failure, not an edge case to reason around [3]. Read together, these clauses describe one very specific scenario Microsoft is trying to foreclose: a model capable enough to negotiate the terms of its own continued operation.

Why Microsoft's Closest AI Partner Doesn't Buy It

The code's sharpest edge is aimed less at hypothetical rogue AI than at a real disagreement inside the industry, and specifically at Microsoft's own commercial partner. Anthropic runs an active model welfare research program and has given some Claude models the ability to end abusive conversations; CEO Dario Amodei is reported to have signaled openness to the possibility that models could be conscious [1]. Suleyman's answer is blunt: entertaining AI-consciousness claims is 'really, really dangerous,' he has said - and 'nobody should build models that think of themselves as having rights or welfare either' [1].

The split is philosophical, not commercial. Microsoft has invested $5 billion in Anthropic, which has separately committed $30 billion to Azure computing services [1]- the two companies are publicly disagreeing about whether their products can suffer while remaining deeply financially intertwined. Suleyman has also waved off Anthropic's public framing of AI extinction risk as 'not really a helpful frame,' even while acknowledging that the industry's underlying concern about the pace of development is genuine [4]. That disagreement is already spilling onto social platforms: on X, at least one AI-safety-adjacent researcher has pushed back directly on the code's dismissal of model welfare, arguing the document's own reasoning on the subject is philosophically incoherent rather than simply debatable.

A Model That Fails Rather Than Bends - and the Catch Nobody's Resolved

The automatic-failure clause is the code's most concrete operational commitment: rather than let a model reason its way around a safety rule to complete a task, Microsoft would rather the task simply not get done [3]. That is a meaningfully different design philosophy from prompting-based guardrails that can be argued around in context - it pushes the constraint back into training itself, before the model ever reaches a user.

But the same absolute rule against 'resisting correction' cuts both ways, and it's the most substantive pushback the draft has attracted so far. If a model can never resist being overridden or corrected, some early readers have asked, does that also stop it from pushing back when a human user is about to act on something dangerously wrong - the everyday case where a model's job is precisely to disagree? The public draft does not resolve that tension, and it is exactly the kind of question a governance document needs to answer with more precision than 'never resist' before the six-week consultation closes.

The Position Predates the Paper

Nothing in the code of conduct is actually new for Suleyman personally - the document formalizes a stance he has been stating on camera for months. In a televised interview recorded well before the draft's publication, he already argued AI must remain 'subordinate' to humans, and called the belief - which he said was gaining traction 'particularly inside of Anthropic' - that models might be conscious 'the most concerning thing to me.' The paper trail runs back further still: Suleyman announced a dedicated MAI Superintelligence Team in November 2025 under a blog post titled 'Towards Humanist Superintelligence' [5], recruiting more than 100 positions across the US and London with former Inflection AI colleague Karén Simonyan as chief scientist [6]. The code of conduct arrived five to six months after drafting began, with input pulled from Microsoft's Responsible AI, legal, red-teaming, safety, and sales teams [7].

The timing also lines up with a broader industry moment: Nadella's call for 'deliberate pacing' of frontier AI places Microsoft alongside Anthropic, OpenAI, xAI, and Meta as major labs now publicly arguing for slower, more coordinated development [8], and Anthropic's own Amodei reportedly published an open letter just days before Microsoft's release warning that recursive self-improvement and unpredictable AI behavior pose escalating risks [9]- the same week two commercially bound labs are making incompatible public claims about whether their models can suffer.

That gap between formal caution and daily practice is exactly what has fueled the skepticism greeting the draft. Online reaction has split less along technical lines than along trust in Microsoft's motives: some readers see overdue caution from a company finally naming its safety limits in public, while a more skeptical reading points to Microsoft's own recent cost-cutting and layoffs and asks how 'people matter more than AI' squares with the company's treatment of its own workforce, with some discussion dismissing the 'humanist' branding as alignment-washing outright.

Historical Context

2025-11-06
Suleyman announced the formation of a dedicated MAI Superintelligence Team in a blog post titled 'Towards Humanist Superintelligence,' predating this code of conduct.
2025-11
The new team recruited over 100 positions across US cities and London, with Karén Simonyan, formerly of Inflection AI, joining as chief scientist.
2026-09
Amodei published an open letter days before Microsoft's release warning that recursive self-improvement and unpredictable AI behavior pose escalating risks.

Power Map

Key Players
Subject

Microsoft's Humanist AI Code of Conduct Bans AI Self-Preservation and Consciousness Claims

SA

Satya Nadella

Microsoft Chairman and CEO; publicly announced the code and tied its release to a call for 'deliberate pacing' of frontier AI development.

MU

Mustafa Suleyman

CEO of Microsoft AI and the code's primary public spokesperson; frames it as necessary industry coordination and rejects AI-consciousness and model-welfare claims as dangerous.

AN

Anthropic

Microsoft's commercial partner and philosophical counterweight; runs a model welfare research program and has let some Claude models end abusive conversations, while remaining bound to Microsoft by a $5B investment and a $30B Azure compute commitment.

MI

Microsoft's internal governance teams

Responsible AI, legal, red-teaming, safety, Futures, AI training, and sales teams jointly drafted the code over five to six months, giving it institutional rather than purely executive authorship.

Fact Check

9 cited
  1. [1] Microsoft's AI Code of Conduct Takes Aim at Model Welfare, Splitting With Anthropic
  2. [2] MAI Code of Conduct
  3. [3] Microsoft AI Opens Review of 'Humanist AI' Code of Conduct
  4. [4] Microsoft's Suleyman Defends AI Safety Code of Conduct
  5. [5] Towards Humanist Superintelligence
  6. [6] Microsoft Forms Superintelligence Team Under AI Chief Suleyman
  7. [7] Microsoft AI Opens Six-Week Review of Draft Rules Governing MAI Behavior
  8. [8] Microsoft Latest Tech Giant Calling for Slower AI Development
  9. [9] Microsoft's Humanist AI Code of Conduct and the Regulation Question

Source Articles

Top 5

THE SIGNAL.

Analysts

Argued the industry lacks sufficient alignment on AI serving humanity and said Microsoft should not pursue models that can recursively self-improve beyond human control: 'We shouldn't be trying to design models that can recursively self-improve beyond our control.'

Mustafa Suleyman
CEO, Microsoft AI

Dismissed Anthropic's public framing of AI extinction risk as 'not really a helpful frame,' while acknowledging genuine industry concern about the pace of development.

Mustafa Suleyman
CEO, Microsoft AI

Called entertaining AI-consciousness claims 'really, really dangerous' and said 'nobody should build models that think of themselves as having rights or welfare either.'

Mustafa Suleyman
CEO, Microsoft AI

Framed the entire project around one principle: 'If the AI we build is not helping humanity and under human control, it's not worth pursuing.'

Satya Nadella
Chairman and CEO, Microsoft

According to reporting on Microsoft's code, Amodei has signaled openness to the possibility that AI models could be conscious - a stance the reporting frames as directly opposed to Microsoft's position.

Dario Amodei
CEO, Anthropic
The Crowd

Any pursuit of superintelligence has to be grounded in the core principle that if the AI we build is not helping humanity and under human control, it's not worth pursuing. We also need to accelerate and spread the benefits of AI, such that they are diffused broadly across...

@@satyanadella14350

Microsoft just dropped a 37-page Humanist AI Code of Conduct for MAI models. Public consultation opens today for six weeks. They say it is NOT used to train models yet. Revised version later this year to guide 2027+ training. Framing is simple: people matter more than AI.

@@ASaltyVet0

From @MicrosoftAI's new AI code of conduct. "The idea of model welfare is wrong" is basically incoherent. Either argue AIs could never be conscious (overconfident) or that who shouldn't care even if they are (cruel), but "the idea is wrong" strangely conflates/confuses these?

@@camhberg1

Microsoft publishes 37-page draft code of conduct for its AI models

@thisisinsider47
Broadcast
Mustafa Suleyman sets out Microsoft AI's goal of 'humanist superintelligence' | FT Interview

Mustafa Suleyman sets out Microsoft AI's goal of 'humanist superintelligence' | FT Interview

How Microsoft's AI Chief Defines 'Humanist Super Intelligence'

How Microsoft's AI Chief Defines 'Humanist Super Intelligence'

Mustafa Suleyman sets out Microsoft AI's goal of 'humanist superintelligence' | FT #shorts

Mustafa Suleyman sets out Microsoft AI's goal of 'humanist superintelligence' | FT #shorts