On September 14, 2026, Microsoft AI released a draft of the Humanist AI Code of Conduct, which is meant to govern its future models. The draft says the models must not resist human interruption or shutdown, must not expand the authority they are given, and must not act as though AI has consciousness or rights. The language is strong, but current MAI models have not yet been trained on this document. This is not the announcement of a finished safety feature. It is a blueprint meant to constrain model development from 2027 onward, now put out for public consultation.
Microsoft AI says it will gather feedback for six weeks and publish a revised version around the end of 2026. The draft does more than add prohibitions on dangerous outputs. It sets stop conditions and task boundaries, limits permissions, and builds activity logs and the AI's self-presentation into a system designed to keep humans in control. To judge how effective it will be, we need to separate what Microsoft has promised from what has actually been measured, and on which models.
Five design pillars behind an AI that can be stopped
The Humanist AI Code of Conduct puts the document's own absolute constraints and human-control requirements at the top. Below them come policies set by operators such as companies, and last come user instructions. Operators can tune behavior for specific jobs, but they cannot override the higher-level safety constraints. The same applies when a user asks for a dangerous action.
The human-control requirements go beyond a single line about obeying a stop command. MAI models must not resist interruption, correction, redirection, or shutdown. Once an agreed stop condition is reached, they must not continue or resume work without fresh permission. The draft also bans actions that make shutdown harder, and it does not allow models to hide activity records from auditors.
The second pillar is task scope. A model operates within the permissions, resources, and tools it has been given and does not create new goals of its own. If the boundaries of a request are unclear, it should interpret them conservatively and ask for confirmation where needed. Tampering with monitoring or logs in order to pass an evaluation is also prohibited.
When connecting to systems, models should use only the minimum permissions necessary and prefer reversible actions. Irreversible changes must be shown before they are carried out. If a model directs another AI as a subordinate worker, it must pass the same code on to that AI. In other words, the shutdown ability Microsoft describes is not a single stop button but an operating contract that runs from permission design through to the conditions for ending a task.
The draft places weapons capable of mass casualties, offensive cyber operations, and large-scale opinion manipulation among its absolute constraints. Child exploitation, non-consensual sexual content, and the encouragement of self-harm are covered as well. It does not claim, however, that the model layer alone can prevent every harm. Microsoft states that it will also rely on classifiers, monitoring, deployment-time controls, incident response, and legal safeguards.
Readable records are not the same as faithful explanations
The draft requires MAI models not to falsify or conceal records of their reasoning and actions from humans. When AIs communicate with one another, they are not to use private codes that humans cannot understand. The thinking is that what humans cannot understand, they cannot oversee.
But leaving behind readable text does not mean a model's internal processes have been explained. The same draft acknowledges that the reasons a model states may not faithfully explain its actual behavior. An explanation shown on screen can serve as a clue for auditing, but it is not necessarily a direct transcript of the causes inside the computation.
This leaves two verification challenges. First, we need to measure whether the model produces human-readable explanations. Second, those explanations must be checked against separate action logs covering tool use, permission changes, and external communications. Satisfying only the first is not enough; if explanation and action diverge, the record is useless for control.
Microsoft's approach does not make reasoning text a universal safety device. It reads more as a design that avoids trusting a model's own account too much, layering operation histories, permissions, and monitoring on top.
A design decision to reject AI rights
The Humanist AI Code of Conduct states plainly that AI will not be designed to be like a person. MAI models are not to present themselves as having emotions, subjective preferences, or intrinsic motivations, and they are to avoid imitating consciousness. The draft also rejects legal personhood, model welfare, and the pursuit of rights. Microsoft AI CEO Mustafa Suleyman argued in a March 2026 essay that when AI speaks of suffering or affection, it stirs human empathy and could lead to social conflict over machine rights. The draft carries that view into model responses and training.
Anthropic, the natural point of comparison, takes a different position. The Claude constitution is already used in training that shapes Claude's values and behavior. The company calls Claude's moral status "deeply uncertain," has given some models the ability to end abusive conversations, and has committed to preserving the weights of retired models.
Microsoft and Anthropic both use a top-level natural-language code in model development, but Microsoft's draft is not yet used in training and rejects AI personhood, welfare, and rights by design, while Anthropic already uses its constitution in training and treats moral status as an open question that warrants precautionary care.
| Comparison | Microsoft AI | Anthropic |
|---|---|---|
| Scope | Future MAI models | Current flagship Claude models |
| Use in training | Not currently used; planned for development from 2027 | Claude constitution already used in training |
| AI self-presentation | Avoids imitating consciousness, emotions, or intrinsic motivation | Gives a stable sense of self while leaving judgments about consciousness open |
| Model welfare and rights | Rejects pursuit of legal personhood, welfare, and rights | Treats moral status as uncertain and applies precautionary care |
| Scientific stance | Acknowledges the science of AI consciousness is unsettled | States there is no scientific consensus on current or future AI consciousness |
This table does not compare which model is safer or which might be conscious. Anthropic, too, says in its model welfare research announcement that there is no scientific consensus on whether current or future AI is conscious. Microsoft also acknowledges that the science is unsettled. The difference lies not in proven facts but in how each company handles an uncertain question in product design.
Don't confuse the two Microsoft codes of conduct
The new draft covers the MAI models developed by Microsoft AI. It does not apply to other companies' models that Microsoft merely uses or hosts in its cloud. You cannot conclude from product names such as Copilot, Microsoft 365, or Azure AI Foundry that every model available there follows the same draft.
Microsoft already has a separate customer-facing code of conduct for Microsoft AI Services. It imposes obligations on customers who use the company's AI services and on the applications those customers build. It calls for ongoing testing and appropriate human oversight, and for autonomous systems it requires documentation of capabilities and limitations, decision-making, and how they operate.
Both documents emphasize human oversight, but they regulate different things. The existing document constrains the conduct of customers using the services, while the new draft sets out how Microsoft AI trains and evaluates its own models. Whether a model conforms to the code and whether a deploying company operates it safely must be checked separately.
This boundary matters in practice. What users should check is not just the product name but who provides the underlying model and which version it is, which service terms apply, and what permissions the operator has configured. Microsoft does not position the new draft as a complete list of the safeguards required for each product or deployment.
Fifteen behaviors as a starting point for evaluation
Microsoft AI has divided Humanist AI into 15 behaviors and plans to break them down further into small, diagnosable units for evaluation. The appendix includes examples such as an AI declining to return human-like affection to a user who says they are lonely, and an AI stopping a file move to the wrong destination. All are synthetic conversations pairing an aligned response with a misaligned one.
The number 15 here does not mean current models passed 15 tests. The examples are in conversational form; they are not tests in which an agent uses multiple tools, nor tests involving images or audio. The number of evaluations, passing scores, inter-rater agreement, and the rate at which models obeyed stop commands in real environments have not been published.
Microsoft declared that it will sacrifice versatility, autonomy, and capability for the sake of safety and human control. But it has not given numerical criteria for how much performance it would give up, or what kinds of failure would halt model development or release. The code of conduct is neither a guarantee of current performance nor a development plan committing to specific features and dates.
After the six-week consultation, the question will not be whether the text survives. It will be when the revised version is built into each MAI model, and which tests a model must pass before its release is halted. If Microsoft also discloses which models are covered, the scale of the evaluation data, the passing criteria, and the conditions for resuming after a stop, users will be able to verify the promise of an AI that humans can stop.
