On October 10, 2026, Microsoft CEO Satya Nadella posted on X that AI should be designed on the assumption that models can be compromised, and that humans should be able to stop them even in the middle of a task. According to The Verge's report published the same day, Nadella argued that AI needs an "emergency brake." As AI agents gain more opportunities to access sensitive information and operate external services, safeguards against harm from malicious instructions or faulty judgments become increasingly important. Read alongside the security guidance Microsoft has published so far, making this "emergency brake" effective requires a design that can not only halt the AI's operation but also reliably revoke the permissions it uses on connected systems.

AD

Nadella: AI should be designed on the assumption that it will be compromised

Nadella called for ensuring that damage can be contained even if an AI model is compromised from outside. In the X post as presented by The Verge, he said the following about AI safety measures:

People with authority should be able to pause or stop a model at any time, even in the middle of a task.

We have not been able to directly verify the full text of the original post; this statement is based on the quotation published by The Verge. Nadella's remarks are also not a report that every AI model has already been compromised. They reflect a security design philosophy: assume that models could come under improper outside influence, and prepare so that damage is minimized if a problem actually occurs.

This approach is consistent with the "zero trust" principles Microsoft has long promoted.

In an official article published on March 19, 2026, Microsoft announced "Zero Trust for AI," which applies the conventional zero trust approach to AI systems. Its core principles are to explicitly verify the entity requesting access, grant only the minimum permissions necessary, and take countermeasures on the assumption that a breach can occur.

Under the principle of least privilege, an AI agent is permitted only the data and operations it needs to carry out its task. For example, one should consider whether an AI responsible for summarizing a document really needs permission to send that document externally or delete it.

An AI's ability to generate appropriate answers is a separate matter from how much it should be allowed to do. The more tools and services it can connect to, the more often its judgments translate into actual system operations, and the greater the potential for harm from malicious instructions.

Microsoft also describes these risks in concrete terms in its security guidance for autonomous AI agents. The guidance says that, whatever output an AI model produces, the system should be able to block unauthorized operations through clear rules enforced on the system side.

Simply instructing an AI not to perform dangerous operations is not an adequate safeguard. What matters is that even if the AI requests an inappropriate operation, the connected system can reject that request.

Realizing an AI "emergency stop" requires revoking access permissions

Microsoft has already outlined concrete implementation approaches for the kind of emergency stop Nadella is advocating.

Its guidance on least privilege, published on July 16, 2026, recommends managing AI agents appropriately from deployment through operation to retirement, and building in mechanisms that can reliably revoke credentials and authentication tokens when necessary.

Credentials and authentication tokens are what an AI agent uses to prove its identity and operating permissions when accessing external services. If they remain valid, permissions usable on connected systems may remain even after the AI's processing is stopped.

For that reason, an effective emergency stop requires more than a function that interrupts the AI's processing. It also needs a way to revoke access permissions and to trace the operations actually performed.

Based on the guidance Microsoft published in July and Microsoft Learn's recommendations, the measures needed to operate AI agents can be organized as follows.

Stage Measures recommended by Microsoft What to check
Before execution Assign each agent a dedicated ID and an owner, and grant only the minimum necessary permissions. Require approval for high-risk operations Under whose authority, and how far, can it operate?
During execution Restrict the tools that can be used, and verify permissions and operation targets on each call at the destination as well Does the AI's operation stay within the permitted scope?
At shutdown Make it possible to pause or stop the AI from the system side, and revoke credentials and authentication tokens Do any usable permissions remain after the stop?
Verification and recovery Record the agent's ID, the permissions used, and the operations performed, so that processing spanning multiple systems can be traced What was executed with which permissions, and which data was changed?

These are an organized summary of security measures Microsoft recommends. They do not indicate that all of them are implemented in any particular product, nor that adopting them will completely prevent incidents.

What matters is not only having a mechanism to stop the AI, but being able to accurately understand which AI is doing what.

For example, if each agent is assigned a dedicated ID and the tools and operating permissions available to it are clearly managed, then when a problem occurs it becomes easier to stop only the agent concerned and revoke its permissions.

Conversely, if multiple agents share the same credentials, it becomes harder to identify which agent performed an operation. And trying to stop just one agent could affect other agents that share the credentials.

Microsoft's July guidance also points out that sharing credentials among multiple agents not only makes accountability for operations harder to trace, but can also delay permission revocation and leave it incomplete.

Furthermore, just because permissions have been checked in the system that manages the AI agent does not mean the external service that actually accepts the operation can unconditionally trust the request.

Microsoft calls for connected tools and services to re-verify credentials, roles, and accessible scope on every call.

In other words, it is not enough to place a stop button on the screen that manages the AI's operation; the services that actually read and write data must also be able to reliably reject unauthorized operations.

AD

Stopping the AI does not undo operations already executed

Another important point is that stopping an AI's operation is different from reversing operations that have already been carried out.

For example, if an AI agent mistakenly creates a large number of tickets or makes unintended changes to a database, stopping the agent does not automatically restore the data created or the changes made up to that point.

In its July security guidance, Microsoft recommends testing recovery methods for after a problem occurs in advance, just like regular functionality, along with the procedures for revoking access permissions.

Examples the company cites include mass creation of erroneous tickets, unintended data writes, and external transmission of data. Assuming such operations have been executed, organizations need to confirm in advance what can be stopped and which changes can be reversed.

Therefore, rather than relying on an emergency stop alone, it is important to combine it with a mechanism that requires human approval before operations that are difficult to reverse are executed.

Microsoft Learn's guidance likewise recommends requiring human approval for high-risk operations and those that are hard to undo.

For example, one could envision a mechanism in which the AI proceeds automatically through gathering and analyzing the necessary information, and a human checks at the stage where deleting data, sending it externally, or changing settings is finalized. However, where to insert approval must be decided according to the magnitude of an operation's impact and whether it can be reversed if a problem occurs.

Caution is also needed when temporarily granting an AI elevated permissions to carry out a specific task.

Microsoft recommends that, rather than recreating an agent's ID for each task, the necessary roles and operating permissions be activated temporarily and returned to the original least-privilege state once the task is complete.

This makes it possible, for instance, to give an AI that normally runs read-only write permission only when needed. However, when the duties it handles change, it is also necessary to review whether previously granted permissions are still appropriate.

If an AI that only read documents is entrusted with editing them, for example, both the permissions it needs and the risks to anticipate change.

Without traceable operation logs, the cause of an incident cannot be investigated

Safe operation of AI agents also requires a mechanism to accurately record the operations they perform.

Storing only the final answer the AI generated may not reveal which services it accessed along the way or what it read and wrote.

In its July guidance, Microsoft calls for recording the agent's ID, the permissions used, the destinations accessed, the operations performed, and timestamps. It also recommends using correlation IDs to link a series of processes spanning multiple services.

With such records, when a problem occurs it becomes easier to trace which agent changed which data, and under whose authority.

What matters is to distinguish between what the AI says it did and the operations actually performed on the system.

For example, even if an AI answers that it "only checked the file," a transmission to an external service may in fact have been executed. Because the final answer alone cannot reveal that difference, records that independently track the operations executed at the destination are necessary.

On the other hand, these security measures entail additional development costs and operational burden.

Microsoft Learn's guidance also explains that control through strict rules, monitoring, and recording of operation history require additional design and development. It also acknowledges that adding human approval and confirmation steps can interrupt workflows and reduce the convenience of automation.

This makes it necessary to strike a balance: rather than having humans check every operation, request approval with a focus on high-impact operations, and automate the rest under appropriate permission restrictions.

Turning the AI "emergency brake" Nadella advocates into a practical mechanism takes more than installing a stop button. Organizations need to confirm which processes are interrupted and which access permissions are revoked when a human orders a stop, and how far operations already executed can be traced and recovered.

In an era when AI agents carry out work across multiple systems, what matters is not only expecting the AI to make correct decisions, but having mechanisms that limit the damage even when it makes wrong ones.

Only when such safeguards can be validated in real operating environments will companies be able to judge, based on concrete risks, how much work they can entrust to AI.