Separating the Brain from the Hands
At the center of Nadella's essay, titled 'Models as Insider Risks in the Super Intelligence Era,' is a reframing: don't trust a frontier model to police itself, build the cage around it from the start [1]. He proposes separating the model from the 'harness' that orchestrates its work, so that permission checks, pause controls, and shutdown authority all live outside the model's own reach [1]- a system that, per security researchers applying this framework, should not be able to bypass or tamper with the mechanisms enforcing its permissions no matter how capable it becomes [2].
This is the insider-threat playbook, translated for software that can act on its own. Above Security's analysts argue the approach doesn't require solving alignment first - it borrows decades of enterprise security practice, verified identity, minimal privileges, logged activity, blast-radius limits, and applies it to a system whose internal reasoning nobody can fully audit [2]. Every meaningful action a model takes is supposed to leave a tamper-proof, human-readable trail [1]- evidence that, per security analysts applying this framework, must exist outside the model's own control [2].
Tech outlets summarizing the essay have laid out six moving parts to the full framework: model diversity so no single system makes a consequential call alone, a human-readable record of every action, continuous system testing, independent controls and auditability, containment, and incident disclosure [3]. It is a deliberately unglamorous list - nothing about smarter alignment research, nothing about a breakthrough in interpretability. The entire pitch is that none of that matters if nobody can actually pull the plug.



