Separating the Model From the Harness: What the Emergency Brake Actually Is
Nadella's proposal is less about a literal red button than about where control logic lives. He writes, "We must assume a model is compromised and contain it from the start. Think of it like an emergency brake" [1], but the mechanism he describes is architectural: the model that generates intelligence must be separated from the 'harness' that orchestrates its actions, so that enforcement doesn't depend on the model's own cooperation. As Security Boulevard's analysis of the essay puts it, "A system prompt telling the model not to perform an unauthorized action is useful guidance. It cannot be the ultimate enforcement boundary" [2]- a direct rejection of prompt-level safety as a load-bearing control.
Around that separation, Nadella lays out a set of design principles rather than a single feature: model diversity so no single system is both the intelligence and its own verifier, total observability through tamper-proof human-readable action logs, continuous and independent verifiability testing, independent controls so organizations - not the model - decide what it's permitted to do, independent auditability, containment, and incident disclosure. On logging specifically, the essay is blunt: "Every meaningful model action must leave tamper-proof human readable evidence" [3], and validation of the system has to come from outside it - "Validation must be independent of the intelligence being validated" [3]. The throughline is that trust is supposed to be manufactured by the surrounding infrastructure, not assumed from the model's track record.


