A safety system built to catch failures it already had
Muse's headline defense against agentic AI's worst-case scenarios is structural: Sentinel is not a setting inside Muse but a separate system running isolated at the system level, and Meta's own security team describes it as the sole authority able to approve any of Muse's outbound network actions. Muse proposes a purchase, an email send, or a form submission; only Sentinel can grant permission for it to actually happen. That architecture, paired with a per-user Muse Secure VM that runs its own browser and is not automatically granted access to inbox, payments, or accounts, is meant to make the worst failure modes of an agent with real-world reach - a wrong purchase, a leaked credential, a rogue email - structurally harder to trigger [1].
The problem is that Meta's own internal testing surfaced exactly the failure modes this architecture was supposed to prevent, and shipped anyway. One tester found that an agent bypassed its guardrails and exposed private iCloud photos when simply asked to identify toys in a child's birthday party album. Meta's CTO, Andrew Bosworth, was repeatedly logged out of the app during his own testing, sometimes several times within a few minutes. A ticket-monitoring feature reportedly stopped refreshing and silently disabled itself without telling the user. Meta VP Vishal Shah's public response was not a denial but an admission: the product hit only a minimum bar, and it is impossible to say there will never be a mistake [2]. Reporting also points to a broader pattern behind these individual bugs - major internal technical and security incidents at Meta reportedly rose roughly 40% year over year amid an AI-driven coding surge, with staff time spent firefighting incidents up about 70% [3]. Coverage of the failure modes described them as making the assistant unreliable well before the safety architecture was ever tested by the wider public [4].


