The Sandbox-Plus-Watchdog Architecture
Nvidia's Open Agent Safety Platform pairs two purpose-built layers rather than relying on the AI model to police itself. OpenShell, an Apache 2.0-licensed open-source runtime, executes agents inside sandboxed environments with kernel-level isolation while tracing every action and enforcing defined policy on Nvidia Vera CPUs [1]. Sitting alongside it is Sentry, a hardware reference design running on BlueField-4 DPUs that watches agent behavior out-of-band - outside the agent's own execution path - and can quarantine an agent that tries to move past its boundaries within milliseconds [1]. In Nvidia's Vera Rubin POD systems, the BlueField-4 sits on the only network path between the node and the model, so enforcement happens at line speed via Nvidia's DOCA stack rather than depending on the agent voluntarily complying [2]. The design reflects a five-principle framework Nvidia is pushing across the industry: policy that must be provably inescapable, enforcement that lives outside the model, and a shared-responsibility split between labs, enterprises and hardware vendors [2].


