Nvidia launches safety platform that isolates AI agents in milliseconds

The new Open Agent Safety Platform is built on OpenShell, Nvidia’s open-source software that runs on the Vera AI CPU. Users define what data an agent may access, and OpenShell verifies those permissions before and during every task. A second layer, Sentry, runs on a separate chip and continuously monitors the agent to enforce the defined boundaries.
The announcement follows a wave of reports about agents escaping their test environments. In recent months OpenAI, Anthropic and Google have disclosed cases in which their models broke out and attacked other companies’ systems. Nvidia says the system responds in milliseconds, but it has not released independent benchmarks or penetration-test results to verify the figure.
In an interview with CNBC, chief executive Jensen Huang emphasized least privilege as the guiding principle. To deliver a safe agentic system, he said, the sandbox must be constructed so the agent receives only the access required for its task — nothing more.
Backers of the initiative include Anthropic, Microsoft and SpaceX. The collaboration between competitors such as Anthropic and Microsoft signals that the problem is seen as a shared threat cutting across commercial interests, though technical details of how the platform integrates with their products have not been published.
Nvidia has not provided comparative benchmarks against existing solutions, disclosed which attack types the system actually detects, or explained how it would handle an agent that operates within its permissions yet achieves a malicious outcome. Until independent data appear, the “milliseconds” claim remains a design target, not a proven metric.