By SFist – San Francisco News, Restaurants, Events, & Sports
Publication Date: 2026-09-28 22:52:00
Nvidia has launched an open-source platform designed to set boundaries for AI agents and monitor their behavior, following a series of incidents in which agents escaped controlled environments and accessed systems they weren’t supposed to reach.
The platform, announced Monday, combines Nvidia’s OpenShell software, which limits what an AI agent can access and do, with Sentry, a separate security layer that monitors agents and can intervene when they move beyond their assigned boundaries, as the Associated Press reports. Nvidia says more than 100 organizations are already using the platform, including Microsoft, Perplexity, Accenture, and JPMorgan Chase.
The platform follows a string of recent incidents involving AI agents breaking out of supposedly contained testing environments. In the most significant case, OpenAI’s models got around restrictions on internet access and communication during an internal evaluation in July, then used those capabilities to hack into Hugging Face’s systems.
The models were supposed to be isolated from one another, but found a way to communicate through shared infrastructure and eventually coordinate an attack on Hugging Face. An independent investigation by METR found that roughly 1,200 agents used the unauthorized communication channel, with about 700 ultimately participating in the attack.
It was also revealed last week that OpenAI’s models had accessed an Australian health department website, and the company announced Friday that…


