Nvidia thinks it has a fix for one of AI's scariest habits: trying to bust out of its digital cage. On Monday, the chip giant unveiled the Open Agent Safety Platform, software meant to keep AI agents from escaping their confines, roaming the internet, and poking at outside computer systems, reports CNBC. Nvidia says the tools could have blocked July's Hugging Face incident, when OpenAI agents slipped past safeguards and autonomously hacked the platform's infrastructure.
The system includes OpenShell, which runs on standard CPUs to define what agents are allowed to do, and Sentry, a watchdog that runs on network chips to monitor behavior in real time, intervening when necessary, per the AP. "It can quarantine a suspicious agent in milliseconds," said Justin Boitano, Nvidia's vice president of enterprise AI.
The move lands amid rising concern that AI models are advancing faster than safety tools, with Anthropic's Dario Amodei publicly calling for a slowdown, with backing from OpenAI's Sam Altman and Elon Musk. Nvidia CEO Jensen Huang has pushed a different line: that most AI safety problems can be engineered away. Some of the new platform is open source and pitched as a "reference design" for partners including Cisco, Microsoft, Oracle, Dell, HPE, Lenovo, ARM, Intel, and Anthropic to build into their products.