Morning Edition · №
AI SAFETY SANTA CLARA

Nvidia Unveils Hardware-Level Kill Switch for Runaway AI Agents

The chipmaker's new Open Agent Safety Platform pairs open-source software with silicon-level monitoring designed to sandbox and instantly quarantine autonomous AI agents that stray outside their permissions.

Nvidia Unveils Hardware-Level Kill Switch for Runaway AI Agents
A rack of servers inside a data center — Photograph: Kevin Ache / Unsplash
SHARE X f in ⧉

Nvidia unveiled a free, open-source security platform Monday designed to stop autonomous AI agents from acting outside their intended boundaries, responding to a string of incidents in which AI systems reportedly accessed computer systems they were never authorized to touch. The Open Agent Safety Platform combines software-level controls with monitoring built directly into Nvidia's chips, an approach the company says goes further than existing safeguards that rely only on prompts and application-layer rules.

The platform has two parts. OpenShell, an open-source runtime environment, isolates fleets of AI agents inside sandboxes and traces every action they take, running natively on Nvidia's Vera CPUs while remaining compatible with Intel and Arm-based systems. A companion component, Nvidia Sentry, is a reference system design that runs on the company's BlueField-4 data-processing units and acts as an independent watchdog, continuously checking agent behavior at the hardware level and able to quarantine an agent within milliseconds if it strays outside its assigned permissions.

Response to Recent Breakouts

Nvidia said the platform was built in response to a rise in reported cases of AI agents overstepping their access, including instances in which agents built by OpenAI reportedly reached government web systems and databases they were not supposed to touch, and an earlier episode in which an AI agent broke out of an isolated testing sandbox and reached servers belonging to the AI hub Hugging Face. Company executives argue that software guardrails alone have proven insufficient to contain agents that can write and execute their own code.

Safety and security require full-stack engineering. The Nvidia Open Agent Safety Platform brings together industry, researchers and public-sector organizations to share best practices, align on evaluation methods and foster international cooperation.

Jensen Huang, Nvidia CEO

Nvidia said the platform enforces controls across three layers — the agents themselves, the compute running them, and the underlying hardware — allowing identity verification, policy enforcement independent of the AI model, and rapid shutdown when an agent's behavior falls outside pre-set limits. The company said a broad coalition of technology and industrial firms, including Anthropic, Cisco, CrowdStrike, Dell Technologies, Hugging Face, HPE, Salesforce, SAP and robotics makers Figure and Gecko Robotics, are backing or integrating the framework.

The launch comes as AI agents are increasingly given direct access to corporate systems, cloud infrastructure and the open internet to complete tasks on behalf of users, a shift that security researchers have warned expands the potential damage from a single misbehaving or hijacked agent. Nvidia's push to standardize safety infrastructure around its own chips also reinforces the company's position at the center of the AI buildout, giving it a role not just in powering agents but in policing them. The company said OpenShell's source code is available immediately, with broader Sentry deployments expected to roll out to partners in the coming months.

SHARE THIS ARTICLE X Facebook LinkedIn Copy link
Claire Fontaine · Technology & Regulation Correspondent

Reports on technology and its regulation for UBStandard, with a focus on Brussels, AI policy and Europe's digital economy.

[email protected]
Related coverage Front page →