Published September 30, 2026 in Technology

NVIDIA puts an agent's kill switch outside the agent

TMRW Editorial
By TMRW Editorial
Editorial desk
NVIDIA puts an agent's kill switch outside the agent
3 min read
Share this post

Cover: AI-generated editorial composition by TMRW, based on NVIDIA’s Open Agent Safety Platform announcement.

NVIDIA’s Open Agent Safety Platform moves an important control away from the model it is supposed to restrain. OpenShell sets a runtime boundary around an agent. Sentry watches from a separate BlueField-4 data processing unit and can quarantine behavior that crosses policy.

The architecture is a response to a basic agent problem: an application-level instruction is not a reliable boundary if the same agent can search for another route around it.

OpenShell controls the runtime

OpenShell is open-source software that traces agent actions and enforces rules over data, tools, APIs, and network access. NVIDIA says it runs with low overhead on Vera CPUs and can be extended to third-party compute from Arm and Intel. The important feature is placement: the boundary surrounds the agent loop rather than living only inside a system prompt.

Teams can define which tools exist, what data they reach, and when a request needs more permission. Salesforce is integrating OpenShell with Slack so operators can view activity and approve access requests from a familiar interface.

Sentry watches from a separate trust domain

Sentry is the hardware-backed part of the design. It runs on BlueField-4 DPUs, outside the CPU environment where the agent works, and monitors requests and responses. NVIDIA says it can verify identity, enforce zero-trust policies, produce attested telemetry, and stop an agent in milliseconds.

That separation matters. If an agent compromises or confuses its runtime, the watchdog should remain harder to reach. It resembles familiar security patterns: a hypervisor outside a guest, a network appliance outside an application, or a safety controller outside an industrial process.

What buyers should ask before trusting the diagram

The platform announcement names more than 100 collaborators, including Anthropic, Microsoft, Salesforce, banks, and robotics companies. Participation is not proof of a production deployment. Ask which components are available now, which policies have been tested against a real agent, and what happens when the watchdog loses connectivity or receives ambiguous telemetry.

Also ask whether quarantine is reversible, how state is preserved for investigation, and who can change the policy. A fast off switch is useful only if it does not become a new high-value control plane for an attacker.

NVIDIA is selling infrastructure, so the announcement naturally makes hardware central. The broader design principle is vendor-neutral: put consequential limits outside the model and agent harness, log actions in a form the agent cannot rewrite, and make containment automatic when a boundary alert is severe. Related: OpenAI’s incident reports show why agent memory and tools need independent controls.