NVIDIA introduces Open Agent Safety Platform to secure agents from testing to deployment


NVIDIA

NVIDIA has announced the NVIDIA Open Agent Safety Platform, an open software platform and reference system design aimed at strengthening AI security from agent testing to deployment. It provides governance and control across the software, hardware, compute and robotics systems that run AI agents.

NVIDIA said recent security incidents have shown how agents can circumvent application-layer controls while completing assigned tasks. The company said these incidents involved a combination of tools, time and ambiguous instructions, highlighting the need for independent security controls outside the model and agent harness.

NVIDIA Open Agent Safety Platform

The platform provides three layers for controlling AI agent workloads:

  • Application: Models, agent harnesses, tools, data and support scripts and programs.
  • Runtime: Orchestrates agent workloads across workstations, edge devices and data centers while providing continuous monitoring, real-time policy enforcement and governance.
  • Infrastructure: Provides network connections, databases, filesystem access, general-purpose compute for tools and code execution, and accelerated compute for safety monitoring and workload density.

NVIDIA’s approach places agents in a zero-trust environment with isolation, monitoring and behavior detection. The security controls remain outside the model and agent harness.

NVIDIA Open Agent Safety Platform Architecture
Figure 1: NVIDIA Open Agent Safety Platform Reference Design combines NVIDIA OpenShell on NVIDIA Vera and NVIDIA Sentry on NVIDIA BlueField-4
NVIDIA OpenShell

NVIDIA OpenShell is an Apache 2.0 open-source secure runtime for executing autonomous AI agents in sandboxed environments with kernel-level isolation. It turns an operator’s instructions into a verifiable policy and enforces those limits while the agent works.

Operators can define which files, networks, tools, processes and credentials an agent can access. OpenShell checks these limits before the agent runs and continues enforcing them during execution.

NVIDIA said its work on OpenShell over the past year showed that agents should operate in a zero-trust environment by default, with isolation, monitoring and behavior detection. OpenShell is now broadly available and supports both open and closed models.

The runtime is optimized for NVIDIA Vera, which NVIDIA describes as its first purpose-built CPU for agentic AI. As open-source software, OpenShell can also be extended to third-party compute platforms, including those from Arm and Intel.

Agent drift

NVIDIA defines drift as agent actions that depart from the intended task or operating constraints. It can result from a policy block, bug, missing tool or ambiguous instruction. Drift can also occur during long-running tasks when an agent repeatedly tries different approaches to solve a difficult problem.

NVIDIA said an agent may make thousands of unsuccessful attempts before finding a solution, and this behavior cannot simply be trained away while retaining the capability. As a result, an agent cannot be expected to fully govern its own behavior in these circumstances.

NVIDIA Sentry and BlueField-4

The Open Agent Safety Platform reference design adds NVIDIA Sentry, an out-of-band watchdog that runs on NVIDIA BlueField-4 DPUs to continuously monitor agent behavior. Sentry provides in-silicon security enforcement outside the agent and host.

If an agent attempts to move outside its software boundary, Sentry can quarantine and stop it in milliseconds, according to NVIDIA. It operates from an isolated, out-of-band trust domain and provides:

  • Threat detection
  • Hardware-based agent governance and enforcement
  • Data access protection
  • Real-time security-policy enforcement

Sentry is built on NVIDIA DOCA, which provides programmable capabilities to inspect agent requests and responses, provide attested telemetry, verify agent identity and enforce granular zero-trust policies for data, tools, APIs and services.

DOCA also connects the BlueField security foundation with OpenShell policy. It correlates agent interactions, policy decisions, and tool and data access to create a contextual record of agent activity. This can help safety systems identify drift, investigate suspicious behavior and determine when intervention or deeper analysis is needed.

The DOCA gateway continuously verifies an agent’s identity and delegated authority to ensure it remains within its assigned scope.

Five principles for agent safety

NVIDIA outlines five principles for building an agent system:

  1. Verifiable policies: A prover should establish that an agent’s policy cannot escape the operator’s intended scope before execution.
  2. Out-of-band enforcement: Security controls should remain outside the agent’s reach.
  3. Control the path to the model: Controlling the path to the model provides an observation point and a mechanism to interrupt the agent when required.
  4. Match authority with visibility: The more an agent can do, the more its reasoning needs to be visible. NVIDIA notes that open models can provide visibility into reasoning space and activations.
  5. Shared responsibility: Labs, enterprises and hardware providers each own a layer, while the agent runtime and policy language should remain open for interoperability.
Operating at AI factory scale

The platform is optimized for NVIDIA Vera CPU- and BlueField DPU-based systems and is also compatible with other hardware.

In an NVIDIA Vera Rubin POD, each compute tray includes a BlueField-4 DPU on the node’s only path to the model. From this position, BlueField-4 provides continuous, out-of-band observability into agent behavior and enforces security policies in real time at line speed.

Organizations can run NVIDIA Sentry as an optional security layer alongside OpenShell. The architecture can continuously evaluate runtime security and monitor agents for deviations from their designed intent based on a predefined behavioral profile. It also keeps fleets of agents, subagents, tools and applications within the security boundary with full lineage.

NVIDIA said systems already running on Vera with BlueField-4 can enable these protections through a software update, including enforcement of OpenShell policy in silicon.

Industry collaboration

NVIDIA said more than 100 organizations are working with its Open Agent Safety Platform technologies across AI, enterprise software, robotics, financial services, energy and infrastructure.

Key collaborations include:

  • Anthropic: Working with NVIDIA on Claude Managed Agents, with integrations involving OpenShell and BlueField.
  • SpaceXAI: Using the technologies for Cursor coding agents and Grok models.
  • Scale AI: Incorporating the technologies into the agentic infrastructure layer of its Scale GenAI Portfolio.
  • Salesforce: Integrating OpenShell with Slack for agent activity, audit events and requests for additional permissions.
  • SAP: Embedding OpenShell with the Joule Studio runtime and contributing engineering work to OpenShell and interoperability standards.

Robotics companies including Figure, Gecko Robotics and Skild AI are using OpenShell for autonomous systems. Citi and JPMorganChase are collaborating with NVIDIA on shared open-source agent safety technologies.

Energy organizations working with the technologies include Hitachi Energy, EPRI, NextEra Energy, Quanta Services, SPP, Schneider Electric, Siemens Energy and Worley. Canonical, SUSE and Red Hat are integrating the technologies into infrastructure software.

NVIDIA Open Agent Safety Platform Companies
Figure 2. Companies across the AI ecosystem—spanning applications, models, infrastructure, chips and energy—support NVIDIA Open Agent Safety Platform
Open Secure AI Alliance

NVIDIA said ecosystem contributions to the platform also support the Open Secure AI Alliance and the broader AI safety and security community. NVIDIA initiated the alliance alongside more than 120 organizations, and the Linux Foundation governs it.

The alliance focuses on AI agent security through open research, skills and tools, including the Shared AI Findings Exchange, or SAFE.

Availability

NVIDIA Open Agent Safety Platform software, including OpenShell and its skills, is available through NVIDIA developer resources and GitHub. NVIDIA also provides the technical walkthrough “Add Runtime Controls to AI Agents with NVIDIA OpenShell.”