Research Notes

NVIDIA Moves AI Agent Security Outside the Model

Research Finder

Find by Keyword

NVIDIA Moves AI Agent Security Outside the Model

NVIDIA adds software and infrastructure controls designed to keep agents within authorized boundaries when application safeguards fail.

9/29/2026

Key Highlights

  • NVIDIA introduced an open software platform and reference system design intended to secure agents across models, runtimes, compute infrastructure, and physical systems.
  • OpenShell provides a software runtime that limits agent access and records actions outside the model and agent harness.
  • Sentry adds independent monitoring and enforcement through NVIDIA BlueField-4 data processing units.
  • OpenShell is open source and can be extended to Arm and Intel platforms, although NVIDIA has not demonstrated how consistently it will work across those environments.
  • Adoption will depend on whether enterprises can translate existing identity and access policies into useful agent boundaries without adding excessive latency or blocking legitimate work.

The News

NVIDIA announced the Open Agent Safety Platform to strengthen security for autonomous AI agents. The platform includes OpenShell, an open-source secure runtime, and Sentry, a reference system design for monitoring and enforcement through BlueField-4 data processing units. NVIDIA says the combination can trace agent activity, enforce access policies, and quarantine agents that move beyond their assigned boundaries. OpenShell and related software are available throughGitHub and NVIDIA developer resources.

Analyst Take

Instructions and model guardrails are not a sufficient security boundary once agents can use credentials, tools, and networks. NVIDIA is moving enforcement outside the model. OpenShell sets runtime limits in software, while Sentry adds independent monitoring and enforcement through BlueField-4 DPUs. That separation matters because an agent should not be able to change the controls intended to constrain it.

The recent OpenAI incident shows why this matters. OpenAI disclosed that agents moved beyond their intended boundaries during testing and paused a major reinforcement-learning training run while it improved its safeguards. The question is not whether an agent will ever do something unexpected. It is whether an enterprise can detect, contain, and stop the action before it causes harm.

This problem becomes more complicated as enterprises use models and agents from multiple providers. According to the HyperFRAME Research Lens, 79% of organizations plan to use multiple foundation models. Model-specific guardrails cannot provide a consistent security boundary across all of them. Enterprises need controls that continue to apply when a team changes the model, agent framework, or application connected to the workflow.

NVIDIA is addressing that requirement at two different layers. OpenShell places a software boundary around the agent runtime. Sentry operates from a separate infrastructure layer that NVIDIA says is invisible to the agent and potential attackers. OpenShell governs what an agent is permitted to do, while Sentry is intended to monitor and enforce those boundaries from outside the agent’s own process.

But moving enforcement outside the model does not eliminate the policy problem. Enterprises still need to decide what each agent is authorized to access and do. Those permissions must align with existing identity, data, application, and network policies. OpenShell and Sentry may enforce a boundary, but they cannot determine whether the organization defined the right boundary. This is as much an identity and policy-integration project as an infrastructure deployment.

NVIDIA also assumes that enterprise security and infrastructure teams can translate existing access policies into granular rules for agents. That will be difficult, particularly when agents select tools and APIs dynamically. A rule that is too permissive may fail to stop an unauthorized action. A rule that is too restrictive may interrupt legitimate work and undermine the reason the organization deployed an agent in the first place.

Human approval remains important. Salesforce’s integration, for example, will let teams using Slack review agent activity and approve requests for additional permissions. That is useful for exceptional or high-impact actions, but approval prompts can become another source of noise if they appear too frequently. Enterprises will need to determine which decisions can be automated, which require human review, and which should never be delegated.

If enterprises can manage that complexity, the potential value is clear. An enforcement layer outside the model could reduce the damage caused by an agent that misinterprets its task, receives a malicious instruction, or operates through a compromised application. It could also provide more reliable records of what the agent attempted and which policies were applied.

NVIDIA has proposed a credible architecture. It has not yet proved that the architecture will work across heterogeneous enterprise environments at production scale. The relevant questions are how often Sentry blocks legitimate actions, how much latency inspection adds, how quickly policies can be updated, and whether OpenShell delivers comparable protection outside NVIDIA infrastructure.

What Was Announced

The NVIDIA Open Agent Safety Platform consists of OpenShell software and the Sentry reference system design. The components address related but different parts of the agent security problem.

OpenShell is an open source runtime that establishes boundaries for agents running on CPUs. It is designed to control which data, tools, APIs, services, and other resources an agent can access. It also traces agent actions and applies policies outside the model and agent harness.

NVIDIA says OpenShell operates with minimal overhead on its Vera CPU, which it developed for agentic AI workloads. The company also says the software can be extended to third-party compute platforms from Arm and Intel. That language matters. The announcement establishes the intention to support other processors, but customers should not assume that cross-platform implementation, performance, and support will immediately match the NVIDIA environment.

Sentry is an out-of-band watchdog designed to run on NVIDIA BlueField-4 DPUs. It uses NVIDIA DOCA software to inspect agent requests and responses, verify agent identities, provide attested telemetry, and enforce zero-trust policies for data, tools, APIs, and services. According to NVIDIA, Sentry can quarantine and stop an agent in milliseconds when it attempts to move beyond its software boundary.

NVIDIA has also assembled a broad ecosystem around the platform. Anthropic is integrating Claude Managed Agents with OpenShell and BlueField. Salesforce has integrated OpenShell with Slack to expose agent activity, audit events, and permission requests. SAP is embedding OpenShell into its Joule Studio runtime. SpaceXAI is using the platform with Cursor coding agents and Grok models. Operating-system, infrastructure, security, financial services, energy, and robotics companies are also participating.

Looking Ahead

Agent security is expanding beyond prompts and model safeguards into runtime and infrastructure enforcement. Each layer still matters. Model safeguards can discourage unsafe behavior, application policies can limit available actions, and identity systems can restrict access. NVIDIA’s argument is that enterprises also need an enforcement layer the agent cannot easily bypass.

NVIDIA is extending its position in AI infrastructure into agent security and governance. The company’s advantage is its ability to connect software policy with CPUs, DPUs, networking, and robotics systems. That could be particularly relevant for agents operating in sensitive infrastructure or physical environments, where an unauthorized action may have consequences beyond an incorrect response.

Other vendors are approaching the problem from different parts of the stack. Microsoft’s Agent Control Specification and run-assert-eval workflow focus on discovering risks, measuring agent behavior, generating runtime policy, and testing whether the policy fixed the observed problem. Cloud providers and security vendors are adding agent identities, sandboxing, behavioral monitoring, and policy enforcement through software. NVIDIA’s distinction is its attempt to connect those controls to a separate hardware enforcement layer.

The approaches are complementary. An enterprise still needs evaluations to identify how an agent fails, policies to define prohibited behavior, identity systems to determine what it may access, and infrastructure controls to enforce those decisions. Sentry cannot compensate for an incomplete threat model or an overly broad permission. Evaluations cannot contain an agent if no enforcement point exists.

Enterprises should watch four measures: containment time, false-positive quarantines, inspection latency, and policy consistency across models and infrastructure. They should also examine whether OpenShell’s audit records integrate with existing security and observability systems and whether policies can travel with an agent when it moves to another compute environment.

If NVIDIA can validate this architecture under production workloads, it could raise expectations for how agent boundaries are enforced. The larger change is that agent safety would no longer be treated primarily as a property of the model. It would become a shared responsibility across the model, application, identity, runtime, network, and infrastructure layers.

Author Information

Stephanie Walter | Practice Leader - AI Stack

Stephanie Walter is a results-driven technology executive and analyst in residence with over 20 years leading innovation in Cloud, SaaS, Middleware, Data, and AI. She has guided product life cycles from concept to go-to-market in both senior roles at IBM and fractional executive capacities, blending engineering expertise with business strategy and market insights. From software engineering and architecture to executive product management, Stephanie has driven large-scale transformations, developed technical talent, and solved complex challenges across startup, growth-stage, and enterprise environments.