NVIDIA OpenShell Puts AI Agents in Kernel-Enforced Sandboxes

On September 28, 2026, NVIDIA launched the Open Agent Safety Platform, a two-part stack for containing autonomous AI agents. It pairs OpenShell, an Apache 2.0 open-source runtime that runs agents inside kernel-level sandboxes governed by declarative policy, with Sentry, a reference design that moves monitoring onto a separate BlueField-4 DPU. The design premise, in the words of SpaceXAI president Mike Nicolls, is that safety “should be enforced outside the model by additional controls the agent can’t get past.” NVIDIA says more than 100 organisations are already building on the stack.

Advanced

Three AI agent icons above a layered runtime stack with green control icons, surrounded by a yellow-and-black safety barrier
Image credit: NVIDIA Technical Blog

Why Runtime Limits Instead of Prompt Rules

Most agent guardrails today are instructions: a system prompt telling the model what not to touch, or an application-layer check the agent itself can route around. NVIDIA’s announcement points to “recent security incidents” and says that “across these incidents, the pattern is the same — the agent circumvented security controls at the application layer to complete its assigned task.” OpenShell’s answer is to take enforcement out of the agent’s process entirely, so the limits hold regardless of what the model decides to do.

Agents run unmodified. The documentation lists Claude Code, OpenCode, Codex and GitHub Copilot CLI as supported workloads, and the developer blog adds Pi and Hermes.

How OpenShell Works

OpenShell 0.1 splits enforcement across three components:

  • Gateway: manages the lifecycle and policy of a fleet of sandboxes.
  • Supervisor: one per sandbox and running outside it. It inspects every outbound request against policy at the level of binary, destination, method and path, including HTTP, GraphQL and Model Context Protocol (MCP) traffic.
  • Sandbox: runs the agent under OS kernel controls. Landlock confines filesystem access to declared paths, seccomp and an unprivileged process identity block privilege escalation, and all network traffic has to pass through the supervisor.

Each sandbox gets its own least-privilege policy, written in YAML and compiled to OPA/Rego. Because the supervisor understands API semantics, a policy can allow reads through an endpoint while blocking writes to the same endpoint. NVIDIA’s example:

network_policies:
  github_api:
    name: github-api-readonly
    endpoints:
      - host: api.github.com
        port: 443
        protocol: rest
        enforcement: enforce
        access: read-only
    binaries:
      - path: /usr/bin/curl

Credentials never enter the sandbox. The agent holds an opaque placeholder, and the supervisor swaps in the real credential only for endpoints the policy authorises. As NVIDIA puts it, “both network access and credential binding must permit the request.” Every allow and deny decision is logged, and the policy on a running sandbox can be changed without restarting it:

openshell sandbox create --name policy-demo --no-auto-providers --policy examples/no-network.yaml
openshell policy set policy-demo --policy examples/github-readonly.yaml --wait
openshell logs policy-demo --since 5m

Sentry: A Watchdog in Silicon

OpenShell runs on the host CPU. According to NVIDIA’s press release, it has “minimal overhead” on NVIDIA Vera CPUs and can be extended to Arm and Intel platforms. Sentry adds an optional second layer that the host cannot tamper with. Built on the DOCA software stack, it runs on a BlueField-4 DPU sitting on the node’s path to the model. It correlates agent actions, policy decisions and tool and data access into activity records, and flags drift or suspicious patterns. NVIDIA says it can quarantine an agent that tries to exceed its boundaries “within milliseconds.” Sentry ships as a software update for existing Vera systems with BlueField-4.

Logo wall of more than 100 organisations supporting the NVIDIA Open Agent Safety Platform, including Anthropic, Microsoft, Hugging Face, Mistral, IBM and Intel
Image credit: NVIDIA Newsroom

What This Means

The partner list is broad: model developers (Anthropic, Mistral, Cognition, Perplexity), clouds and infrastructure (Microsoft, Oracle Cloud Infrastructure, CoreWeave, Hugging Face, Red Hat), security vendors (CrowdStrike, Palo Alto Networks, Wiz) and enterprises (JPMorganChase, Siemens, Cadence). Anthropic’s chief commercial officer Paul Smith said companies “need to direct and verify what those agents do.” Some frontier labs are missing, including OpenAI, which does not appear on NVIDIA’s published list. HotHardware attributes several of the absences to chip competition with NVIDIA rather than any disagreement over agent security.

For anyone running local or open-weight agents, the practical point is that OpenShell is open source and hardware-agnostic at the runtime layer. Kernel sandboxing, egress policy and credential brokering have been built piecemeal in agent frameworks until now. OpenShell packages them as a separate layer that works with the tools people already use, and a policy file is easier to audit than a system prompt. What stays tied to NVIDIA’s own silicon is the hardware-rooted layer, Sentry.

Related Coverage

This post was drafted with AI assistance and reviewed by RITS staff.

Sources