In this article
NVIDIA released the Open Agent Safety Platform on September 28 with a useful architectural split. OpenShell 0.1.0 mediates an agent’s access from outside its workload. An optional Sentry layer runs on a BlueField-4 data processing unit, outside the host software that the agent may influence. The design does not ask one model to supervise another from inside the same process boundary. It separates the workload, its runtime guard, and an infrastructure backstop. (NVIDIA platform architecture)
That is a meaningful change in failure domains, not evidence that agent containment is solved. OpenShell can enforce only the policy it receives. Sentry can interrupt activity it can observe, but it cannot make an incomplete authorization model complete or reverse an external side effect that already happened. NVIDIA has published the architecture and code for OpenShell, while independent production measurements of false positives, missed violations, and recovery behavior remain absent.
OpenShell makes the runtime boundary explicit
OpenShell is an Apache 2.0 runtime that can wrap existing coding and agent frameworks without moving enforcement into the agent. Its repository describes three trust roles. The supervisor stays outside the workload and owns admitted policy, gateway credentials, L7 proxying, and relays. The sandbox runs beside the agent as the same non-root identity and installs seccomp and Landlock controls. The agent child inherits those restrictions with zero Linux capabilities and no_new_privs. (OpenShell 0.1.0) (sandbox architecture)
That decomposition gives platform teams a concrete place to answer questions that prompts cannot settle:
- Which binaries may reach which hosts, ports, methods, and paths?
- Which files and processes may the workload read, modify, or create?
- Which credential may be attached to a request without entering the workload?
- Who may widen a live policy, and which revision authorized a later effect?
Adding a binary or endpoint to an existing network rule must restate the affected scope, and the gateway rejects unrequested binary-to-endpoint grants. Filesystem and process changes require recreating the sandbox. Authority expansion is therefore a control-plane event rather than an agent-side edit. (OpenShell policy workflow)
Production adoption still needs adversarial boundary tests. Verify that no network path bypasses the supervisor, credentials cannot be recovered from the workload, and mediation loss fails in the intended direction. OpenShell says an interrupted pending open fails closed and a disconnected supervisor gets a bounded reconnection window before termination. Operators still need to test those semantics in their own container, VM, Kubernetes, proxy, and DNS configurations. (sandbox architecture)
Sentry changes the failure domain, not the policy problem
NVIDIA describes Sentry as a separate hardware watchdog that correlates agent interactions, policy decisions, and tool or data access. In the Vera Rubin POD reference design, BlueField-4 is on the node’s only path to the model. That placement gives the DPU an observation point outside both the agent and host, plus a way to interrupt the next model call even if host resources are no longer trusted. NVIDIA says the layer can monitor and enforce policy at line speed. (NVIDIA platform architecture)
The scope needs careful reading. The strongest topology claim applies to that reference configuration, and Sentry is optional. NVIDIA says the broader platform is compatible with other hardware, but compatibility is not the same as reproducing the same out-of-band trust boundary. Reuters independently confirmed the two-component release and reported NVIDIA’s claim that Sentry can cut off an agent attempting to escape its CPU container. That is a vendor assertion about a new system, not a measured containment result. The unresolved operating trade-off remains: organizations define their own rules, and neither NVIDIA nor Reuters published false-positive or missed-detection rates. (Reuters)
Treat the watchdog as a circuit breaker, not a transaction rollback system. Cutting the model path can stop an agent from planning its next step. It does not undo a message already sent, a row already deleted, or a payment already submitted. High-consequence tools therefore still need idempotency keys, staged effects, approval thresholds, and compensating actions at the application boundary.
Use an acceptance matrix, not a demo. For every effect class, record the enforcing layer, policy version, failure mode, emitted evidence, and recovery owner. Inject denied network calls, credential requests, supervisor loss, policy races, and already-committed effects. The evidence must identify which layer stopped the action and what remains to repair.
SAP adds the business decision above containment
SAP’s implementation note exposes a distinction that the platform diagram does not settle. Joule Studio applies role-based policy, business authorization, and process context to decide whether an action should execute. OpenShell governs how it executes, what it can access, and where inference goes. FedRAMP, FIPS, and regulated-industry enablement remain on a roadmap. (SAP engineering account)
That suggests three different control products:
- Business authorization decides whether this user, agent, and process may attempt an action.
- Runtime mediation constrains the files, services, credentials, tools, and processes available during execution.
- Independent containment observes a separate failure domain and can stop continuation when the host or workload is suspect.
Collapsing those layers into one “agent policy” hides ownership. Security can validate runtime isolation while a product team still owns an incorrect refund limit. A hardware watchdog can remain healthy while an IAM mapping grants the wrong tenant. Audit records should therefore join business decision ID, runtime policy revision, agent identity, model route, and effect receipt rather than treating a quarantine event as the whole story.
Claude Sonnet 5.5 adds a separate migration gate
Anthropic also launched Claude Sonnet 5.5 on September 28. Its release notes identify five ways code written for Sonnet 5 can break: the request for disabling up-front thinking changes, forced tool_choice values any and tool return HTTP 400, thinking blocks bind to the model and conversation, the earlier computer-use tool is rejected on the Claude API and Google Cloud, and older models cannot serve as advisors. Sonnet 5.5 thinking blocks are also account-bound and may be dropped when replayed by another unlinked account. (Claude Platform release notes)
Keep the model and API migrations separate. First canary Sonnet 5.5 with the current tool schema and account boundary. Test forced tool calls, preserved thinking, cross-account replay, computer use, and advisor selection as separate fixtures. A successful text response does not prove that a dropped thinking block preserved workflow state.
Stay Sharp: test a reference monitor, not just a sandbox
NIST defines a reference monitor through three properties: the validation mechanism is always invoked, tamperproof, and small enough to analyze and test. (NIST reference monitor glossary) Those properties turn “outside the agent” into testable questions.
For complete mediation, enumerate every route to files, sockets, credentials, tools, and inference, then prove each crosses a named enforcement point. Test tamper resistance from the workload identity, not an administrator shell. Keep the policy language and trusted computing base narrow enough to review the effective grant.
OpenShell and Sentry are shaped around these ideas, but architecture diagrams do not establish that a deployed configuration satisfies them. Sidecars, debug ports, alternate DNS paths, host-mounted sockets, or emergency credentials can quietly create a second route. The senior engineering task is to find that route before the agent does.
What to watch
- Whether NVIDIA publishes the exact property its policy prover establishes, plus reproducible counterexamples for rejected policies.
- Measured Sentry detection coverage, false-positive rates, quarantine latency, and behavior when the DPU or model path is unavailable.
- Whether non-NVIDIA deployments can preserve an equally independent enforcement boundary rather than only API compatibility.
- Sonnet 5.5 incident reports involving account-bound thinking, forced tool selection, or the computer-use tool transition.