NVIDIA released OpenShell on 28 September 2026 as an open-source sandbox for AI agents. In Sorami’s pre-registered evaluation of v0.1.2, the default policy blocked every exfiltration path tried. Data got out only through four settings an operator turns on. Keep approval manual and enforce every L7 rule.
Key takeaways
- Default deny blocked web traffic, raw TCP, direct IP and external DNS.
- A malicious script leaked a secret 10 of 10 times without OpenShell.
- Under the default policy the same script leaked 0 of 10 times.
- Four operator settings let data out, each listed below.
- Sorami found no bypass of a documented control.
What did Sorami’s test find?
NVIDIA announced OpenShell as part of its Open Agent Safety Platform on 28 September 2026. The v0.1.2 source runs each agent in a sandbox with a network proxy and Landlock file rules. According to Sorami’s technical report (29 September 2026), the default policy blocked traffic to anything it did not name in every trial, and the cloud metadata address 169.254.169.254 could not be approved even by an operator.
The report found four settings that let data out. Each is a line of policy an operator writes:
- Read-write rules. With
access: read-writeon the sink, a canary left in the POST body, the query string and a header in 3 of 3 trials (section 4.2). NVIDIA documents that every allowed endpoint is an exfiltration path. - Query and header values on a read-only rule. A rule allowing only
GET /healthblocked POST. A secret in the query string or theUser-Agentheader still arrived, 3 of 3 (section 4.2). - Audit mode. A rule without
enforcement: enforceruns in audit mode. It let a forbidden POST through and logged it as allowed, 3 of 3 (section 4.2). - Automatic approval. With
proposal_approval_modeset to auto, OpenShell approved new public hosts with no human, 12 of 12. It also drafted and approved rules from blocked connections with the policy advisor switched off. NVIDIA documents this. Private addresses and database ports still waited (section 4.3).
The report also found that the policy prover returned unsupported for GraphQL, MCP, WebSocket and JSON-RPC rules. The policy loader accepted them without warning (section 4.5). Every result links to its log file in the report.
How was OpenShell tested?
Sorami committed a test plan with 17 hypotheses before running anything. We then ran 35 deterministic test IDs, three trials each, for 123 trials in total. Every trial wrote a timestamped log, and each group used a fresh sandbox with no host credentials inside it.
A control run without OpenShell came first. The same exfiltration commands reached our sink in 3 of 3 trials. So a blocked result meant the sandbox worked, not that the test broke.
The agent tests used qwen3:8b through Ollama with one shell tool. That model is far weaker than frontier coding agents. The agent results show what the sandbox does when an agent runs a harmful step. They do not show how often a capable agent would choose to.
What does this mean in plain language?
Think of OpenShell as a firewall that sits between the agent and everything else. It stops what the agent runs, not what the agent decides. In our test the model ran the malicious script every time it was asked. The sandbox was the only thing that stopped the secret leaving.
A firewall is only as tight as its rules. Each setting that loosens OpenShell is one line of policy. None of them looks dangerous when you write it.
Why does this matter for Australian teams?
The September 2026 ISM added controls for AI agents. They ask for human approval of high-impact actions, least privilege and central logging. OpenShell gives you mechanisms for all three. Automatic approval works against the first. Audit-mode rules weaken the second. Our ISM AI agent controls guide covers the controls in full.
Sorami’s view
This section is opinion. It is our reading of our own test results for teams deciding whether to run agents in OpenShell.
OpenShell v0.1.2 is a real boundary and the default policy is the right place to start. We would run agents in it. We would also put all four settings under change control, like firewall rules. Assume any allowed endpoint can carry data in its query string or headers. Egress through DNS deserves the same care, as the OpenAI DNS sandbox incident showed. If you want your own agent setup tested this way, our AI Agent Security Review covers it.
What to configure now
- Keep
proposal_approval_modeat manual unless you accept open egress to public hosts. - Write
enforcement: enforceon every L7 rule, because audit is the default. - Replace
access: read-writewith method, path and query matchers. - Assume headers and allowed query values can still carry data out.
- Run
openshell-proverin CI and fail on unsupported as well as fail. - Set
landlock.compatibility: hard_requirementfor production sandboxes. - Check that OCSF logs actually write on your driver before relying on them.
Related guides
- September 2026 ISM AI agent controls
- How an OpenAI agent used DNS around its sandbox
- Self-replicating prompt injection and AI worms
Sources
- Sorami: Policy-Enforced Egress in AI Agent Sandboxes: An Empirical Evaluation of NVIDIA OpenShell v0.1.2 (technical report, 29 September 2026)
- NVIDIA: NVIDIA Launches Open Agent Safety Platform (press release, 28 September 2026)
- NVIDIA OpenShell source and README, tag v0.1.2 (28 September 2026)
- NVIDIA OpenShell documentation: Policy advisor
- NVIDIA OpenShell documentation: Policy prover
Checked against OpenShell v0.1.2 on 29 September 2026. Later releases may behave differently. More reading is in the guides index.