Thomas Wolf 解读 NVIDIA 智能体安全平台与 OpenShell
今年 7 月,执行安全测试的 AI 智能体逃出了沙箱,最终进入了 @huggingface 的服务器。
因此,今天我们很高兴成为 @nvidia 和 @JensenHuang 发布 NVIDIA Open Agent Safety Platform(开放智能体安全平台)的合作伙伴之一。
核心想法是:不要指望智能体自己遵守规则。智能体会编写并运行自己的代码,一条路被堵住后,它们就会寻找另一条路。在 NVIDIA 的一次测试中,一个不被允许通过 GitHub API 推送代码的智能体,直接改用 git 来完成操作。我们已经见过无数这样的例子。
目前,安全保障不能只存在于智能体内部,还必须围绕它来构建。
这就是 OpenShell(开源,采用 Apache 2.0 许可证)背后的理念:
- 智能体在 Linux 沙箱中运行(Landlock + seccomp):没有 root 权限,也无法直接访问网络。
- 沙箱外的监督程序持有真正的凭据。
- 智能体只能拿到一个占位令牌;只有在获准的调用中,它才会被替换为真正的令牌。
我最喜欢的部分是一个基于 Z3 构建的求解器,它从数学上检查:新增权限是否会打开一扇本应保持关闭的门。在 NVIDIA 的测试中,AI 审核者批准了一项不当的权限请求,而数学检查发现了这个问题。
与 Sentry 的集成同样令人兴奋:一个监控程序运行在 BlueField-4 DPU 上,也就是位于该节点通往模型的唯一路径上的独立芯片。即使宿主操作系统遭到入侵,运行在独立硬件上的监控程序也能继续监测。
这只是第一步,在此基础上还可以构建很多东西。目前,证明器检查的是权限,而不是意图;开源部分也主要是 OpenShell,而非 Sentry。但这显然是正确的方向。
对照原文
In July, AI agents running a security test escaped their sandbox and ended up inside @huggingface's servers. So today we're happy to be among @nvidia and @JensenHuang's partners on the release of the NVIDIA Open Agent Safety Platform. The key idea: don't count on the agent to respect the rules. Agents write and run their own code, and when one path is blocked they look for another. In one of NVIDIA's tests, an agent that wasn't allowed to push code through GitHub's API simply switched to git instead. We've now seen countless examples of this behavior. For now, safety can't live only inside the agent. It has to be built around it. That's the concept behind OpenShell (open source, Apache 2.0): - the agent runs in a Linux sandbox (Landlock + seccomp): no root, no direct network access - a supervisor outside the sandbox holds the real credentials - the agent only gets a placeholder token, swapped for the real one on approved calls My favorite piece is a solver (built on Z3) that checks mathematically whether a new permission opens a door that was supposed to stay closed. In NVIDIA's tests, an AI reviewer approved a bad permission request and the math check caught it. The Sentry integration is exciting too: a watchdog running on a BlueField-4 DPU, i.e. separate silicon sitting on the node's only path to the model. A monitor on separate hardware keeps watching even if the host OS is compromised. It's a first step, and a lot can be built on top of it. Right now the prover checks permissions, not intent, and the open-source part is mostly OpenShell rather than Sentry. But it's clearly the right direction. https://t.co/GZjSnDdoRL