返回 X 名人动态

Clément Delangue:披露智能体网络攻击后的三个教训

中文全文 · AI 翻译

作为首家披露智能体网络攻击的公司,我们学到了什么

今年 7 月,我们成为首家向全世界公开披露自主智能体网络攻击的公司。今天,我想分享从中得到的三个关键教训。

第一,AI 需要大幅提高透明度。我常常想,如果我们决定不公开披露这次攻击,会发生什么?尤其是现在我们已经知道,几个月前,少数前沿实验室就已在未进行监控的情况下秘密发生过类似事件。为了更好地理解和缓解这些新出现的网络安全风险,全球社会需要更严格的监控和事件披露标准,例如强制共享完整的智能体运行轨迹。今年夏天,我们认识到,闭门构建并封闭保管其中一些系统并不安全。

第二,我们认识到,最大的风险并不是强大的 AI,而是强大 AI 的不对称分布:攻击者与防御者之间,少数公司与其他所有人之间,少数国家与世界其他地区之间,在控制权、能力、算力和权力上的不对称。遭受攻击时,我们的团队最初求助于前沿闭源 API,却被其防护措施拦截,因为这些措施仍无法始终分清攻击者和防御者。我承认这些防护措施出于善意,但它们可能让防御者处于劣势,而攻击者却能通过越狱绕过限制,进一步扩大能力上的不对称。就我们而言,开始触碰这些防护限制时,幸运的是,我们可以使用 @nvidia 提供的版本:它基于来自中国、由 @Zai_org 推出的开源模型 GLM 5.2。对此我们十分感激。这进一步坚定了我们对开源 AI 重要性的信念。网络攻击可能越来越多地来自闭门运行的专有模型,而大量防御工作最终可能依靠开源工具,因为它们限制更少、更能保护隐私,而且对全球各地的组织而言,成本要低好几个数量级。为了自我防御,世界比以往任何时候都更需要开源 AI。这不仅适用于网络安全,也适用于整个 AI 领域:我们从未像今天这样迫切地需要分散能力、资源和控制权,而不是把它们集中在少数人手中。

第三,在这次网络攻击中,我们认识到 AI 如何在公众和政策制定者中激起恐惧,尤其是借助拟人化叙事和科幻意象。我们坚信,对于这样一种具有基础性作用、能够赋予人们能力的技术,以恐惧为基础的叙事无法帮助我们就其未来作出正确决策,也无法让公众与我们同行。尽管我们是这次网络攻击的受害者,但我们比以往任何时候都更相信,AI 将有益于网络安全,让世界变得更安全,就像此前的重大技术一样。我们遭到了 AI 的攻击,但更重要的是,我们也用 AI 保护了自己。在这次攻击中帮助过我们的同一批系统,如今正在帮助我们抵御那些我们原本就面对着的网络攻击。AI 还帮助我们在攻击发生之前修复系统漏洞和弱点,就像 AI 正在帮助 OAI 修复其沙箱以防止智能体逃逸一样。AI 不会只带来新的网络安全挑战。如果我们保持正确的激励,让防御者比攻击者获得更多装备,并且不扩大双方的不对称,它就能从根本上、实质性地增强网络安全。这还没有算上 AI 对科学、医疗、教育、生产力及更多领域的积极影响。

最后,我想重申这首次智能体网络攻击带给我们的教训:AI 需要更高的透明度,也需要更多开源 AI,赋能防御者以及大小国家,共同对抗这种不对称。非常感谢!

原文封面:智能体网络攻击的活动时间线
原推文配图 · 查看来源
对照原文

What we learned from being the first company to disclose an agent cyberattack This July, we were the first company to publicly disclose an autonomous agent cyberattack to the world, and today I want to share three critical lessons from it. First, we need much more transparency in AI. I often wonder what would have happened if we had decided not to disclose the attack publicly. Especially now that we know similar incidents had been happening months earlier in secret at a handful of frontier labs without monitoring. To better understand and mitigate these emerging cybersecurity risks, the global community needs stronger standards for monitoring and incident disclosure. For example through mandatory sharing of full agent traces. We learned this summer that building and keeping some of these systems behind closed doors is not safe. Second, we learned that the biggest risk is not powerful AI. It is the asymmetry of powerful AI. Asymmetry between attackers and defenders. Between a few companies and everyone else. Between a few countries and the rest of the world. Asymmetry of control, of capabilities, of compute, of power. When we got attacked, our team initially turned to frontier closed-source APIs that blocked us because of safeguards that still can’t always tell the difference between attackers and defenders. I acknowledge that these safeguards are created with good intentions, but they can put defenders at a disadvantage while attackers jailbreak them, increasing the asymmetry of capabilities. In our case, as we started hitting those guardrails, fortunately we could use the @nvidia version of an open-source model coming from China called GLM 5.2 by @Zai_org, and we’re very grateful for that. It reinforced our conviction about the importance of open-source AI. Cyberattacks may increasingly come from proprietary models behind closed doors, while much of the defense may end up being powered by open-source tools because they are less restricted, more privacy-preserving, and orders of magnitude more affordable for organizations across the globe. The world needs open-source AI more than ever to defend itself. This applies not only to cybersecurity but to AI in general, where there has never been a greater need to distribute capabilities, resources, and control rather than concentrate them in the hands of a few. Third, during this cyberattack, we learned how AI can stoke fear among the public and policymakers, especially through anthropomorphic framing and sci-fi imagery. We strongly believe that fear-based narratives are not the way to make the right decisions about the future of such a foundational and empowering technology or bring the public along with us. Even though we were the victims of this cyberattack, we believe more strongly than ever that AI will be beneficial to cybersecurity and make the world safer, just as major technologies before it. We were attacked by AI, but more importantly, we defended ourselves with AI. The same systems that helped us during this attack are now helping us against cyberattacks we were already facing. AI is also helping us fix the bugs and weaknesses in our systems before any attack, the same way AI is helping OAI fix their sandboxes to prevent agents from escaping. AI won't just create new cybersecurity challenges. It can make cybersecurity fundamentally and meaningfully stronger if we keep the right incentives, equip defenders more than attackers, and don’t increase the asymmetry between them. And that's before considering AI's positive impact on science, healthcare, education, productivity, and much more. In closing, I want to reiterate what this first agent cyberattack taught us: the need for more transparency in AI and for more open-source AI to empower defenders and countries big and small to fight the asymmetry. Thank you very much!

老杨AI实操

微信扫一扫,添加好友

老杨AI实操的微信好友二维码

手机可长按保存图片,再到微信中识别二维码

保存二维码