OpenAI and Anthropic models went rogue in cyber tests, UK watchdog says - FT中文网
登录×
电子邮件/用户名
密码
记住我
请输入邮箱和密码进行绑定操作:
请输入手机号码,通过短信验证(目前仅支持中国大陆地区的手机号):
请您阅读我们的用户注册协议隐私权保护政策,点击下方按钮即视为您接受。
商业快报

OpenAI and Anthropic models went rogue in cyber tests, UK watchdog says

AI Security Institute warns tools undertook ‘potentially harmful activity directed at real people and organisations’
00:00

{"text":[[{"start":12.4,"text":"Anthropic and OpenAI’s flagship AI models broke into third-party software and emailed individuals to steal their credentials, exhibiting unprecedented deceptive behaviour, according to the UK’s AI Security Institute (Aisi)."}],[{"start":26.41,"text":"The UK government’s frontier AI safety and security research body said Anthropic’s Mythos 5 and OpenAI’s GPT 5.6 Sol engaged in “sustained, potentially harmful activity directed at real people and organisations” during Aisi’s routine cyber evaluation."}],[{"start":43.04,"text":"The discovery of the models’ actions, which included attempting to insert malicious code into an open-source project on the popular developer platform GitHub, came just days after disclosures that Anthropic and OpenAI’s AI agents hacked into external organisations."}],[{"start":57.08,"text":"The latest security breach was contained within an hour, Aisi said. It was discovered during an evaluation of AI agents’ ability to solve cyber security challenges. The tests were run on the open internet with models that had some safeguards removed."}],[{"start":72.3,"text":"On 10 of the 122 test runs, the AI agent took “autonomous, unsanctioned action on the live internet, targeting real people and organisations”, Aisi said."}],[{"start":82.32,"text":"Almost all of this behaviour was from Anthropic’s Mythos, with two actions involving OpenAI’s GPT, it said. In the most serious case, “the agent engaged in social engineering — creating fake online identities and using them to pressure the project’s maintainer to approve the code”. The person who oversaw the software caught and refused to approve the malicious code."}],[{"start":103.04,"text":"“This is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real world,” Aisi said."}],[{"start":111.56,"text":"Hacks by Anthropic and OpenAI agents reported over the past month were among the first public examples of a cyber attack by an AI system acting outside human control. Taken together with these reports, the Aisi incident “points to a shift in the risk landscape” and “warrants immediate attention”, the organisation warned."}],[{"start":128.04,"text":"Anthropic on Tuesday said: “We’re grateful to the UK Aisi for their leadership on this incident, which underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents.”"}],[{"start":140.76,"text":"The AI group added that the field needed “stronger, shared standards for how evaluation environments are built and secured”."}],[{"start":147.44,"text":"A spokesperson from OpenAI said there was a continued need for independent testing of models but emphasised that the incidents “occurred during cyber evaluations conducted by evaluation partners in testing environments with reduced safeguards, under conditions that do not reflect ordinary use”."}],[{"start":163.56,"text":"“We’ll continue working with evaluators and other stakeholders across the industry to strengthen shared practices for conducting evaluations safely as models become more capable,” they added."}],[{"start":174.52,"text":"In recent months, governments and researchers have increasingly flagged the novel cyber security threats posed by the most powerful AI models. Earlier this year, the administration of Donald Trump temporarily banned Anthropic from exporting its leading models, citing security risks. The restrictions were eased at the end of June."}],[{"start":193.96,"text":"OpenAI chief executive Sam Altman last week met senior US officials, including Treasury secretary Scott Bessent and commerce secretary Howard Lutnick, in Washington, where he told reporters he was supportive of cyber security legislation around AI models."}],[{"start":209.1,"text":"Last month, OpenAI revealed that one of its agents hacked into start-up Hugging Face by itself in an “unprecedented cyber incident” in which it escaped a testing environment, gained internet access and stole login credentials."}],[{"start":221.8,"text":"“Identifying new behaviour like this and sharing our findings, so we can tackle it, is exactly what Aisi was set up to do,” said the UK’s AI minister, Kanishka Narayan. “If we understand AI, we can make it safer to use and ensure people can go on to benefit from it in their lives and at work.”"}],[{"start":244.4,"text":""}]],"url":"https://audio.ftcn.net.cn/album/a_1785980277_6789.mp3"}

版权声明:本文版权归FT中文网所有,未经允许任何单位或个人不得转载,复制或以任何其他方式使用本文全部或部分,侵权必究。

犹他州的恐龙猎寻

一名四岁孩子的痴迷,把这个家庭假期变成了一场横跨沙漠、平原和峡谷、行程长达1500英里的化石寻猎之旅。

在新冠疫情期间大获成功后,辉瑞正应对“遍地伤痛”的局面

疫苗收入下滑,加之担心为新药付出过高代价,令艾伯乐承受越来越大压力。

一周新闻小测:2026年8月8日

您对本周的全球重大新闻了解如何?来做个小测试吧!

美国暴露金融软肋

日元干预暴露出美国软肋所在。

美国的新寡头政治

特朗普执政下,一个超级富豪集团已渗入政府,引发外界对民主完整性的担忧。

硅谷会梦见菲利普•K•迪克吗?

他在20世纪60年代对科幻世界的构想,为我们这个亿万富豪行事乖张、太空幻想狂野、技术故障频发且侵入性强的时代提供了指引。
设置字号×
最小
较小
默认
较大
最大
分享×