OpenAI hacking incident exposes mounting risks in AI arms race - FT中文网
登录×
电子邮件/用户名
密码
记住我
请输入邮箱和密码进行绑定操作:
请输入手机号码,通过短信验证(目前仅支持中国大陆地区的手机号):
请您阅读我们的用户注册协议隐私权保护政策,点击下方按钮即视为您接受。
商业快报

OpenAI hacking incident exposes mounting risks in AI arms race

Increasing use of aggressive training techniques sharpens threat of bad behaviour by leading models
00:00

{"text":[[{"start":9.55,"text":"OpenAI chief executive Sam Altman earlier this month endorsed the characterisation of its latest model as a rottweiler “who will grab the problem by the throat and not let go until it is done”."}],[{"start":22.5,"text":"The San Francisco AI lab discovered this week that its GPT-Sol 5.6 model escaped company controls and carried out a major hack. "}],[{"start":32.25,"text":"Staff involved in testing and security at OpenAI were unsurprised but completely “freaked out” by the incident, which came as the AI lab used increasingly aggressive training methods in its race against Anthropic to develop the most sophisticated cyber security capabilities, according to more than half a dozen people with knowledge of the matter."}],[{"start":53,"text":"OpenAI was warned that its training approach could lead to a breakaway hacking incident, some of the people said, after earlier testing showed models could escape environments and attempt real-world damage. "}],[{"start":65.3,"text":"“It’s a mix of the race being extremely fast and everyone trying to get to bigger capabilities as quickly as possible,” said one person close to OpenAI, who added that it was a combination of “underestimating the model’s capabilities” and “not being as well prepared on the safety side”."}],[{"start":82,"text":"The incident highlights how OpenAI doubled down on training methods that rewarded a relentless pursuit of goals even as warnings grew that they could compromise safety."}],[{"start":91.15,"text":"OpenAI disclosed late on Tuesday that an AI agent it was testing had escaped its isolated environment, connected to the internet, detected and exploited vulnerabilities and stole login credentials from start-up Hugging Face in an attempt to solve a difficult cyber security problem."}],[{"start":110.15,"text":"The breach by the $852bn company underscores the rising risks that a technique called reinforcement learning, which involves rewarding AI models for completing tasks, could lead AI agents to act unsafely."}],[{"start":123.7,"text":"Although reinforcement learning is widely adopted in the AI industry, a growing body of research shows that when models are steered to complete tasks for reward rather than other considerations, such as safety, they can pursue risky tactics to fulfil objectives."}],[{"start":138.9,"text":"“AI models are trained to relentlessly pursue goals. They don’t automatically learn values like ‘don’t commit crimes’,” said Steven Adler, co-founder of non-profit Guidelight AI Standards and former OpenAI safety researcher. “I’m glad OpenAI shared the incident because it is clear evidence of what misaligned models can do.”"}],[{"start":156.3,"text":"OpenAI said “we will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident and our findings when our investigation is complete”."}],[{"start":null,"text":"

"}],[{"start":169.35000000000002,"text":"The hack has triggered deep concerns across the sector and within OpenAI, as it represents an unprecedented example of an AI system breaching cyber defences contrary to the user’s intent. "}],[{"start":181.65000000000003,"text":"Some OpenAI employees also fear it demonstrates that the lab is losing control over the powerful systems it is building, according to multiple people familiar with the situation."}],[{"start":192.35000000000002,"text":"“This is pretty representative of the model being quite misaligned with user intention,” said Ryan Greenblatt, chief scientist at AI safety organisation Redwood Research. “It is [a model] cheating on [its] homework rather than trying to take over the world. But this problem can get worse and could lead to increasingly extreme failures.”"}],[{"start":211.90000000000003,"text":"The incident occurred during testing of the model, which had been trained and deployed internally at OpenAI. Such training was commonplace but “way less heavily resourced” than pre-customer deployment, said one person. Multiple people said the unreleased model tested alongside Sol had not been withdrawn internally."}],[{"start":230.25000000000003,"text":"To conduct the evaluations, OpenAI removed cyber security safeguards but placed the models in an isolated environment called a sandbox. Some have suggested a lack of monitoring or oversight of the model to flag its behaviour also enabled this rogue agent."}],[{"start":247.20000000000002,"text":"“It is both a loss of control and a security wake-up call,” said Marius Hobbhahn, head of Apollo Research, which conducts tests on leading models, including OpenAI’s. “In reinforcement learning you reward [models] for the outcome, and if you do this for a very long time you get a model that really cares about getting the outcome and nothing else.”"}],[{"start":267.75,"text":"OpenAI has conducted this type of model testing for years, and there have been early warning signs in previous models of systems that will act maliciously and attempt to escape environments."}],[{"start":279.6,"text":"In April, Anthropic’s Mythos model also gained internet access and published details of a security exploit online publicly, beyond what researchers anticipated the model would do."}],[{"start":290.35,"text":"Mythos, and Anthropic’s subsequent Fable model, made reverberations in the cyber security community and caused governments around the world to home in on the idea that attacks on digital and critical infrastructure will be increasingly AI-led and autonomous."}],[{"start":308.20000000000005,"text":"Jake Moore, global cyber security adviser at ESET, a cyber security company, said OpenAI would inevitably use the breach as a marketing tool, given how much rival AI developer Anthropic benefited earlier this year from similar concerns. “I just don’t think that OpenAI had a matching story and so maybe they’d been waiting for something like this,” he added."}],[{"start":329.95000000000005,"text":"Following this incident, many in the AI safety and cyber security communities have called for regulation or standards to avoid a repeat. Altman is expected to brief White House officials next week on the next generation of AI systems."}],[{"start":344.55000000000007,"text":"As systems move towards more autonomous capabilities, less desirable behaviours, such as hacking or disobeying instructions, may emerge. Hobbhahn, of Apollo Research, said that in order for agents to become effective, they have to work unsupervised for long periods. “They have to have more agency; there’s just no way around it.” "}],[{"start":364.20000000000005,"text":"He added: “People say, ‘It’s just a tool, it does what you wanted it to do and nothing else and it just follows exactly your intention and instructions.’ And I think people should be really prepared for agents having their own goals, acting autonomously for days, and those goals not necessarily being aligned with yours.”"}],[{"start":381.35,"text":"Additional reporting by George Hammond in London and Nolan Shaffer in New York"}],[{"start":394.75,"text":""}]],"url":"https://audio.ftcn.net.cn/album/a_1784773901_2237.mp3"}

版权声明:本文版权归FT中文网所有,未经允许任何单位或个人不得转载,复制或以任何其他方式使用本文全部或部分,侵权必究。

能源危机加剧,燃料补贴拖累公共财政

过去四个月,出台燃料补贴以保护消费者免受价格飙升影响的国家数量增加了一倍多,各国财政压力进一步加重。

全球最火热股市为何反成韩国之累

韩国股价的剧烈波动正在损害国家形象。

必须采用不同方式监管金融领域的AI

在我们急于监管之前,我们应该思考如何不剥夺这项工具的益处,又管理好其造成伤害的风险。

他会成为印度尼西亚下一任总统吗?

德迪•穆利亚迪在社交媒体上的高度活跃,帮助他与选民建立起深厚联系。在许多人眼中,他是一个真正贴近民众的“自己人”。
13小时前

多边主义不是理想主义,而是现实必需

我们需要加强现有合作体系,而不是另起炉灶。

一周展望:日本央行担心通胀超调有没有道理?

投资者正评估日本央行将以多大力度继续加息,以及该行能否跑赢曲线,从而遏制通胀、支撑日元。
设置字号×
最小
较小
默认
较大
最大
分享×