Can ChatGPT win a Fields Medal? | ChatGPT 能否获得菲尔兹奖? - FT中文网
登录×
电子邮件/用户名
密码
记住我
请输入邮箱和密码进行绑定操作:
请输入手机号码,通过短信验证(目前仅支持中国大陆地区的手机号):
请您阅读我们的用户注册协议隐私权保护政策,点击下方按钮即视为您接受。
FT英语电台

Can ChatGPT win a Fields Medal?
ChatGPT 能否获得菲尔兹奖?

New AI models could soon pose a threat to the world’s top mathematicians
新型人工智能模型在解决数学难题方面的能力正迅速提升。其在极具挑战性的全新问题上的表现,已对全球顶尖数学家构成威胁。
00:00

undefined

The writer is a science commentator

When Yang-Hui He, a fellow at the London Institute for Mathematical Sciences, received an invitation to an all-expenses paid weekend held in Berkeley, California, last month, it was a no-brainer. The trip would afford the Oxford university lecturer, an expert in algebraic geometry and string theory, insider access to a potentially historic moment for his discipline.

Plus, the brief sounded fun: working with other top mathematicians to find out if the most advanced AI models, when confronted with brand new problems, could rival or exceed the collaborative reasoning abilities of the best human minds. The answer? The machines did better than expected. “I’m not saying we felt existentially threatened but there was a general feeling of awe,” He told me. He also flew back $1,500 richer after dreaming up a problem that stumped the AI. 

Using AI to crack maths puzzles is not new. In early 2024, Google DeepMind unveiled technology that could hold its own in high-school student maths competitions. But interacting with the latest AI models last month felt more “like working with a very, very good graduate student”.

This moment could potentially change the profession. While the prospect of a machine securing a Fields Medal — widely regarded as mathematics’ equivalent of a Nobel Prize — still feels reassuringly distant, one can envision an unsettling future in which graduate maths programmes are pruned, university departments are shuttered, and the torch of Pythagoras and Euclid passed to a faceless silicon successor.

The weekend in mid-May was organised by Epoch AI, a US-based non-profit organisation that benchmarks AI capabilities. In an initiative set up last autumn called FrontierMath, Epoch paid professional mathematicians to submit novel problems along with their solutions, proofs and derivations, that could be used to challenge AI models.

These specially crafted conundrums, earning their creators up to $1,000 apiece and graded into three tiers of difficulty (including undergraduate and research level), were collected by Epoch via the secure messaging app Signal, so that they could not be inadvertently included in AI training data scraped from the internet. By April this year, Scientific American reported, an OpenAI model had confounded expectations by solving around a fifth of them.

And so it was time for tier 4 challenges: super-tricky problems that would take top academics weeks or months to solve collaboratively — and designed to resist AI guesswork or brute force number-crunching. Thirty academic experts, including He, met at Epoch’s Berkeley offices to brainstorm some new problems in person. Again, secrecy prevailed: lunches and dinners were brought in; attendees signed non-disclosure agreements and He recalled needing security cards to visit the toilets.

The full results of how the AI model performed on 50 tier 4 problems are yet to be disclosed. But He was struck by how much the tech has improved since 2022, when “ChatGPT couldn’t even find the tenth digit of seven divided by 13 . . . now it’s beginning to do something more intelligent.”

He explained how the AI, called o4-mini, was able to solve some of the problems in minutes, writing mathematical scripts and drawing on external specialist software. Most impressive, he said, were detailed literature searches, turning up obscure but critical papers and coding shortcuts. Another attendee, Ken Ono, a University of Virginia mathematician and freelance consultant for Epoch, called the results “frightening”. 

The project is not without controversy: in January, Epoch apologised for initially failing to disclose OpenAI’s financial backing of FrontierMath, leading to suspicions that the company’s AI models, including o4-mini, would have favoured access to some of the unseen maths problems used for benchmarking.

AI models cannot yet tackle the hardest maths challenges. Even so, one can imagine the next generation of machines thinning out the next generation of human mathematicians. That could shrink the pool from which future Fields medallists are drawn; there might be fewer hopefuls to attack famous unsolved problems like the Riemann Hypothesis, one of six carrying a $1mn bounty.

While the use of prime numbers in encryption shows the practical use of mathematics, there is something quite profound about living in a universe filled with dazzling concepts like zero, infinity and imaginary numbers. Perhaps fretting over whether the addition of AI might subtract from this human endeavour is not that irrational after all.

版权声明:本文版权归FT中文网所有,未经允许任何单位或个人不得转载,复制或以任何其他方式使用本文全部或部分,侵权必究。

谷歌为Anthropic打造的2000亿美元华尔街融资机器

私募信贷、芯片租赁和数据中心担保,支撑起AI支出的全新庞大模式。

“日元干预”等于“美国自保”

美国联手日本支撑日元,不只是出于盟友情谊,更是为了避免日本加息或美国国债遭抛售、导致美债收益率进一步走高。

“诅咒之岛”:科技游民与诈骗犯藏身的千亿美元奢华开发项目

警方的突击搜查再次打击了马来西亚陷入困境的中资“森林城市”项目的声誉。

问题不在因凡蒂诺

马杜罗:应该将国际足联的监管职能与商业活动分开,其治理应真正做到包容并具有代表性,监督必须真正独立。

俄罗斯扩大“影子”液化天然气船队,应对欧盟禁令

随着明年制裁进一步收紧,越来越多的“影子”船舶将帮助俄罗斯继续出口液化天然气。

阿斯利康与百时美施贵宝:大药企有时也不够大

当资产负债表规模扩大、能够押注潜在重磅药物时,规模才会带来优势。
设置字号×
最小
较小
默认
较大
最大
分享×