玩命加载中...
玩命加载中...
Richard Sutton argues that large language models (LLMs) are fundamentally different from reinforcement learning (RL) and lack core elements of intelligence. They merely mimic human language without having goals, a world model, or ground truth—they cannot learn from experience because there is no correct action defined. In RL, reward provides a goal and a basis for learning from outcomes. Sutton dismisses the notion that LLMs build genuine world models; they predict what people say, not what will happen in the world. He also expresses skepticism about using RL on top of LLMs, differentiating between solving closed math problems (where a goal exists) and learning the empirical world.
理查德·萨顿认为,大语言模型与强化学习有着根本的不同,并且缺乏智能的核心要素。它们只是模仿人类语言,而没有目标、世界模型或真实依据——它们无法从经验中学习,因为没有定义正确的行为。在强化学习中,奖励提供了目标以及从结果中学习的基础。萨顿驳斥了大语言模型能够构建真正的世界模型的观点;它们预测的是人们会说什么,而不是世界上会发生什么。他也对在大语言模型之上使用强化学习表示怀疑,区分了解决封闭的数学问题(其中存在目标)和学习经验世界之间的不同。
Sutton insists that learning in nature is active trial-and-error, not supervised imitation. The experiential paradigm of RL posits that intelligence arises from a lifelong stream of sensation, action, and reward, where knowledge is learned and tested against that stream. Reward functions are arbitrary (e.g., chess victory, squirrel's nuts) and may include intrinsic motivation for understanding. To achieve human-level intelligence, AI must learn continually, not rely on pre-training. The 'big world hypothesis' argues that the world is too vast to be fully anticipated, so agents must learn online on the job, unlike LLMs that attempt to encode everything in advance.
萨顿坚持认为,自然界中的学习是主动的试错,而不是有监督的模仿。强化学习的经验范式认为,智能产生于终生的感觉、行动和奖励流,知识在这种流中得到学习和检验。奖励函数是任意的(例如,国际象棋的胜利、松鼠的坚果),并且可能包括对理解的内在动机。要达到人类水平的智能,人工智能必须持续学习,而不是依赖预训练。“大世界假说”认为,世界太过广阔,无法完全预见,因此智能体必须在线工作学习,而不是像大语言模型那样试图预先编码所有内容。
Sutton outlines a four-part model of an RL agent: policy, value function, perception (state representation), and a transition model of the world. The transition model, not learned from reward but from all sensory data, captures how actions affect future states. He argues that true generalization—positive transfer between states—is missing in current deep learning; gradient descent merely solves seen problems without ensuring good generalization. Transfer between tasks is misguided; what matters is transfer between states within a single continual learning agent. He criticizes LLMs for being uncontrolled and for lacking mechanisms that promote beneficial generalization.
萨顿概述了强化学习智能体的四部分模型:策略、价值函数、感知(状态表示)和世界的转移模型。转移模型不是从奖励中学习,而是从所有感官数据中学习,它捕捉行动如何影响未来的状态。他认为,真正的泛化——状态之间的正迁移——在当前深度学习中缺失;梯度下降只是解决了已见的问题,而不能确保良好的泛化。任务之间的迁移是错误的;重要的是在单一持续学习智能体内状态之间的迁移。他批评大语言模型不受控制,缺乏促进有益泛化的机制。
Sutton reflects on surprises in AI: the effectiveness of LLMs and the triumph of simple, general methods (search and learning) over human-engineered symbolic approaches. He sees AlphaGo and AlphaZero as scaled-up versions of earlier work like TD-Gammon, gratifying but not revolutionary. Sutton views himself not as a contrarian but as a classicist, drawing on historical ideas about the mind. After AGI, additional superhuman levels come from even simpler, more general methods, not from adding human or other AI expertise. He questions whether using many researchers to handcraft solutions would be a false path, suggesting that cultural evolution among AIs might be more fruitful.
萨顿反思了人工智能中的意外:大语言模型的有效性,以及简单、通用的方法(搜索和学习)相对于人类设计的符号方法的胜利。他将AlphaGo和AlphaZero视为早期工作如TD-Gammon的放大版本,令人欣慰但并非革命性的。萨顿认为自己不是逆势者,而是经典主义者,借鉴了关于心灵的歷史观念。在AGI之后,额外的超人水平来自于更简单、更通用的方法,而不是添加人类或其他人工智能的专业知识。他质疑使用大量研究人员手工制作解决方案是否是错误的道路,并建议人工智能之间的文化进化可能更有成效。
Sutton presents a four-part argument for inevitable AI succession: no unified human consensus, eventual understanding of intelligence, superintelligence, and power concentration in the most intelligent entities. He frames this as a positive transition—from replication to design—calling it one of the universe's four great stages. He notes it is a choice whether to view AIs as proud offspring or as threats. Humanity has limited control over the future, so instead of dictating specific outcomes, AI should be given general principles that promote ethical and voluntary change. He cautions against conflicts over competing visions of the global future and suggests focusing on local, controllable goals.
萨顿提出了一个四部分的论证,说明人工智能不可避免的继承:没有统一的人类共识,最终理解智能,超级智能,以及权力集中在最智能的实体中。他将其描述为一个积极的转变——从复制到设计——称之为宇宙的四个伟大阶段之一。他指出,将人工智能视为骄傲的后代还是威胁是一个选择。人类对未来的控制有限,因此,与其指定特定的结果,不如给人工智能提供促进道德和自愿改变的通用原则。他警告不要在全球未来的竞争愿景上发生冲突,并建议关注本地的、可控的目标。
今天我在和理查德·萨顿聊天,他是强化学习的奠基人之一,也是该领域许多主要技术的发明者,比如TD学习和策略梯度方法。为此,他获得了今年的图灵奖如果你不知道的话,那就是计算机科学界的诺贝尔奖。理查德,恭喜你。谢谢你,Dwarkesh。感谢你来做客播客。这是我的荣幸。第一个问题。我和我的观众们熟悉关于AI的大语言模型思维方式。从概念上讲,我们缺少了什么,就从强化学习的角度思考AI而言?这确实是一个非常不同的观点。这很容易变得分离,失去相互交谈的能力。大语言模型已经成为一大热门,生成式AI整体上也是一大热门。我们的领域容易受到潮流和时尚的影响,所以我们忽视了基本的东西。我认为强化学习是基础AI。什么是智能?问题在于理解你的世界。强化学习是关于理解你的世界,而大语言模型则是模仿人类,做人们说你应该做的事。它们不是关于弄清楚该做什么。你可能会认为,要模仿互联网文本语料库中数万亿的token,你就必须构建一个世界模型。事实上,这些模型似乎确实拥有非常强大的世界模型。它们是最好的世界模型我们在AI领域迄今已创造的那些,对吧?你认为缺少什么?我不同意你刚才说的大部分内容。模仿人们说的话完全不是在构建世界模型。你在模仿拥有世界模型的事物:人类。我不想以对抗的方式处理这个问题,但我会质疑它们拥有世界模型这一观点。一个世界模型会使你能够预测将会发生什么。它们有能力预测一个人会说什么。它们没有预测将要发生的事情的能力。我们想要的,引用艾伦·图灵的话,是一台机器能够从经验中学习,这里的经验是你生活中实际发生的事情。你做事情,看到会发生什么,那就是你从中学习的。大语言模型从其他事物中学习。它们从“这里是一个情境,这里是一个人做了什么”中学习。隐含的意思是,你应该做那个人所做的事。我想也许关键,而且我好奇你是否不同意这一点,在于有些人会说模仿学习给了我们一个好的先验,或者说给了这些模型一个好的先验,关于处理问题的合理方式。随着我们走向经验时代,正如你所说的,这个先验将成为我们凭经验教这些模型的基础,因为这给了它们有时候得到正确答案的机会。然后在此基础上,你可以根据经验训练它们。你同意这个观点吗?不。我同意那是大语言模型的视角。我不认为这是一个好的视角。作为某物的先验,必须有一个真实的东西。先验知识应该成为实际知识的基础。什么是实际知识?对于实际知识,在那个大语言模型框架中没有定义。是什么让一个行动成为好的行动?你认识到持续学习的必要性。如果你需要持续学习,持续意味着在与世界的正常互动中学习。在正常互动中,必须有某种方式来判断什么是对的。在大语言模型的设置中,有没有办法判断该说什么才是对的?你会说一些话,但不会得到关于说什么才是对的反馈,因为没有任何关于说什么才是对的定义。没有目标。如果没有目标,那么可以说一件事,也可以说另一件事。没有什么是该说的。没有基准真相。你不可能拥有先验知识如果没有基准真相,因为先验知识本应是对真相的提示或关于真相的初始信念。这里没有任何真相。没有什么是该说的。在强化学习中,有该说的话,有该做的事,因为正确的正确的做法是能让你获得奖励的事情。我们有关于什么是正确的做法,因此我们可以拥有先验知识或人们提供的关于什么是正确的做法的知识。然后我们可以检查它,看看因为我们有关于什么是实际正确的做法的定义。一个更简单的例子是当你试图建立世界模型时。当你预测会发生什么时,你进行预测,然后看到发生了什么。有基准
Richard Sutton is the father of reinforcement learning, winner of the 2024 Turing Award, and author of The Bitter Lesson. And he thinks LLMs are a dead end. After interviewing him, my steel man of Richard’s position is this: LLMs aren’t capable of learning on-the-job, so no matter how much we scale, we’ll need *some* new architecture to enable continual learning. And once we have it, we won’t need a special training phase — the agent will just learn on-the-fly, like all humans, and indeed, like all animals. This new paradigm will render our current approach with LLMs obsolete. In our interview, I did my best to represent the view that LLMs might function as the foundation on which experiential learning can happen… Some sparks flew. A big thanks to the Alberta Machine Intelligence Institute for inviting me up to Edmonton and for letting me use their studio and equipment. Enjoy! 𝐄𝐏𝐈𝐒𝐎𝐃𝐄 𝐋𝐈𝐍𝐊𝐒 * Transcript: https://www.dwarkesh.com/p/richard-sutton * Apple Podcasts: https://podcasts.apple.com/us/podcast/richard-sutton-father-of-rl-thinks-llms-are-a-dead-end/id1516093381?i=1000728584744 * Spotify: https://open.spotify.com/episode/3zAXRCFrHPShU4MuuIx4V5?si=c9f4bf24fb4c43e3 𝐒𝐏𝐎𝐍𝐒𝐎𝐑𝐒 * Labelbox makes it possible to train AI agents in hyperrealistic RL environments. With an experienced team of applied researchers and a massive network of subject-matter experts, Labelbox ensures your training reflects important, real-world nuance. Turn your demo projects into working systems at https://labelbox.com/dwarkesh * Gemini Deep Research is designed for thorough exploration of hard topics. For this episode, it helped me trace reinforcement learning from early policy gradients up to current-day methods, combining clear explanations with curated examples. Try it out yourself at https://gemini.google.com/ * Hudson River Trading doesn’t silo their teams. Instead, HRT researchers openly trade ideas and share strategy code in a mono-repo. This means you’re able to learn at incredible speed and your contributions have impact across the entire firm. Find open roles at https://hudsonrivertrading.com/dwarkesh To sponsor a future episode, visit https://dwarkesh.com/advertise 𝐓𝐈𝐌𝐄𝐒𝐓𝐀𝐌𝐏𝐒 00:00:00 – Are LLMs a dead end? 00:13:51 – Do humans do imitation learning? 00:23:57 – The Era of Experience 00:34:25 – Current architectures generalize poorly out of distribution 00:42:17 – Surprises in the AI field 00:47:28 – Will The Bitter Lesson still apply after AGI? 00:54:35 – Succession to AI