François Chollet 谈大语言模型与推理模型的范式差异
基础大语言模型(2024年及更早)与现代大推理模型之间的关键区别不在于符号化工具使用。它是从转导范式(直觉推断查询的答案)到归纳范式的转变(直觉推断产生查询答案的程序/指令)。
它们被训练为归纳式的,并且它们执行测试时归纳,即测试时预测自然语言程序/推理链。这解锁了全新的能力——特别是流动智力。基础大语言模型至今仍拥有约0的流动智力。大推理模型拥有大量的流动智力。
大语言模型在ARC 1(2019年的基准测试)上的表现至今仍约为10-15%。将它们扩展约100,000倍,使其从0%提高到10%。与此同时,相同大小或更小的大推理模型在2025年已饱和ARC 1。
对照原文
The critical distinction between base LLMs (2024 and earlier) and modern LRMs is not symbolic tool use. It's the switch from a transductive paradigm (intuit the answer to the query) to an inductive paradigm (intuit the program/instructions that produce the answer to the query). They're trained to be inductive, and they perform test-time induction, i.e. test-time prediction of a NL program / reasoning chain. This unlocks entirely new capabilities -- in particular fluid intelligence. Base LLMs, to this day, have ~0 fluid intelligence. LRMs have substantial levels of fluid intelligence. The performance of LLMs on ARC 1 (a benchmark from 2019) remains ~10-15% today. Scaling them up by a factor ~100,000x got them from 0% to 10%. Meanwhile LRMs the same size or smaller saturated ARC 1 in 2025.