François Chollet 追问数学代码训练能否提升通用能力
如果锯齿状的前沿主要就是数学 + 代码(你可以用 RLVR 将其任意推进),而其他一切都开始趋于平稳,因为它们仍然受限于人类生成的数据瓶颈,那会怎样?
模型在不可验证领域的性能一直稳步提升,尽管速度远慢于数学和代码。但这种稳步提升是更高 G(本身由 RLVR 驱动)的副作用,还是仅仅取决于注入训练的新人类数据量(这仍然在巨大规模上持续发生)?
许多事情都取决于这个问题的答案
改写:G 是从数学 + 代码 RLVR 中“涌现”出来的吗?还是你只会获得更高的数学 + 代码技能,而没有其他?
也许 G 是一种技能,并且它与数学 + 代码问题解决是同构的?
对照原文
What if the jagged frontier is mainly math + code (which you can push arbitrarily far with RLVR), and everything else starts to plateau because it is still bottlenecked by human generated data? Model performance in non-verifiable areas has kept improving steadily, albeit much slower than for math and code. But is that steady improvement a side effect of a higher G (itself driven by RLVR), or only a function of the amount of new human data getting injected into training (which is still continually happening on a massive scale)? A lot of things depend on the answer to this question
Rephrased: does G "emerge" from math + code RLVR? Or do you only get higher math + code skill and nothing else? Maybe G is a skill and it is isomorphic to math + code problem solving?