Jerry Liu:评估与强化学习环境中的七个难题
昨天,我与来自 @SnorkelAI 的 @vincentsunnchen 举办了一次有趣的晚餐对话,主题是评估和强化学习环境。
“数据和强化学习环境”公司(如 Snorkel)在过去几年中经历了巨大的增长。评估领域的兴趣呈爆炸式增长。与此同时,模型在每次新发布时都在基准测试中大放异彩。我们讨论了每个人都在定义的评估、尚未解决的评估、哪些事情交给前沿模型、哪些依靠自己掌握的智能,以及更多内容:
-
构建强化学习环境的一个重大挑战是“公平性”——当模型在给定环境中失败时,你能否将其归因于输入、测试框架还是奖励模型?
-
构建适当的奖励很困难。有些任务不易量化。你还想阻止钻奖励机制的空子。同时,你也不想对中间奖励规定得过于死板。
-
长时程评估仍然极其困难,一些业务流程在最终结果出现前可能需要数周或数月。
-
大多数受监管的行业仍然需要人类参与循环来保证近 100% 的准确性,“80%”的准确性是不够的。
-
随着模型变得更智能,将会出现两极分化的局面:精品数据供应商(如任何中小企业)和大规模数据提供商。
-
模型仍然表现出“锯齿状智能”,在长尾边缘案例上仍然会失败。
-
对于特定任务,始终有机会收集独特数据,并以更低的成本和更高的准确性对模型进行后训练。
这是我们创始人晚餐系列的 #003 期。我们下次应该讨论什么话题?请在下面告诉我们你的想法!

对照原文
Yesterday I hosted a fun dinner conversation with @vincentsunnchen from @SnorkelAI on evals and RL environments. The "data and RL env" companies (like Snorkel) have seen massive growth in the past few years. There's been an explosion of interest in evals. At the same time, models are ripping through benchmarks with each new release. We talked about the evals everyone is defining, what evals are still left unsolved, what’s left up to frontier models vs. intelligence that you own, and more: - A big challenge for building RL environments is “fairness” - when the model fails on a given environment, can you attribute it to the input, harness, or reward model? - Building proper rewards is hard. Some tasks are not easily quantifiable. You also want to discourage reward hacking. At the same time, you don’t want to be too prescriptive with intermediate rewards. - Long horizon evals are still extremely hard, some business processes can take up to weeks or months before the final outcome - Most regulated industries still need human in the loop to guarantee ~100% accuracy, “80%” accuracy is not good enough - As models get more intelligent, there will be a barbell of boutique data vendors (e.g. any SMB) any scaled up data providers. - Models still exhibit “jagged intelligence” where they still fail on a long tail of edge cases. - There might always be opportunities to gather unique data for a given task and posttrain models for lower cost and higher accuracy. This marks #003 in our founder dinner series. What topic should we discuss next? Let us know your thoughts below!