返回 X 名人动态

推理时扩展教程:采样、自一致性与 best-of-N

中文全文 · AI 翻译
推文 1 / 2

推理扩展第1部分。 从一个修改后的文本生成函数开始(温度缩放、top-p过滤、多项分布采样),用于生成多样化输出以实现自一致性和最佳N选1(通过这种方式将答案准确率提高>2倍)

00:00 引言和回顾 00:31 训练时和推理时的扩展 07:52 我们将要实现的内容 11:47 笔记本设置和模型加载 17:43 构建灵活的文本生成函数 24:40 思维链提示 28:26 采样和输出多样性 33:43 下一个 token 的 logits和贪婪解码 38:20 温度缩放逐步解析 42:46 Softmax和token 概率 47:42 多项分布采样 54:51 将温度采样添加到文本生成中 59:31 Top-p过滤逐步解析 1:10:23 将top-p过滤添加到文本生成中 1:13:43 采样和LLM水印 1:16:01 自一致性和多数投票 1:20:36 实现自一致性 1:29:02 MATH-500结果 1:35:01 准确性和计算权衡 1:36:50 下一步和自精炼

推文 2 / 2

以及 YouTube 上视频的链接:https://youtu.be/t5y-kS9nNxU

对照原文

Inference scaling part 1. Starting with a modded text generation function (temperature scaling, top-p filtering, multinomial sampling) to generate diverse outputs for self-consistency and best-of-N (improving answer accuracy by>2x) 00:00 Introduction and recap 00:31 Training-time and inference-time scaling 07:52 What we'll implement 11:47 Notebook setup and model loading 17:43 Building a flexible text generation function 24:40 Chain-of-thought prompting 28:26 Sampling and output diversity 33:43 Next-token logits and greedy decoding 38:20 Temperature scaling step by step 42:46 Softmax and token probabilities 47:42 Multinomial sampling 54:51 Adding temperature sampling to text generation 59:31 Top-p filtering step by step 1:10:23 Adding top-p filtering to text generation 1:13:43 Sampling and LLM watermarking 1:16:01 Self-consistency and majority voting 1:20:36 Implementing self-consistency 1:29:02 MATH-500 results 1:35:01 Accuracy and compute tradeoffs 1:36:50 Next steps and self-refinement

And a link to the video on YouTube: https://t.co/9Y5WCVELZ4