Sebastian Raschka 讲解对数概率评分与自我改进
从零开始推理,第五轮!这次,讨论对数概率评分(也是预训练和蒸馏中交叉熵等损失函数的一个很好的基础概念)以及自我改进。
00:00 引言和推理时扩展回顾 05:02 加载预训练的 LLM 08:00 比较和评分模型答案 10:18 构建基于规则的评分器 17:53 词元概率和序列似然 26:47 在 PyTorch 中计算词元概率 30:12 词元索引和移位目标 37:27 对数概率和数值稳定性 45:57 用平均对数概率评分答案 56:24 自我改进的工作原理 59:07 生成评议和修订答案 1:01:00 实现自我改进循环 1:05:57 MATH-500 评估结果 1:07:35 要点和下一步
以及 YouTube 上视频的链接: https://www.youtube.com/watch?v=TVMyOJ_3Gxo
对照原文
Reasoning from scratch, round number 5! This time, talking about log-probability scoring (also a great fundamental concept for loss functions like cross-entropy in pre-training and distillation) and self-refinement. 00:00 Introduction and inference-time scaling recap 05:02 Loading the pretrained LLM 08:00 Comparing and scoring model answers 10:18 Building a rule-based scorer 17:53 Token probabilities and sequence likelihood 26:47 Computing token probabilities in PyTorch 30:12 Token indexing and shifted targets 37:27 Log probabilities and numerical stability 45:57 Scoring answers with average log probabilities 56:24 How self-refinement works 59:07 Generating critiques and revised answers 1:01:00 Implementing the self-refinement loop 1:05:57 MATH-500 evaluation results 1:07:35 Takeaways and next steps
And a link to the video on YouTube: https://t.co/rDDvPaABO6