Raschka 分享从零实现 RLVR 与 GRPO 的教程
从零开始推理,第6轮!
可验证奖励强化学习(RLVR)和群体相对策略优化(GRPO)的介绍(及实现)。
00:00 引言
01:54 什么让推理模型与众不同?
04:25 推理轨迹与模型能力
08:29 准确性和格式奖励
11:34 顿悟时刻与DeepSeek-R1训练
14:41 推理努力与答案长度
18:38 RLHF与RLVR
23:04 GRPO对比PPO
26:40 用烹饪类比解释GRPO
31:43 KL项与简化GRPO
35:04 加载预训练模型
36:07 加载MATH训练数据
39:26 采样模型响应
46:30 计算可验证奖励
49:55 计算优势
51:54 标记和序列对数概率
55:29 实现序列对数概率
57:37 修复推理模式错误
1:02:24 计算GRPO损失
1:04:37 整合GRPO步骤
1:09:19 GRPO训练循环
1:12:57 训练设置、日志记录和检查点
1:17:24 运行训练并检查输出
1:19:28 加载和评估检查点
1:22:33 MATH-500结果与训练稳定性
1:24:05 内存需求与后续步骤
以及 YouTube 版本的链接:https://www.youtube.com/watch?v=237Hf7Q3lgg
对照原文
Reasoning from scratch, round number 6! An introduction (and implementation) of Reinforcement Learning with Verifiable Rewards (RLVR) and Group Relative Policy Optimization (GRPO). 00:00 Introduction 01:54 What makes a reasoning model different? 04:25 Reasoning traces and model capability 08:29 Accuracy and format rewards 11:34 Aha moments and DeepSeek-R1 training 14:41 Reasoning effort and answer length 18:38 RLHF and RLVR 23:04 GRPO vs. PPO 26:40 GRPO explained with a cooking analogy 31:43 The KL term and simplified GRPO 35:04 Loading the pretrained model 36:07 Loading the MATH training data 39:26 Sampling model responses 46:30 Computing verifiable rewards 49:55 Computing advantages 51:54 Token and sequence log probabilities 55:29 Implementing sequence log probabilities 57:37 Fixing the inference-mode error 1:02:24 Computing the GRPO loss 1:04:37 Putting the GRPO step together 1:09:19 The GRPO training loop 1:12:57 Training settings, logging, and checkpoints 1:17:24 Running training and inspecting outputs 1:19:28 Loading and evaluating checkpoints 1:22:33 MATH-500 results and training stability 1:24:05 Memory requirements and next steps
And a link to the YouTube version: https://t.co/K1cSuA01iG