【AI 核心深度 M5-043】比较 PPO、DPO 与 GRPO 的依赖与适用场景。(Systematic Comparison of PPO, DPO, and GRPO)深度数理推导与工程落地解析

所属模块:M5 · NLP 与大语言模型 (NLP & Large Language Models) | 专题分类:DPO 家族 (DPO Family (DPO / KTO / ORPO)) | 难度等级:Easy

一、核心一句话结论 (One-Sentence Summary)

PPO 需 RM+Critic+在线采样;DPO 只需偏好对与参考模型;GRPO 去 Critic、用组内基线、需可验证奖励。

ADVERTISEMENT · 赞助推荐

PPO uses an explicit Reward Model and Critic network for online policy gradients, DPO uses offline preference pairs to optimize an implicit reward without reinforcement learning, and GRPO eliminates the Critic by computing relative advantages across prompt rollout groups.

二、核心考点要义 (Key Insights)

  • 📌 PPO:最强但最复杂(4 模型、在线采样)
  • 📌 DPO:最简单(无需 RM/RL),但无在线探索
  • 📌 GRPO:介于两者(无 Critic,但需采样与奖励)

English Insights:
– PPO (Schulman et al. / OpenAI): 4 models (Actor, Critic, Reference, RM); online rollouts; highly flexible and gold standard for open-ended exploration, but complex and unstable
– DPO (Rafailov et al. / Stanford): 2 models (Actor, Reference); 100% offline static preference pairs; stable supervised cross-entropy training, but limited by offline data distribution
– GRPO (DeepSeekMath / DeepSeek-R1): 2 models (Actor, Reference); online rollouts with group sampling ($G$ outputs per prompt); eliminates Critic by computing normalized intra-group advantage; dominant for math and code reasoning

三、核心数学原理与机理推导 (Mathematical Principles & Derivation)

$$text{PPO}: text{RM}+text{Critic}+text{on-policy};quad text{DPO}: text{pairs}+pi_{text{ref}};quad text{GRPO}: text{group baseline}, text{no Critic}$$

数学机理:三者的依赖与机制。PPO(RLHF 标准)——依赖:(a) 奖励模型(RM,需偏好数据训练);(b) Critic(价值网络,需训练);(c) Reference(参考模型,算 KL);(d) 在线采样(每轮用当前策略生成)。优点——最强的能力(在线探索、可迭代改进、可处理任意奖励)。缺点——工程复杂(4 模型、显存大)、训练不稳、采样慢、超参多。DPO——依赖:(a) 偏好对数据;(b) 参考模型(通常是 SFT 模型)。机制——直接把偏好对做对比损失(见推导),无需 RM、无需 RL 循环、无需采样。优点——简单稳定(像 SFT)、成本低、易复现。缺点——(a) 无在线探索(用固定数据,无法探索数据外的区域);(b) 假设数据来自参考策略(否则效果下降);(c) 无法利用’未通过验证’的样本(不像 RL 可用部分奖励)。GRPO(Group Relative Policy Optimization)——依赖:(a) 奖励函数(可为 RM 或可验证奖励如答案对错);(b) 在线采样(每个 prompt 采样一组回答);去掉 Critic(用组内平均奖励作为基线)。机制——对每个 prompt 采样 G 个回答,计算各自奖励,用 (r_i − mean)/std 作为优势(组内归一化),再用 PPO 式损失优化(但无 Critic)。优点——比 PPO 省一个模型、比 DPO 有在线探索能力;特别适合可验证奖励(数学/代码:奖励是’答案是否正确’)。缺点——需采样(慢)、奖励需可靠、仍有 RL 的不稳定性。选择逻辑——(a) 有偏好数据、求简单稳定 → DPO;(b) 有可验证奖励(数学/代码)、求推理能力 → GRPO/RLVR;(c) 需最强对齐、有充足资源 → PPO(+ 迭代 RM);(d) 实践路径:DPO 起步 → 不足再上 GRPO/PPO。

📖 查看英文严格数学推导 (English Mathematical Derivation)

Mathematical Mechanism: Structural Comparison Matrix: 1. PPO: $$mathcal{L}_{text{PPO}} = hat{mathbb{E}}_t left[ min(r_t hat{A}_t^{text{GAE}}, text{clip}(r_t, 1-epsilon, 1+epsilon) hat{A}_t^{text{GAE}}) right], quad hat{A}_t = Q(s_t, a_t) – V_psi(s_t)$$ Requires Critic network $V_psi$ to learn state-value baselines. 2. DPO: $$mathcal{L}_{text{DPO}} = -mathbb{E}_{(x, y_w, y_l)} left[ log sigmaleft( beta log frac{pi_theta(y_w mid x)}{pi_{text{ref}}(y_w mid x)} – beta log frac{pi_theta(y_l mid x)}{pi_{text{ref}}(y_l mid x)} right) right]$$ Zero online sampling; trains directly on offline dataset $mathcal{D}$. 3. GRPO (Group Relative Policy Optimization): Samples $G$ completions ${y_1, dots, y_G} sim pi_{theta_{text{old}}}(cdot mid x)$ per prompt $x$. Advantage for sample $i$ is normalized directly against the group mean and standard deviation: $$hat{A}_i = frac{R_i – text{mean}({R_1, dots, R_G})}{text{std}({R_1, dots, R_G}) + epsilon}$$ No Critic model needed! Slashes training memory and eliminates Critic value divergence completely.

四、工业级落地权衡与工程考量 (Industrial Trade-offs)

深度剖析与工程权衡:① ‘在线探索’是 PPO/GRPO 相对 DPO 的核心优势——在线采样使策略能探索’偏好数据之外’的回答空间,从而发现更好的回答;DPO 只能在固定数据的’支撑集’内优化(这也是’DPO 难以超越数据质量’的原因)。② ‘奖励的来源’决定方法选择——(a) 有人类偏好 → DPO/PPO;(b) 有可验证奖励(答案对错、测试通过) → GRPO/RLVR(最可靠);(c) 有AI 反馈 → RLAIF 系列。故’先问奖励从哪来’是选择方法的第一步。③ GRPO 的’组内基线’为何有效——同一 prompt 的多个回答共享相同的’难度’,故组内比较能消去’题目难度’这一混淆因素;这使基线估计更准(比 Critic 更简单且往往更好)。④ 与’推理模型’的关系——推理任务(数学/代码)天然有可验证奖励,故 GRPO/RLVR 成为推理模型训练的主流(DeepSeek-R1 等)。⑤ 成本对比——DPO(无采样、无 RM、无 Critic)<< GRPO(需采样,无 Critic)< PPO(需采样 + RM + Critic)。故’先用 DPO,不够再升级’是经济的选择。⑥ 面试要点——被问’PPO/DPO/GRPO 怎么选’,应给出’依赖(RM/Critic/采样)+ 在线探索能力 + 奖励类型 + 成本‘的四维对比与选择逻辑;能指出’奖励来源决定方法选择’是深度理解的标志。

⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)

Deep Dive & Engineering Trade-offs: ① Infrastructure Memory Footprint: – PPO: Actor + Critic + Ref + RM $implies sim 4times$ base parameters in VRAM. – DPO: Actor + Ref $implies sim 2times$ base parameters (or $1times$ if Ref is offloaded/quantized). – GRPO: Actor + Ref $implies sim 2times$ base parameters, but runs dynamic on-policy generation rollouts. ② Applicability Domains: – DPO is best for general alignment (conversational tone, safety, persona) where pairwise comparison data is readily available. – GRPO is best for verifiable reasoning (math, coding, logical puzzle solving) where rule-based execution environments provide scalar rewards and group exploration discovers novel reasoning paths. – PPO remains competitive in multi-objective complex environments where reward models are non-stationary and continuous RL exploration is mandatory. ③ Why GRPO Replaced PPO in DeepSeek-R1: Training a Critic model for long reasoning chains (10,000+ tokens) is virtually impossible because value predictions oscillate wildly at each intermediate step. GRPO bypasses the Critic entirely by using group baseline normalization. ⑤ Interview Strategy: Provide the 3-way taxonomy table across Models Required, Online vs Offline, Advantage Estimation Method, and Primary Use Cases, highlighting GRPO’s role in modern reasoning models.

五、常见面试避坑陷阱 (Common Pitfalls & Traps)

  • ⚠️ 认为 DPO 全面优于 PPO(DPO 无在线探索)
  • ⚠️ 在可验证任务上用 DPO 而非 GRPO/RLVR

English Pitfalls:
– Believing DPO is an online reinforcement learning algorithm (standard DPO is an offline supervised classification algorithm)
– Attempting to use GRPO without group sampling (group size $G ge 4$ is mandatory to compute intra-group standard deviation)
– Assuming PPO is obsolete (PPO remains superior in complex game-theoretic and multi-turn interactive environments)

六、高频深度面试追问与预测 (Follow-Up Questions)

  1. 什么场景该用 PPO 而非 DPO?
  2. Why is training a Critic network especially intractable for 10,000-token Chain-of-Thought reasoning models?
  3. GRPO 为什么适合推理任务?
  4. Under what conditions does offline DPO suffer compared to online GRPO exploration?

七、知识图谱对齐 (Knowledge Graph Anchor)

  • 🔗 关联底层卡片:直接偏好优化 (DPO):无奖励模型对齐闭式解、KTO 与 ORPO 对比 (Direct Preference Optimization (DPO), KTO & ORPO)
  • 🗺️ 知识图谱模块:大语言模型全景图谱

🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)

本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。

👉 前往 TalentMe 交互式研读本题 (M5-043) →


Discover more from AirSOTA – Air School Of Thoughts AtoZ

Subscribe to get the latest posts sent to your email.