【AI 核心深度 M5-054】写出 GRPO 的优势估计与损失。(Mathematical Formulation of GRPO Advantage Estimation and Loss)深度数理推导与工程落地解析

所属模块:M5 · NLP 与大语言模型 (NLP & Large Language Models) | 专题分类:GRPO 与推理模型 (GRPO & Reasoning Models) | 难度等级:Easy

一、核心一句话结论 (One-Sentence Summary)

对每个 prompt 采样一组回答,用组内奖励的均值/标准差归一化作为优势(替代 Critic),再做 PPO 式裁剪优化。

ADVERTISEMENT · 赞助推荐

Group Relative Policy Optimization (GRPO) samples a group of candidate outputs for each prompt and estimates advantage using normalized intra-group reward z-scores, eliminating the Critic value network and slashing training VRAM.

二、核心考点要义 (Key Insights)

  • 📌 对每个 prompt 采样 G 个回答(一个 group)
  • 📌 优势 = (r_i − 组内均值)/组内标准差
  • 📌 去掉 Critic,用组内统计作为基线

English Insights:
– Critic elimination: PPO requires a massive Critic model ($V_psi$) to estimate baseline values; GRPO completely discards the Critic network, saving $>50%$ training memory
– Group advantage estimation: samples $G$ completions ${y_1, dots, y_G}$ per prompt $x$; computes relative advantage by normalizing rewards using intra-group mean and standard deviation
– Core engine of DeepSeekMath and DeepSeek-R1: enabled training reasoning models over 10,000+ token chains where Critic value networks fail to converge

三、核心数学原理与机理推导 (Mathematical Principles & Derivation)

$$hat A_i=frac{r_i-mathrm{mean}({r_j})}{mathrm{std}({r_j})};qquad mathcal{L}_{text{GRPO}}=-mathbb{E}left[frac1Gsum_imin!left(rho_ihat A_i, mathrm{clip}(rho_i,1pmepsilon)hat A_iright)-betamathrm{KL}right]$$

数学机理:GRPO(Group Relative Policy Optimization,DeepSeekMath) 的核心创新是用’组内相对奖励’替代 Critic。(1) 采样组(group)——对每个 prompt x,用当前策略采样 G 个回答(如 G=8~64):{y_1, …, y_G};(2) 计算奖励——对每个回答算奖励 r_i(可为 RM 打分或可验证奖励如答案对错);(3) 组内归一化作为优势——Â_i=(r_i − mean(r))/std(r);这表示’该回答比同组平均好多少(以标准差为单位)’;(4) PPO 式优化——用 Â_i 作为优势,做与 PPO 相同的裁剪优化(含 KL 惩罚)。为什么组内均值是好基线——(a) 消去’题目难度’——同一 prompt 的 G 个回答共享相同难度,故组内比较自动扣除难度(比 Critic 估计的价值更准);(b) 无需 Critic——省一个模型(显存与训练复杂度大降);(c) 低方差——组内 G 个样本的均值是比单样本回报更好的基线。与 PPO 的对比——PPO 用 Critic V(s) 估计基线(需训练价值网络、且状态空间巨大难以准确);GRPO 用’同 prompt 的样本均值’(无需训练、天然准确)。与 RLOO 的关系——RLOO 用’留一法均值’(对第 i 个样本用其余样本的均值作基线);GRPO 用’全组均值 + 标准差归一化’(更简单且归一化使尺度一致)。优势的广播——Â_i 是序列级的(整个回答一个值),广播到该回答的所有 token(因为奖励是序列级的)。G 的选择——G 越大基线越准但采样成本越高(∝G);常用 8~64。注意——组内标准差为 0 时(所有回答奖励相同)优势为 0(无梯度),这是’全对或全错’的常见情况(需用其他技巧,如 DAPO 的动态采样)。

📖 查看英文严格数学推导 (English Mathematical Derivation)

Mathematical Mechanism: 1. Group Sampling & Advantage Estimation: For each prompt $x$, sample a group of $G$ outputs from the old policy: ${y_1, y_2, dots, y_G} sim pi_{theta_{text{old}}}(cdot mid x)$. An outcome verifier or reward model computes scalar rewards ${r_1, r_2, dots, r_G}$. The advantage for output $i$ is computed via intra-group z-score normalization: $$hat{A}_i = frac{r_i – text{mean}({r_1, dots, r_G})}{text{std}({r_1, dots, r_G}) + epsilon} = frac{r_i – frac{1}{G}sum_{j=1}^G r_j}{sqrt{frac{1}{G}sum_{j=1}^G (r_j – bar{r})^2} + epsilon}$$ 2. GRPO Clipped Objective: Optimize policy $pi_theta$ by maximizing the clipped surrogate loss with an explicit token-level KL divergence penalty against reference policy $pi_{text{ref}}$: $$mathcal{L}_{text{GRPO}}(theta) = frac{1}{G} sum_{i=1}^G frac{1}{|y_i|} sum_{t=1}^{|y_i|} left[ minleft( frac{pi_theta(y_{i,t} mid x, y_{i,<t})}{pi_{theta_{text{old}}}(y_{i,t} mid x, y_{i,<t})} hat{A}_i, , text{clip}left(frac{pi_theta}{pi_{theta_{text{old}}}}, 1-epsilon, 1+epsilonright) hat{A}_i right) – beta D_{text{KL}}(pi_theta , | , pi_{text{ref}}) right]$$ Where the KL divergence uses the unbiased estimator: $D_{text{KL}}(pi_theta | pi_{text{ref}}) = frac{pi_{text{ref}}(y_{i,t})}{pi_theta(y_{i,t})} – log frac{pi_{text{ref}}(y_{i,t})}{pi_theta(y_{i,t})} – 1$.

四、工业级落地权衡与工程考量 (Industrial Trade-offs)

深度剖析与工程权衡:① ‘组内相对’的优雅性——它把’绝对奖励’转为’相对排名’,天然处理了’奖励尺度不定’与’难度差异’两个问题;这是 GRPO 在 LLM 场景优于 PPO 的关键。② G 的成本-收益——采样 G 个回答的成本 ∝G(自回归生成),但 G 越大优势估计越准;实践中 G=8~16 常见(G=64 用于小模型或困难任务)。③ ‘全对/全错’的退化——若某 prompt 的所有 G 个回答都对(或都错),则组内奖励无差异、优势为 0(无梯度);这浪费了采样成本。DAPO 的’动态采样’正是解决此问题(过滤掉这类 prompt)。④ 与可验证奖励的天然契合——GRPO 的奖励可以是’答案是否正确’(0/1),此时组内归一化直接给出’相对好坏’;这使 GRPO 成为 RLVR 的主流算法(DeepSeek-R1 采用)。⑤ KL 惩罚仍需——GRPO 保留了 KL 惩罚(防偏离参考模型与语言退化);这是与 DPO 的重要区别(DPO 的 KL 是隐式的)。⑥ 面试要点——被问’GRPO 是什么’,应写出’组内采样 → 组内均值/标准差归一化作为优势 → PPO 式裁剪‘并解释’组内基线消去难度、省去 Critic‘;能指出’全对/全错时优势为 0(需动态采样)’是深度理解的标志。

⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)

Deep Dive & Engineering Trade-offs: ① Why GRPO Outperforms PPO in Reasoning: In multi-step mathematical reasoning (10,000+ tokens), a Critic network must predict expected return $V(s_t)$ at every intermediate token. If the model makes a tiny error on token 5,000, value prediction collapses, injecting massive noise into TD residuals. GRPO bypasses intermediate token baselines entirely by using the empirical group mean of identical prompts as the baseline. ② Optimal Group Size $G$: Typically $G in [8, 16]$. If $G$ is too small (e.g., $G=2$), standard deviation estimation is noisy; if $G$ is too large (e.g., $G=64$), rollout generation compute dominates training time. ③ Zero-Variance Edge Case: If all $G$ outputs receive identical rewards (e.g., all 0 on a difficult problem the model cannot solve, or all 1 on a trivial problem), $text{std}({r}) = 0$. In this case, $hat{A}_i = 0$ for all samples, safely producing zero gradient updates and preventing noise injection. ④ Memory Footprint: Without the Critic network and optimizer states, training a 70B model requires only 2 models (Actor + Reference), enabling training on standard 8-GPU nodes without multi-node pipeline parallelism. ⑤ Interview Strategy: Write down the intra-group z-score advantage formula, explain why eliminating the Critic stabilizes long-chain reasoning, and detail the GRPO clipped loss formulation.

五、常见面试避坑陷阱 (Common Pitfalls & Traps)

  • ⚠️ 以为 GRPO 完全不需要基线(用组内均值作基线)
  • ⚠️ 忽略’全对/全错’导致优势为 0 的退化

English Pitfalls:
– Attempting to run GRPO with group size $G=1$ (intra-group advantage calculation requires $G ge 2$, practically $G ge 4$)
– Forgetting to add $epsilon$ in the advantage denominator (causes division by zero when all outputs in a group receive identical reward)
– Assuming GRPO requires training a separate Value network (GRPO is strictly Critic-free)

六、高频深度面试追问与预测 (Follow-Up Questions)

  1. 为什么组内均值是好的基线?
  2. Why is the empirical group mean an unbiased baseline for policy gradient estimation?
  3. G 取多大合适?
  4. What happens to GRPO advantage calculations when all $G$ generated completions fail a mathematical test?

七、知识图谱对齐 (Knowledge Graph Anchor)

  • 🔗 关联底层卡片:GRPO 组相对策略优化:DeepSeek-R1 纯强化学习推理与长思考链涌现 (GRPO: Group Relative Policy Optimization & DeepSeek-R1)
  • 🗺️ 知识图谱模块:大语言模型全景图谱

🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)

本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。

👉 前往 TalentMe 交互式研读本题 (M5-054) →


Discover more from AirSOTA – Air School Of Thoughts AtoZ

Subscribe to get the latest posts sent to your email.