所属模块:
M5 · NLP 与大语言模型 (NLP & Large Language Models)| 专题分类:对齐与 RLHF (Alignment & RLHF)| 难度等级:Hard
一、核心一句话结论 (One-Sentence Summary)
动作是 token 但奖励是序列级;实现上 ratio 与 KL 按 token 算、优势广播到 token;只对回答 token 算损失。
PPO in LLMs operates over a hybrid token-sequence hierarchy where rewards are evaluated at the sequence level, KL penalties are computed at the token level, and advantages are attributed per-token via the Critic value network.
二、核心考点要义 (Key Insights)
- 📌 动作 = 每个 token;奖励 = 整个回答一个标量
- 📌 ratio 与 KL 是 token 级;优势是序列级广播到 token
- 📌 损失只算回答 token(指令部分 mask 掉)
English Insights:
– Sequence-level reward: human preference models evaluate the complete response as a single atomic scalar $R_phi(x, y)$
– Token-level KL penalty: divergence against reference policy is computed at every token step: $text{KL}t = log frac{pitheta(y_t)}{pi_{text{ref}}(y_t)}$
– Credit assignment challenge: how to distribute a single sequence reward $R_phi$ backward across hundreds of intermediate tokens (resolved via Critic temporal difference learning)
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$r_t=frac{pi_theta(y_t|x,y_{<t})}{pi_{text{old}}(y_t|x,y_{<t})};qquad A_t=hat A_{text{seq}} text{(broadcast)};qquad text{loss on response tokens only}$$
数学机理:LLM 中 RL 的形式化——把生成过程建模为 MDP:(a) 状态 s_t=(x, y_{<t})(prompt + 已生成的 token);(b) 动作 a_t=y_t(下一个 token);(c) 奖励 r 只在序列末尾给出(奖励模型对整个回答打分)——这是关键的特殊性(稀疏奖励);(d) 转移是确定的(拼接 token)。实现细节:(1) ratio 是 token 级——r_t=πθ(y_t|x,y{<t})/πold(y_t|x,y{<t}),即每个 token 的概率比;PPO 的 clip 作用在每个 token 上。(2) 优势的处理——严格地,优势 A_t 应反映’在 s_t 采取 a_t 的长期收益’;但 LLM 中奖励只在末尾,且 γ=1 时中间步骤的 TD 残差主要由 Critic 给出。实践中常见两种简化:(a) 用 Critic + GAE(标准做法,但 Critic 难训);(b) 直接把序列级优势(或’奖励−基线’)广播到所有 token(简化,GRPO/部分实现采用)。(3) KL 是 token 级——序列级 KL 分解为逐 token KL 之和,故实现为’每个 token 的奖励减 β·log(π_θ/π_ref)’。(4) 损失只算回答 token——prompt 部分 mask 掉(与 SFT 一致),因为只有回答是’策略生成的’。(5) 序列级聚合——最终损失是’所有回答 token 的 PPO 损失之和/平均’(可选按长度归一化)。与其他实现差异——(a) GRPO:去掉 Critic,用同一 prompt 的多个回答的平均奖励作为基线,优势 = (r_i − mean)/std(组内归一化),再广播到该回答的所有 token;(b) RLOO:用’留一法’的均值作为基线;(c) DPO:完全不做 RL 循环,直接用偏好对做对比损失。关键工程点——(a) 采样用 inference engine(vLLM)加速、训练用 trainer(权重同步);(b) 计算 log-prob 需要’重新前向’(因为采样时的 log-prob 可能未保存);(c) 数值稳定(log-ratio 的精度)。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Mathematical Mechanism: 1. Hybrid Reward Assignment: For generated sequence $y = [y_1, dots, y_T]$, the composite per-token reward signal $r_t$ fed into PPO is structured as: $$r_t = begin{cases} -beta log frac{pi_theta(y_t mid x, y_{<t})}{pi_{text{ref}}(y_t mid x, y_{<t})} & text{if } t < T \ R_phi(x, y) – beta log frac{pi_theta(y_T mid x, y_{<T})}{pi_{text{ref}}(y_T mid x, y_{<T})} & text{if } t = T end{cases}$$ 2. Critic Temporal Credit Assignment: The Critic network $V_psi(s_t)$ predicts the expected cumulative remaining reward from token position $t$: $$V_psi(s_t) approx mathbb{E}left[ sum_{l=t}^T r_l right]$$ The 1-step TD residual is: $delta_t^V = r_t + gamma V_psi(s_{t+1}) – V_psi(s_t)$ (with $gamma = 1.0$). Generalized Advantage Estimation (GAE) propagates this backward: $$hat{A}_t = delta_t^V + lambda hat{A}_{t+1}$$ Tokens that set up a breakthrough reasoning step or toxic deviation receive high or low advantage estimates through the learned Critic baseline.
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
深度剖析与工程权衡:① ‘奖励稀疏’是 LLM RL 的核心特殊性——标准 RL 每步都有奖励,LLM 只有末尾一个;这使 (a) Critic 的训练更难(需从稀疏信号学价值)、(b) ‘把序列优势广播到 token’的近似更常见。② ‘token 级 KL’的必要性——若只在序列末尾加 KL,则中间的偏离不受约束;逐 token KL 提供了’每步的约束’,是防止语言退化的关键。③ ‘重新前向算 log-prob’的开销——采样时通常不保存每个 token 的 log-prob(为省显存),故训练时需用当前/旧策略重新前向计算;这增加了一次前向的开销(是 RLHF 慢的原因之一)。④ 与’长度归一化’的关系——序列级损失若按’token 数平均’,则长回答的每个 token 权重更小;若按’序列求和’,则长回答影响更大。选择会影响’是否鼓励长输出’。⑤ GRPO 的简化价值——去掉 Critic 省了一个模型与一套训练逻辑,且’组内均值基线’在 LLM 中效果良好;这使 GRPO 成为当前推理模型 RL 的主流(DeepSeek-R1 采用)。⑥ 面试要点——被问’LLM 里的 PPO 与标准 PPO 有何不同’,应给出’动作=token、奖励=序列级、ratio/KL 为 token 级、优势广播到 token、只算回答 token‘,并说明’GRPO 用组内均值替代 Critic’;能指出’需重新前向算 log-prob’这一工程细节是深度理解的标志。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
Deep Dive & Engineering Trade-offs: ① Pure Sequence-Level RL (REINFORCE): Assigning the identical sequence reward $R$ to every token without a Critic: $hat{A}_t = R – bar{R}$. This avoids the complex Critic model entirely, but suffers from extreme gradient variance. Works well only when paired with large group sampling (as in GRPO). ② Critic Value Head Architecture: In modern LLM PPO, the Critic shares the Transformer backbone with the Actor (adding a linear value head) or runs as an independent identical-scale model. Independent models prevent Actor representation degradation, but double VRAM footprint. ③ Value Function Loss Clipping: Similar to policy clipping, PPO clips value updates: $L^{text{VF}} = maxleft((V – R)^2, (text{clip}(V, V_{text{old}}-epsilon, V_{text{old}}+epsilon) – R)^2right)$, preventing drastic value function jumps on noisy rewards. ④ Position-Dependent Bias: Initial tokens often receive near-zero advantage because the Critic’s prediction $V(s_1)$ closely matches the mean return $bar{R}$, while late tokens experience sharper variance based on specific concluding choices. ⑤ Interview Strategy: Write down the hybrid reward formula (token KL on $t < T$, scalar RM on $t=T$), explain how Critic TD residuals perform temporal credit assignment, and compare Critic-based PPO with Critic-free GRPO.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 把奖励当作 token 级(实际是序列级)
- ⚠️ 只在序列末尾加 KL(中间不受约束)
English Pitfalls:
– Broadcasting the final scalar reward $R_phi$ uniformly to all token positions without subtracting the baseline Critic value
– Omitting the token-level KL penalty during intermediate generation steps
– Failing to clip the value function loss during Critic training
六、高频深度面试追问与预测 (Follow-Up Questions)
- 为什么优势要广播到每个 token?
- Why does Critic-free advantage estimation (like GRPO) require large group sampling to match PPO’s stability?
- token 级 KL 与序列级 KL 的关系?
- What causes value function clipping to protect against Critic gradient explosion?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
RLHF 人类反馈强化学习:Bradley-Terry 奖励模型、PPO 剪切目标与 KL 惩罚(RLHF: Bradley-Terry Reward Modeling, PPO & KL Penalties) - 🗺️ 知识图谱模块:
大语言模型全景图谱
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。