【AI 核心深度 M5-038】解释 RLHF 的 on-policy 采样与 off-policy 数据。(On-Policy Sampling vs. Off-Policy Data in RLHF)深度数理推导与工程落地解析

所属模块:M5 · NLP 与大语言模型 (NLP & Large Language Models) | 专题分类:对齐与 RLHF (Alignment & RLHF) | 难度等级:Hard

一、核心一句话结论 (One-Sentence Summary)

PPO 需 on-policy 采样(当前策略生成)保证重要性采样正确;RM 训练常用 off-policy 数据;迭代式 RLHF 混合两者。

ADVERTISEMENT · 赞助推荐

RLHF requires on-policy sampling so that policy updates correspond strictly to the current distribution of generated tokens, while off-policy buffer reuse introduces severe distribution shifts and advantage estimation bias.

二、核心考点要义 (Key Insights)

  • 📌 PPO 的 ratio 需样本来自采样时的策略(on-policy)
  • 📌 RM 可用任意策略的数据(off-policy,更省)
  • 📌 迭代式 RLHF:用最新策略的数据更新 RM(近似 on-policy)

English Insights:
– On-policy requirement: policy gradient theorems require expectations taken over trajectories generated by the current policy: $nabla_theta J(theta) = mathbb{E}{y sim pitheta} [nabla log pi_theta(y mid x) A(x, y)]$
– Off-policy drift: if rollouts are collected from an older policy $pi_{text{old}}$ and reused for multiple PPO epochs, importance sampling weights $r_t(theta) = frac{pi_theta}{pi_{text{old}}}$ explode or vanish, destabilizing updates
– Engineering tension: on-policy sampling requires continuous, expensive generation rollouts during training, making RLHF $5times$ to $10times$ slower than offline supervised learning

三、核心数学原理与机理推导 (Mathematical Principles & Derivation)

$$text{PPO}: text{samples}simpi_{text{old}} (text{on-policy});qquad text{RM}: text{data from various policies} (text{off-policy})$$

数学机理:on-policy vs off-policy 的含义——on-policy 指’用当前策略生成的数据’更新策略;off-policy 指’用其他策略(如旧策略、人类、其他模型)生成的数据’。为什么 PPO 需 on-policy——PPO 的目标含重要性采样比 r=πθ/π_old,其无偏性要求样本来自 π_old(采样时的策略);若样本来自很旧的策略,则 r 的方差会爆炸(因为 πθ 与采样策略差异大)。故 PPO 每轮采样后只能做有限次更新(受 clip 限制),然后必须重新采样(这就是’on-policy、采样昂贵’的来源)。RM 训练可用 off-policy——奖励模型是监督学习(拟合偏好标签),不依赖重要性采样,故可用任意来源的偏好数据:人类标注的(来自多种模型/人工)、历史策略的、其他模型的;这使 RM 训练更省(可离线、可复用数据)。但 on-policy 数据对 RM 更有效——因为 RM 需在’策略会访问的区域’可靠;若 RM 只用’旧策略/人类’的数据训练,而新策略探索到新区域,RM 在那里不可靠(导致过优化)。故 iterative/online RLHF 用最新策略的输出做偏好标注来更新 RM(近似 on-policy),使 RM 的可靠区域’跟随’策略。三者的组合——(a) PPO:严格 on-policy 采样(每轮重新生成);(b) RM 初始训练:off-policy(用现有偏好数据);(c) RM 迭代更新:on-policy(用最新策略输出 + 人工/AI 标注)。成本权衡——on-policy 更有效但更贵(需实时采样 + 实时标注);off-policy 便宜但可能失配。实践常’off-policy 冷启动 + on-policy 迭代’。

📖 查看英文严格数学推导 (English Mathematical Derivation)

Mathematical Mechanism: 1. Importance Sampling Correction: When evaluating policy $pi_theta$ using trajectories sampled from behavior policy $pi_{beta}$ (off-policy): $$mathbb{E}_{y sim pi_theta}[R(x, y)] = mathbb{E}_{y sim pi_beta}left[ prod_{t=1}^T frac{pi_theta(y_t mid x, y_{<t})}{pi_beta(y_t mid x, y_{<t})} R(x, y) right]$$ In language models where $T in [500, 2048]$, the cumulative importance weight is a product of hundreds of ratios: $$w(y) = prod_{t=1}^T frac{pi_theta(y_t)}{pi_beta(y_t)}$$ If $pi_theta$ drifts slightly from $pi_beta$, $w(y)$ suffers from exponential variance explosion: either collapsing to zero for almost all samples or exploding to infinity on a single outlier sample, destroying gradient updates. 2. PPO Off-Policy Epoch Horizon: PPO allows mild off-policy reuse by collecting a batch with $pi_{text{old}}$, training for 1-4 gradient epochs with clipping ($r_t in [1-epsilon, 1+epsilon]$), and immediately discarding the batch to collect fresh on-policy samples.

四、工业级落地权衡与工程考量 (Industrial Trade-offs)

深度剖析与工程权衡:① ‘on-policy 采样昂贵’是 PPO 的主要成本——每轮都要用当前策略自回归生成(慢)、且生成的数据用一次就废(受 clip 限制);故 RLHF 的吞吐远低于 SFT。这是 DPO(用固定偏好数据,无采样)更经济的原因。② ‘RLAIF 让 on-policy 变便宜’——用 AI(而非人工)标注新策略的输出,使 on-policy RM 更新的成本大幅下降;这是 iterative RLHF / online DPO 能规模化的关键。③ 与’分布偏移’的关系——off-policy 数据与当前策略的分布差异越大,RM 的可靠性越差(因为训练分布与使用分布不匹配);这与监督学习的’训练-测试分布一致’原则同理。④ DPO 的 off-policy 本质——DPO 用固定的偏好数据集训练,且其推导假设’偏好数据来自参考模型’;若数据来自其他模型(如人类写的),则 DPO 的假设被违反,效果下降(这是 DPO 的局限之一,催生了 online DPO)。⑤ ‘数据效率’的对比——on-policy 数据的’每样本价值’更高(因为它直接反映当前策略的弱点),但获取成本高;off-policy 数据便宜但可能浪费在’已学会的’区域。⑥ 面试要点——被问’为什么 PPO 要 on-policy’,应给出’重要性采样比的无偏性要求样本来自采样时的策略‘,并说明’RM 可 off-policy 但 on-policy 数据更有效 → 迭代式 RLHF‘;能指出’DPO 用固定数据是其局限、online DPO 修正它’是深度理解的标志。

⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)

Deep Dive & Engineering Trade-offs: ① Rollout Freshness vs GPU Throughput: To maximize throughput, engineers want to train on a generated batch for many epochs. In LLMs, training for $>2text{–}4$ epochs on the same rollout batch leads to immediate policy collapse due to importance weight divergence. Batches must be refreshed constantly. ② Decoupled Asynchronous Rollouts: Frameworks like Ray-RLHF run dedicated inference engines (vLLM) that continuously generate rollouts into a shared queue while trainer nodes pull the freshest rollouts, overlapping generation and training without falling behind on-policy freshness. ③ Contrast with DPO (Offline Off-Policy): DPO is fully offline: it trains on static preference datasets without any on-policy generation rollouts. However, offline DPO suffers from distribution shift because the student never learns to correct its own generation errors (addressed by Online DPO). ④ Experience Replay Buffers in RLHF: Traditional RL experience replay buffers (storing samples across thousands of past steps) fail completely in LLM RLHF due to the exponential length product $w(y)$. Replay buffers are strictly restricted to 1 step. ⑤ Interview Strategy: Write down the cumulative sequence importance sampling formula $prod_t frac{pi_theta}{pi_beta}$, explain why variance explodes exponentially with sequence length $T$, and justify why PPO restricts batch reuse to 1-2 epochs.

五、常见面试避坑陷阱 (Common Pitfalls & Traps)

  • ⚠️ 用很旧的策略数据做 PPO(重要性采样方差爆炸)
  • ⚠️ RM 只用 off-policy 数据(分布外不可靠)

English Pitfalls:
– Attempting to reuse rollout data across dozens of PPO training epochs to save compute (causes immediate divergence)
– Using a traditional deep Q-learning experience replay buffer for full autoregressive LLM trajectories
– Assuming DPO is on-policy (standard DPO is completely offline and off-policy)

六、高频深度面试追问与预测 (Follow-Up Questions)

  1. 为什么 PPO 必须 on-policy?
  2. Why does the product of token-level importance weights make multi-epoch off-policy training practically impossible in LLMs?
  3. on-policy RM 数据为什么更有效?
  4. How does Online / Iterative DPO bridge the gap between offline simplicity and on-policy freshness?

七、知识图谱对齐 (Knowledge Graph Anchor)

  • 🔗 关联底层卡片:RLHF 人类反馈强化学习:Bradley-Terry 奖励模型、PPO 剪切目标与 KL 惩罚 (RLHF: Bradley-Terry Reward Modeling, PPO & KL Penalties)
  • 🗺️ 知识图谱模块:大语言模型全景图谱

🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)

本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。

👉 前往 TalentMe 交互式研读本题 (M5-038) →


Discover more from AirSOTA – Air School Of Thoughts AtoZ

Subscribe to get the latest posts sent to your email.