【AI 核心深度 M5-030】描述 RLHF 的三阶段流程。(The Three-Stage RLHF Pipeline)深度数理推导与工程落地解析

所属模块:M5 · NLP 与大语言模型 (NLP & Large Language Models) | 专题分类:对齐与 RLHF (Alignment & RLHF) | 难度等级:Easy

一、核心一句话结论 (One-Sentence Summary)

① SFT 得到初始策略;② 用人类偏好比较训练奖励模型;③ 用 PPO 在奖励模型上优化策略(带 KL 约束)。

ADVERTISEMENT · 赞助推荐

The canonical RLHF pipeline consists of three sequential phases: Supervised Fine-Tuning to establish instruction following, Reward Modeling on human preference pairs, and PPO Reinforcement Learning with a KL penalty to optimize policy generation.

二、核心考点要义 (Key Insights)

  • 📌 阶段一:SFT(监督微调)建立’会说人话’的起点
  • 📌 阶段二:奖励建模(用偏好对训练打分器)
  • 📌 阶段三:PPO 优化(带 KL 约束防偏离)

English Insights:
– Stage 1: Supervised Fine-Tuning (SFT) establishes a competent initial policy $pi_{text{SFT}}$ capable of following instructions and generating well-formatted responses
– Stage 2: Reward Modeling (RM) trains a scoring network $r_phi(x, y)$ on human pairwise preferences $(y_w succ y_l)$ using the Bradley-Terry ranking loss
– Stage 3: Reinforcement Learning (PPO) optimizes policy $pi_theta$ to maximize reward $r_phi(x, y)$ while constraining policy divergence via a per-token KL penalty against $pi_{text{SFT}}$

三、核心数学原理与机理推导 (Mathematical Principles & Derivation)

$$text{SFT}totext{RM}: mathcal{L}=-logsigma(r_phi(x,y_w)-r_phi(x,y_l))totext{PPO}: maxmathbb{E}[r_phi]-betamathrm{KL}(pi|pi_{text{ref}})$$

数学机理:三阶段流程(InstructGPT 范式)。阶段一:SFT(Supervised Fine-Tuning)——用(指令,示范回答)对做监督微调,得到一个’能遵循指令、会说人话’的初始策略 πSFT。为什么必要——(a) 纯预训练模型只会’续写’,无法生成有意义的回答,故 PPO 的采样质量太差、无法启动;(b) SFT 提供了’回答的格式与基本行为’,为 RL 提供了合理的起点。阶段二:奖励建模(Reward Modeling)——让人类对同一 prompt 的多个回答做偏好比较(哪个更好),得到偏好对 (x, y_w, y_l);用 Bradley-Terry 模型假设’更好的回答有更高的潜在分数’,训练奖励模型 rφ 使 L=−log σ(r_φ(x,y_w)−r_φ(x,y_l)) 最小化。为什么用比较而非打分——人类做’哪个更好’比’打绝对分数’更一致、更可靠(绝对分数在不同人之间不可比)。阶段三:PPO 强化学习——用奖励模型作为环境奖励,用 PPO 优化策略 πθ:maxθ E_{x,y∼πθ}[rφ(x,y)] − β·KL(π_θ‖π_ref),其中 KL 项约束策略不偏离参考模型(π_ref,通常是 SFT 模型)太远。为什么需要 KL 约束——(a) 防止奖励黑客(策略钻奖励模型的空子);(b) 防止灾难性遗忘与’语言退化’(过度优化会让输出变得不自然);(c) 保持输出多样性。结果——InstructGPT 显示 RLHF 后的模型(1.3B)在人类偏好上优于大 100 倍的 GPT-3(175B),说明对齐的杠杆作用巨大。

📖 查看英文严格数学推导 (English Mathematical Derivation)

Mathematical Mechanism: 1. Stage 1: SFT Warmup: Standard cross-entropy optimization over demonstration dataset $mathcal{D}_{text{SFT}}$: $$pi_{text{SFT}} = argmax_theta mathbb{E}_{(x, y) sim mathcal{D}_{text{SFT}}} left[ sum_{t} log pi_theta(y_t mid x, y_{<t}) right]$$ 2. Stage 2: Bradley-Terry Reward Modeling: Given prompt $x$ and human-labeled preference $y_w succ y_l$, train scalar reward model $r_phi(x, y)$ via negative log-likelihood: $$mathcal{L}_{text{RM}}(phi) = -mathbb{E}_{(x, y_w, y_l)} left[ log sigmaleft( r_phi(x, y_w) – r_phi(x, y_l) right) right]$$ 3. Stage 3: PPO Policy Optimization: Optimize policy $pi_theta$ to maximize expected reward subject to KL divergence regularization: $$max_theta mathbb{E}_{x sim mathcal{D}, y sim pi_theta(cdot mid x)} left[ r_phi(x, y) – beta D_{text{KL}}left(pi_theta(y mid x) , | , pi_{text{SFT}}(y mid x)right) right]$$ Where per-token KL divergence acts as an immediate penalty: $r_{text{total}}(x, y) = r_phi(x, y) – beta sum_{t} log frac{pi_theta(y_t mid x, y_{<t})}{pi_{text{SFT}}(y_t mid x, y_{<t})}$.

四、工业级落地权衡与工程考量 (Industrial Trade-offs)

深度剖析与工程权衡:① ‘为什么用比较而非打分’是核心设计——人类的绝对打分受主观尺度影响(不同人、不同时间的’7 分’含义不同);比较是相对判断,更一致、更可复现。这使偏好数据的质量显著高于打分数据。② KL 系数 β 的作用——β 越大越保守(贴近参考模型,但可能学不到新行为);β 越小越激进(可能奖励黑客、语言退化)。实践中 β 常取 0.01~0.1,且可动态调整(当 KL 过大时增大 β)。③ 三阶段的必要性——(a) 无 SFT:PPO 无法启动(采样质量差);(b) 无 RM:无法把人类偏好转为可优化的标量奖励;(c) 无 PPO:只能模仿(无法超越演示质量)。三者缺一不可。④ 成本结构——三阶段中阶段二(人工偏好标注)最贵(需大量人工比较);这也是 DPO/RLAIF 等’减少人工’方法出现的动机。⑤ 与 DPO 的关系——DPO 把阶段二与三合并(直接用偏好对优化策略,无需显式奖励模型与 RL);代价是失去’在线采样’与’迭代改进’的能力(见 DPO 题)。⑥ 面试要点——被问’RLHF 流程’,应给出’SFT → RM → PPO(带 KL)‘三阶段与每阶段为什么必要,并解释’为什么用比较而非打分’与’KL 的双重作用’;这是 RLHF 类问题的基本盘。

⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)

Deep Dive & Engineering Trade-offs: ① Why Comparative Preference over Absolute Scoring: Absolute ratings (1 to 7 scale) suffer from severe inter-annotator bias (one annotator’s ‘5’ is another’s ‘7’). Pairwise comparison ($y_w succ y_l$) is intuitive, consistent, and achieves significantly higher inter-annotator agreement. ② The Role of the KL Penalty $beta$: Without the KL divergence penalty ($beta = 0$), the policy quickly exploits reward model blind spots, generating ungrammatical, verbose gibberish that scores high reward (reward hacking). $beta$ keeps the policy anchored within the natural language distribution of $pi_{text{SFT}}$. ③ Infrastructure Complexity of PPO: Stage 3 requires hosting four large models in GPU memory simultaneously: Actor (policy $pi_theta$), Critic (value function $V_psi$), Reference Model (frozen $pi_{text{SFT}}$), and Reward Model ($r_phi$), requiring complex multi-GPU distributed orchestration (DeepSpeed-Chat, Ray, vLLM). ④ Direct Preference Optimization (DPO) as an Alternative: DPO mathematically reparameterizes the reward model directly into the policy loss, eliminating the need to train a separate reward model or run PPO loops. ⑤ Interview Strategy: Diagram the three stages chronologically, write down the Bradley-Terry and PPO-KL optimization formulas, and explain the essential stabilizing role of the KL constraint.

五、常见面试避坑陷阱 (Common Pitfalls & Traps)

  • ⚠️ 跳过 SFT 直接做 PPO(采样质量差无法启动)
  • ⚠️ 去掉 KL 约束(导致奖励黑客与语言退化)

English Pitfalls:
– Attempting PPO optimization without a KL constraint (causes immediate reward hacking and language degradation)
– Skipping Stage 1 SFT and attempting to run RL on a raw base model
– Assuming reward model absolute scalar scores have meaningful units (only the relative difference $r(y_w) – r(y_l)$ is identifiable)

六、高频深度面试追问与预测 (Follow-Up Questions)

  1. 为什么需要 SFT 作为起点?
  2. Why is human pairwise comparison significantly more consistent than absolute scalar Likert-scale scoring?
  3. RLHF 的三阶段能否合并?
  4. How does DPO mathematically eliminate the need for an explicit Reward Model and PPO training loop?

七、知识图谱对齐 (Knowledge Graph Anchor)

  • 🔗 关联底层卡片:RLHF 人类反馈强化学习:Bradley-Terry 奖励模型、PPO 剪切目标与 KL 惩罚 (RLHF: Bradley-Terry Reward Modeling, PPO & KL Penalties)
  • 🗺️ 知识图谱模块:大语言模型全景图谱

🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)

本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。

👉 前往 TalentMe 交互式研读本题 (M5-030) →


Discover more from AirSOTA – Air School Of Thoughts AtoZ

Subscribe to get the latest posts sent to your email.