【AI 核心深度 M5-025】解释 SFT 与 RLHF/DPO 的分工。(Division of Labor Between SFT and RLHF/DPO)深度数理推导与工程落地解析

所属模块:M5 · NLP 与大语言模型 (NLP & Large Language Models) | 专题分类:指令微调与 SFT (Instruction Tuning & Supervised Fine-Tuning) | 难度等级:Hard

一、核心一句话结论 (One-Sentence Summary)

SFT 用演示教’格式与基本行为’(模仿);RLHF/DPO 用偏好教’质量与取舍’(优化),后者需前者提供起点。

ADVERTISEMENT · 赞助推荐

SFT establishes foundational instruction following, domain knowledge, and response formatting, whereas RLHF/DPO refines preference boundaries, suppresses low-probability tail errors, and aligns model outputs with human values.

二、核心考点要义 (Key Insights)

  • 📌 SFT:模仿演示(教’怎么说’),受演示质量上限约束
  • 📌 RLHF/DPO:优化偏好(教’说得多好’),可超越演示
  • 📌 顺序:预训练 → SFT → 偏好优化

English Insights:
– SFT role: builds the baseline capability to understand user prompts and respond in coherent, structured markdown; unlocks latent pre-training capabilities
– RLHF / DPO role: explores the solution space, optimizes relative preference nuances, suppresses hallucinations and toxic completions, and resolves edge-case ambiguity
– Complementary duality: SFT sets the high-quality starting policy $pi_{text{SFT}}$; RL cannot function without a competent SFT base to explore from

三、核心数学原理与机理推导 (Mathematical Principles & Derivation)

$$text{SFT}: maxlog p(y|x);qquad text{RLHF/DPO}: max mathbb{E}[r(x,y)] text{s.t.} text{KL}(pi|pi_{text{ref}})ledelta$$

数学机理:两者的目标函数不同。SFT 最大化演示数据的似然(模仿学习):L=−Σ log p(y|x);它学到的是’人类示范的样子’,故其上限是演示数据的质量(若演示都不好,学不出更好)。RLHF/DPO 优化偏好(而非模仿):给定(更好 y_w,更差 y_l)的偏好对,RLHF 用奖励模型 + PPO 最大化期望奖励(带 KL 约束),DPO 直接用偏好对做对比损失。关键差异:(a) 信号类型——SFT 是’应该这样写’(正例),偏好优化是’这样比那样好’(相对比较);(b) 信息效率——偏好比较更容易获得(人类比较两个回答比’写出完美回答’容易得多),且可表达 SFT 无法表达的’取舍’(如’更简洁但略不完整’vs’更完整但冗长’);(c) 上限——SFT 受演示质量限制,偏好优化可以超越演示(因为只需’判断哪个更好’,不需’演示最好’);(d) 行为——SFT 教格式与基本能力,偏好优化教质量、风格、安全、诚实(如’承认不知道’)。分工与顺序——(a) SFT 先行(教模型’会说人话’、能遵循指令),否则 RLHF 的起点太差(策略无法生成有意义的回答);(b) 偏好优化在后(提升质量与对齐);(c) 两者可迭代(SFT → DPO → 用新模型生成偏好数据 → 再 DPO,即 iterative/online DPO)。能否跳过 SFT——若基座模型已具备指令遵循能力(如经过指令预训练的模型),可直接做 DPO;但对纯预训练模型,直接 DPO 效果差(起点太弱)。实证——Tulu、Zephyr 等显示’SFT + DPO’在同等数据下优于’SFT only’与’RLHF’(成本更低、效果接近)。

📖 查看英文严格数学推导 (English Mathematical Derivation)

Mathematical Mechanism: 1. SFT as Maximum Likelihood Estimation (MLE): SFT maximizes data likelihood over positive demonstrations: $$max_theta mathbb{E}_{(x, y) sim mathcal{D}_{text{demo}}} left[ log pi_theta(y mid x) right]$$ MLE treats all demonstration tokens equally. It has no negative feedback mechanism: it cannot explicitly penalize toxic, hallucinated, or sub-optimal answers. 2. RLHF / DPO as Contrastive Optimization: Given prompt $x$ and pair $(y_w succ y_l)$ (winning vs losing response): $$max_theta mathbb{E}_{(x, y_w, y_l)} left[ log sigmaleft( beta log frac{pi_theta(y_w mid x)}{pi_{text{ref}}(y_w mid x)} – beta log frac{pi_theta(y_l mid x)}{pi_{text{ref}}(y_l mid x)} right) right]$$ RLHF/DPO simultaneously pushes up the likelihood of preferred completions while actively pushing down the likelihood of dispreferred completions, carving sharp decision boundaries.

四、工业级落地权衡与工程考量 (Industrial Trade-offs)

深度剖析与工程权衡:① ‘偏好比演示更容易获得’是核心经济性——写出’完美回答’需要专家(贵);而’比较两个回答哪个好’可由普通标注者或 AI 完成(便宜、快)。故偏好数据的规模通常远大于演示数据。② ‘可超越演示’的机制——偏好优化在回答空间上做搜索(通过策略采样 + 奖励/偏好信号),故可发现’比任何演示都好的回答’;这是 SFT 做不到的(SFT 只能在演示分布的支撑集内插值)。③ KL 约束的双重作用——RLHF/DPO 中的 KL 约束(或 DPO 的隐式参考模型)同时 (a) 防奖励黑客(不让策略偏离太远以钻奖励空子)、(b) 防灾难性遗忘(保持原能力)。④ ‘SFT 阶段该多强’的权衡——SFT 过强(多 epoch、高 lr)会导致遗忘与僵化,且可能’锁死’在演示分布(不利于后续偏好优化探索);故实践中 SFT 常’适度’(1~3 epoch),把质量提升留给偏好优化。⑤ 与’拒绝采样微调(RFT)’的关系——RFT 用模型自己生成、筛选正确的做 SFT;它是’SFT 与 RL 的中间形态’(用自生成数据扩大演示集),常用于推理模型。⑥ 面试要点——被问’SFT 和 RLHF 的分工’,应给出’目标函数(模仿 vs 优化偏好)+ 信号(正例 vs 相对比较)+ 上限(受演示限制 vs 可超越演示)+ 顺序(SFT 先、偏好优化后)‘,并说明’偏好数据更易获得’与’KL 约束的双重作用’;能指出’SFT 过强不利于后续探索’是深度理解的标志。

⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)

Deep Dive & Engineering Trade-offs: ① Why RL Cannot Replace SFT: If you initialize RLHF or DPO directly from a raw pre-trained base model, the action space (all vocabulary tokens across length $L$) is astronomically vast. The policy generates unformatted, rambling text that receives near-zero reward, resulting in random gradient noise and optimization failure. SFT provides the vital warm-start policy $pi_{text{ref}}$. ② Why SFT Cannot Replace RL: In complex tasks, human evaluators can easily pick the better of two answers ($y_w succ y_l$), but struggle to write a perfect reference answer from scratch. RLHF leverages pairwise comparison data to surpass the demonstration ability of average human annotators. ③ Mode Collapse vs Exploration: SFT suffers from exposure bias (teacher forcing). RL allows the model to sample its own generations autoregressively and receive feedback, learning error correction. ④ The Alignment Tax: Excessive RLHF/DPO can degrade creative writing diversity and multi-step reasoning (the alignment tax), requiring careful KL regularization against $pi_{text{SFT}}$. ⑤ Interview Strategy: Contrast MLE (positive-only) vs Contrastive RL (positive and negative), explain why SFT is the prerequisite warm-start for RL, and articulate the complementary strengths of both phases.

五、常见面试避坑陷阱 (Common Pitfalls & Traps)

  • ⚠️ 认为 SFT 数据够好就无需偏好优化
  • ⚠️ 跳过 SFT 直接对纯预训练模型做 DPO

English Pitfalls:
– Attempting to run RLHF or DPO directly on a raw base model without SFT warm-start
– Assuming SFT can learn negative constraints effectively (cross-entropy cannot explicitly penalize dispreferred answers without contrastive objectives)
– Pushing RL optimization so aggressively that the model suffers the ‘alignment tax’ (loss of reasoning and creative diversity)

六、高频深度面试追问与预测 (Follow-Up Questions)

  1. 为什么 SFT 不能超越演示质量?
  2. Why is human pairwise comparison significantly more reliable and cheaper to collect than expert demonstration writing?
  3. 什么情况下可以跳过 SFT 直接 DPO?
  4. What causes the ‘alignment tax’ during RLHF, and how does the KL divergence penalty mitigate it?

七、知识图谱对齐 (Knowledge Graph Anchor)

  • 🔗 关联底层卡片:指令微调 (SFT):Loss Masking 掩码、Data Packing 样本打包与灾难性遗忘 (Supervised Fine-Tuning: Loss Masking & Data Packing)
  • 🗺️ 知识图谱模块:大语言模型全景图谱

🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)

本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。

👉 前往 TalentMe 交互式研读本题 (M5-025) →


Discover more from AirSOTA – Air School Of Thoughts AtoZ

Subscribe to get the latest posts sent to your email.