所属模块:
M5 · NLP 与大语言模型 (NLP & Large Language Models)| 专题分类:指令微调与 SFT (Instruction Tuning & Supervised Fine-Tuning)| 难度等级:Hard
一、核心一句话结论 (One-Sentence Summary)
SFT 用演示教’格式与基本行为’(模仿);RLHF/DPO 用偏好教’质量与取舍’(优化),后者需前者提供起点。
SFT establishes foundational instruction following, domain knowledge, and response formatting, whereas RLHF/DPO refines preference boundaries, suppresses low-probability tail errors, and aligns model outputs with human values.
二、核心考点要义 (Key Insights)
- 📌 SFT:模仿演示(教’怎么说’),受演示质量上限约束
- 📌 RLHF/DPO:优化偏好(教’说得多好’),可超越演示
- 📌 顺序:预训练 → SFT → 偏好优化
English Insights:
– SFT role: builds the baseline capability to understand user prompts and respond in coherent, structured markdown; unlocks latent pre-training capabilities
– RLHF / DPO role: explores the solution space, optimizes relative preference nuances, suppresses hallucinations and toxic completions, and resolves edge-case ambiguity
– Complementary duality: SFT sets the high-quality starting policy $pi_{text{SFT}}$; RL cannot function without a competent SFT base to explore from
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$text{SFT}: maxlog p(y|x);qquad text{RLHF/DPO}: max mathbb{E}[r(x,y)] text{s.t.} text{KL}(pi|pi_{text{ref}})ledelta$$
数学机理:两者的目标函数不同。SFT 最大化演示数据的似然(模仿学习):L=−Σ log p(y|x);它学到的是’人类示范的样子’,故其上限是演示数据的质量(若演示都不好,学不出更好)。RLHF/DPO 优化偏好(而非模仿):给定(更好 y_w,更差 y_l)的偏好对,RLHF 用奖励模型 + PPO 最大化期望奖励(带 KL 约束),DPO 直接用偏好对做对比损失。关键差异:(a) 信号类型——SFT 是’应该这样写’(正例),偏好优化是’这样比那样好’(相对比较);(b) 信息效率——偏好比较更容易获得(人类比较两个回答比’写出完美回答’容易得多),且可表达 SFT 无法表达的’取舍’(如’更简洁但略不完整’vs’更完整但冗长’);(c) 上限——SFT 受演示质量限制,偏好优化可以超越演示(因为只需’判断哪个更好’,不需’演示最好’);(d) 行为——SFT 教格式与基本能力,偏好优化教质量、风格、安全、诚实(如’承认不知道’)。分工与顺序——(a) SFT 先行(教模型’会说人话’、能遵循指令),否则 RLHF 的起点太差(策略无法生成有意义的回答);(b) 偏好优化在后(提升质量与对齐);(c) 两者可迭代(SFT → DPO → 用新模型生成偏好数据 → 再 DPO,即 iterative/online DPO)。能否跳过 SFT——若基座模型已具备指令遵循能力(如经过指令预训练的模型),可直接做 DPO;但对纯预训练模型,直接 DPO 效果差(起点太弱)。实证——Tulu、Zephyr 等显示’SFT + DPO’在同等数据下优于’SFT only’与’RLHF’(成本更低、效果接近)。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Mathematical Mechanism: 1. SFT as Maximum Likelihood Estimation (MLE): SFT maximizes data likelihood over positive demonstrations: $$max_theta mathbb{E}_{(x, y) sim mathcal{D}_{text{demo}}} left[ log pi_theta(y mid x) right]$$ MLE treats all demonstration tokens equally. It has no negative feedback mechanism: it cannot explicitly penalize toxic, hallucinated, or sub-optimal answers. 2. RLHF / DPO as Contrastive Optimization: Given prompt $x$ and pair $(y_w succ y_l)$ (winning vs losing response): $$max_theta mathbb{E}_{(x, y_w, y_l)} left[ log sigmaleft( beta log frac{pi_theta(y_w mid x)}{pi_{text{ref}}(y_w mid x)} – beta log frac{pi_theta(y_l mid x)}{pi_{text{ref}}(y_l mid x)} right) right]$$ RLHF/DPO simultaneously pushes up the likelihood of preferred completions while actively pushing down the likelihood of dispreferred completions, carving sharp decision boundaries.
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
深度剖析与工程权衡:① ‘偏好比演示更容易获得’是核心经济性——写出’完美回答’需要专家(贵);而’比较两个回答哪个好’可由普通标注者或 AI 完成(便宜、快)。故偏好数据的规模通常远大于演示数据。② ‘可超越演示’的机制——偏好优化在回答空间上做搜索(通过策略采样 + 奖励/偏好信号),故可发现’比任何演示都好的回答’;这是 SFT 做不到的(SFT 只能在演示分布的支撑集内插值)。③ KL 约束的双重作用——RLHF/DPO 中的 KL 约束(或 DPO 的隐式参考模型)同时 (a) 防奖励黑客(不让策略偏离太远以钻奖励空子)、(b) 防灾难性遗忘(保持原能力)。④ ‘SFT 阶段该多强’的权衡——SFT 过强(多 epoch、高 lr)会导致遗忘与僵化,且可能’锁死’在演示分布(不利于后续偏好优化探索);故实践中 SFT 常’适度’(1~3 epoch),把质量提升留给偏好优化。⑤ 与’拒绝采样微调(RFT)’的关系——RFT 用模型自己生成、筛选正确的做 SFT;它是’SFT 与 RL 的中间形态’(用自生成数据扩大演示集),常用于推理模型。⑥ 面试要点——被问’SFT 和 RLHF 的分工’,应给出’目标函数(模仿 vs 优化偏好)+ 信号(正例 vs 相对比较)+ 上限(受演示限制 vs 可超越演示)+ 顺序(SFT 先、偏好优化后)‘,并说明’偏好数据更易获得’与’KL 约束的双重作用’;能指出’SFT 过强不利于后续探索’是深度理解的标志。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
Deep Dive & Engineering Trade-offs: ① Why RL Cannot Replace SFT: If you initialize RLHF or DPO directly from a raw pre-trained base model, the action space (all vocabulary tokens across length $L$) is astronomically vast. The policy generates unformatted, rambling text that receives near-zero reward, resulting in random gradient noise and optimization failure. SFT provides the vital warm-start policy $pi_{text{ref}}$. ② Why SFT Cannot Replace RL: In complex tasks, human evaluators can easily pick the better of two answers ($y_w succ y_l$), but struggle to write a perfect reference answer from scratch. RLHF leverages pairwise comparison data to surpass the demonstration ability of average human annotators. ③ Mode Collapse vs Exploration: SFT suffers from exposure bias (teacher forcing). RL allows the model to sample its own generations autoregressively and receive feedback, learning error correction. ④ The Alignment Tax: Excessive RLHF/DPO can degrade creative writing diversity and multi-step reasoning (the alignment tax), requiring careful KL regularization against $pi_{text{SFT}}$. ⑤ Interview Strategy: Contrast MLE (positive-only) vs Contrastive RL (positive and negative), explain why SFT is the prerequisite warm-start for RL, and articulate the complementary strengths of both phases.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 认为 SFT 数据够好就无需偏好优化
- ⚠️ 跳过 SFT 直接对纯预训练模型做 DPO
English Pitfalls:
– Attempting to run RLHF or DPO directly on a raw base model without SFT warm-start
– Assuming SFT can learn negative constraints effectively (cross-entropy cannot explicitly penalize dispreferred answers without contrastive objectives)
– Pushing RL optimization so aggressively that the model suffers the ‘alignment tax’ (loss of reasoning and creative diversity)
六、高频深度面试追问与预测 (Follow-Up Questions)
- 为什么 SFT 不能超越演示质量?
- Why is human pairwise comparison significantly more reliable and cheaper to collect than expert demonstration writing?
- 什么情况下可以跳过 SFT 直接 DPO?
- What causes the ‘alignment tax’ during RLHF, and how does the KL divergence penalty mitigate it?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
指令微调 (SFT):Loss Masking 掩码、Data Packing 样本打包与灾难性遗忘(Supervised Fine-Tuning: Loss Masking & Data Packing) - 🗺️ 知识图谱模块:
大语言模型全景图谱
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。