所属模块:
M4 · 序列与 Transformer (Sequences & Transformers)| 专题分类:模型压缩与蒸馏 (Model Compression & Distillation)| 难度等级:Medium
一、核心一句话结论 (One-Sentence Summary)
off-policy 用教师生成的数据(分布是教师的);on-policy 用学生自己生成的数据(分布是学生的),后者缓解分布不匹配。
Off-policy distillation trains the student on static teacher-generated text distributions, while on-policy distillation trains the student on its own self-generated responses evaluated by teacher feedback, eliminating distribution shift during autoregressive deployment.
二、核心考点要义 (Key Insights)
- 📌 off-policy:用教师的输出(或固定数据集)训练学生
- 📌 on-policy:用学生自己的输出,由教师打分(或蒸馏)
- 📌 on-policy 缓解’训练-推理分布不匹配’(同 exposure bias)
English Insights:
– Off-policy distillation: student learns from pre-generated teacher corpora via standard supervised cross-entropy; fast and stable, but suffers from exposure bias and covariate shift
– On-policy distillation (SeqKD / MiniLLM): student generates completions from prompts, and teacher provides per-token feedback or reverse KL divergence ($D_{text{KL}}(P_{text{student}} | P_{text{teacher}})$)
– Distribution alignment: on-policy distillation forces the student to recover from its own generation mistakes, yielding significantly higher fluency and lower hallucination rates
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$text{off-policy}: xsimmathcal{D}{text{teacher}};qquad text{on-policy}: xsimpi$$}
数学机理:off-policy 蒸馏——学生在一个固定的数据集上学习(数据来自教师或原始语料),即训练分布是教师的分布 D_teacher。问题——(a) 分布不匹配:学生推理时会遇到’自己生成的数据分布’(含自己的错误),但训练时只见过教师/真值分布;这与 exposure bias 同源(训练用 teacher forcing、推理用自回归);(b) 学生可能在教师的’容易样本’上过拟合,而在’自己的错误状态’上无监督。on-policy 蒸馏——让学生自己生成序列(或部分序列),再由教师在这些’学生自己的状态’上给出软标签(或评分)。为什么更好——(a) 训练分布与推理分布一致(都来自学生),消除分布不匹配;(b) 教师纠正学生的错误——教师对’学生走错的状态’给出正确的分布,直接教学生’如何从错误中恢复’;(c) 缓解 exposure bias(见 M4 的相关题)。代价——(a) 成本高:需在线让学生生成(自回归、串行),再让教师前向(额外计算),无法用固定数据集缓存;(b) 训练不稳定:学生分布随训练变化,需持续采样;(c) 教师可能对学生的奇怪状态给出不可靠的标签(教师也没见过这些状态)。实例——(a) GKD(Generalized KD):在学生的自生成序列上蒸馏;(b) MiniLLM:用反向 KL 在学生的分布上蒸馏(避免教师分布覆盖学生不擅长的区域);(c) RLHF 中的 on-policy 采样:用当前策略生成、再用奖励模型/人类反馈评分(本质是 on-policy 的偏好蒸馏)。权衡——实践中常混合:先用 off-policy 冷启动(快速收敛),再用 on-policy 精调(对齐分布)。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Mathematical Mechanism: 1. Off-Policy Distillation (Forward KL / Supervised Learning): Training objective samples sequences $y$ from teacher distribution $P_T(y | x)$: $$mathcal{L}_{text{off-policy}} = mathbb{E}_{x sim mathcal{D}, y sim P_T} left[ -log P_S(y | x) right] = mathbb{E}_{x} left[ D_{text{KL}}(P_T , | , P_S) right]$$ Forward KL is mode-covering: the student must assign non-zero probability everywhere the teacher places mass. If the student has lower capacity, it spreads probability mass thinly, generating blurry or ungrammatical text. 2. On-Policy Distillation (Reverse KL / Policy Gradient): Training samples sequences $y$ from the student’s current policy $P_S(y | x)$: $$mathcal{L}_{text{on-policy}} = mathbb{E}_{x sim mathcal{D}, y sim P_S} left[ D_{text{KL}}(P_S , | , P_T) right] = sum_{t} P_S(y_t | y_{<t}, x) log frac{P_S(y_t | y_{<t}, x)}{P_T(y_t | y_{<t}, x)}$$ Reverse KL is mode-seeking: the student concentrates its capacity on the major modes of the teacher distribution where it is confident, avoiding generating nonsensical hallucinations in regions it cannot model.
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
深度剖析与工程权衡:① 与 RL 的 off-policy/on-policy 术语一致——off-policy 用’行为策略’的数据、on-policy 用’当前策略’的数据;蒸馏的这两种模式与 RL 完全对应(且 RLHF 就是 on-policy 蒸馏)。理解这一统一视角很有价值。② 反向 KL vs 前向 KL 的选择——off-policy 蒸馏常用前向 KL(KL(p_T‖p_S),让学生覆盖教师的所有模式);on-policy 蒸馏(如 MiniLLM)用反向 KL(KL(p_S‖p_T),让学生聚焦自己会遇到的区域,避免在教师的高概率但学生不擅长的区域浪费容量)。这是’模式覆盖 vs 模式聚焦’的经典权衡(见 M3 的 KL 题)。③ 成本-收益的权衡——on-policy 蒸馏的质量更好但成本高(需在线采样);对’一次性蒸馏到小模型’的场景,额外成本可接受(因为推理时的成本节省会长期累积)。④ 与’数据蒸馏’的关系——数据蒸馏(用教师生成训练数据)通常是 off-policy(数据固定);若用学生生成、教师筛选(如 rejection sampling / STaR),则是 on-policy。⑤ 与’自我改进’的关系——on-policy 蒸馏是’自我改进’(self-improvement)的基础范式:模型生成 → 筛选/评分 → 训练自己;这被广泛用于推理能力的提升(如 STaR、ReST、以及 RLHF 的迭代)。⑥ 面试要点——被问’蒸馏的数据选择’,应给出’off-policy(固定数据,便宜但有分布不匹配)vs on-policy(学生自生成 + 教师评分,贵但消除不匹配)‘,并联系到 exposure bias 与 RLHF;能提到’前向 KL vs 反向 KL 的选择’是深度理解的标志。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
Deep Dive & Engineering Trade-offs: ① Covariate Shift & Exposure Bias: In off-policy training, the student is conditioned strictly on ground-truth teacher prefixes ($y_{<t}^{text{teacher}}$). During real-world autoregressive inference, the student conditions on its own prior generations ($y_{<t}^{text{student}}$). Any small error at step $t$ sends the student out-of-distribution, cascading into complete generation collapse. On-policy training exposes the student to its own mistakes during training, teaching it error recovery. ② Computational Complexity: Off-policy training is simple standard gradient descent over static offline datasets. On-policy training requires active online token sampling from the student, followed by teacher forward passes and policy gradient / PPO-like updates, multiplying training compute by $3text{–}5times$. ③ Modern Industrial Applications: Training compact reasoning models (e.g., DeepSeek-R1-Distill, Qwen-2.5-Math) starts with off-policy pre-distillation on millions of teacher tokens, followed by on-policy reinforcement learning or on-policy distillation to sharpen output modes. ④ Optimization Stability: On-policy RL-based distillation suffers from high gradient variance; techniques like MiniLLM or Generalized Knowledge Distillation (GKD) stabilize optimization. ⑤ Interview Strategy: Contrast mode-covering (Forward KL) with mode-seeking (Reverse KL), explain exposure bias, and justify when to pay the higher computational cost of on-policy distillation.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 忽略 off-policy 蒸馏的分布不匹配问题
- ⚠️ on-policy 蒸馏时忽略成本与不稳定性
English Pitfalls:
– Assuming off-policy distillation on teacher text is equivalent to on-policy generation training (ignores covariate shift during autoregression)
– Using Reverse KL on a model without sufficient pre-training (leads to mode collapse into repetitive single phrases)
– Overlooking that on-policy distillation requires generating tokens dynamically during the training loop
六、高频深度面试追问与预测 (Follow-Up Questions)
- 为什么 on-policy 蒸馏更好?
- Why does Forward KL lead to mode-covering behavior while Reverse KL leads to mode-seeking behavior?
- on-policy 的代价是什么?
- How does Generalized Knowledge Distillation (GKD) interpolate between on-policy and off-policy data distributions?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
知识蒸馏 (Knowledge Distillation):温度超参、软标签损失与学生网络(Knowledge Distillation: Temperature Scaling & Soft Targets) - 🗺️ 知识图谱模块:
大语言模型全景图谱
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。