所属模块:
M4 · 序列与 Transformer (Sequences & Transformers)| 专题分类:Seq2Seq 与注意力起源 (Seq2Seq & Attention Origins)| 难度等级:Medium
一、核心一句话结论 (One-Sentence Summary)
teacher forcing 训练时用真实前缀,推理时用模型自己的输出;训练-推理输入分布不一致导致误差累积(exposure bias)。
Teacher forcing trains models by feeding ground-truth historical prefixes; exposure bias arises because inference feeds self-generated tokens, causing errors to compound rapidly.
二、核心考点要义 (Key Insights)
- 📌 训练用真值前缀,推理用自生成前缀(分布不同)
- 📌 一步错误会改变后续输入分布,导致误差累积
- 📌 缓解:scheduled sampling、序列级训练、RL/最小风险训练
English Insights:
– Teacher Forcing: training step $t$ receives ground-truth target $y_{t-1}^$ as input; enables parallel training across all time steps
– Exposure Bias: during inference, the model receives its own potentially corrupted prediction $hat{y}_{t-1}$; out-of-distribution drift
– Compounding errors: a single erroneous token early in generation throws subsequent hidden states into unexplored regions*
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$text{train}: p(y_t|y^*{<t},x);qquad text{infer}: p(y_t|hat y$$},x);qquad text{dist mismatch
数学机理:teacher forcing 是自回归模型的标准训练方式:第 t 步的输入是真实的前缀 y{<t}(来自训练数据),而非模型自己的预测;这使得每一步的训练信号都是’干净’的、且所有时间步可并行计算(因为输入不依赖模型输出)。exposure bias(暴露偏差) 指训练与推理的输入分布不一致:训练时模型只见过’正确前缀’,推理时(第 t 步)输入是模型自己生成的 ŷ,一旦某步生成错误,后续步骤的输入就落入了训练时从未见过的分布区域——模型在这个分布外区域的预测质量无保证,可能继续出错,形成误差累积(error accumulation)。形式上,训练优化的是 Σt log p(yt|y*),两者条件分布不同。序列长度的影响——误差累积的概率随长度增长(每步有 ε 的概率偏离,T 步后偏离概率约 1−(1−ε)^T),故长序列上 exposure bias 更严重。缓解手段:(a) scheduled sampling(Bengio 等 2015)——训练时以概率 p 用模型自己的输出替代真值,p 随训练从 0 退火到某上界;(b) 序列级训练(最小风险训练 MRT、BLEU 作为目标);(c) 强化学习(用序列级奖励做 policy gradient,如 SCST);(d) RLHF/DPO(用偏好数据直接优化生成质量)。}),而推理时的目标是 Σ_t log p(ŷ_t|ŷ{<t
📖 查看英文严格数学推导 (English Mathematical Derivation)
Mathematical Formulations and Dynamics:
① Teacher Forcing Training Regime:
Target likelihood: $max_theta sum_{t=1}^T log P(y_t^* mid y_{1:t-1}^*, X; theta)$.
– Advantage: Clean gradients, rapid convergence, and full temporal parallelization (in Transformers, causal masking computes all $T$ positions in a single forward pass).
– Flaw: The network is never trained on its own mistakes; it always assumes past history is 100% grammatically and semantically perfect.
② Inference Reality (Free Running):
Generated sequence: $hat{y}_t sim P(cdot mid hat{y}_{1:t-1}, X; theta)$.
– Distributional Shift: Let $P_{text{data}}$ be the distribution of ground-truth prefixes, and $P_{text{model}}$ be the prefix distribution induced by sampling. In inference, inputs are drawn from $P_{text{model}} ne P_{text{data}}$. If the model generates an awkward word at step 5, it enters an unfamiliar state not seen during training, triggering cascading hallucination.
③ Mitigation: Scheduled Sampling (Bengio et al., 2015):
At step $t$, input token is selected via coin flip: $x_t = begin{cases} y_{t-1}^* & text{with probability } epsilon_i \ hat{y}_{t-1} & text{with probability } 1 – epsilon_i end{cases}$.
Decay schedule gradually anneals $epsilon_i$ from $1.0$ (pure teacher forcing) to $0$ (free running).
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
深度剖析与工程权衡:① 为什么不能完全用自回归训练——若一开始就用模型自己的输出,早期模型输出质量极差、梯度信号混乱,训练无法启动;故需要’从 teacher forcing 逐步过渡’(scheduled sampling 的退火)或’两阶段’(先 teacher forcing 预训练、再序列级微调)。② scheduled sampling 的副作用——它引入了训练/推理的另一种不一致(随机混合),且梯度估计有偏(每步依赖采样);实践中效果不稳定,故更常用’两阶段 + 序列级微调’。③ 现代 LLM 的立场——LLM 的预训练与 SFT 几乎全用 teacher forcing(因为规模足够大、数据足够多样,模型对’自己生成的分布’也见得多);exposure bias 主要通过 (a) 海量数据、(b) RLHF/DPO 的偏好对齐、(c) 推理时的解码策略(beam search、温度)缓解。这说明’exposure bias 的重要性随数据规模下降’。④ 与 teacher forcing 并行的效率优势——teacher forcing 使所有时间步可并行(输入已知),这是 Transformer 训练高效的前提;纯自回归训练(每步依赖上一步输出)无法并行。⑤ 与’off-policy/on-policy’的类比——teacher forcing 相当于 off-policy(用行为策略的数据),自回归推理是 on-policy(用自身策略);RLHF 的核心正是’在 on-policy 数据上优化’。⑥ 面试要点——被问’exposure bias’,应给出’训练用真值前缀、推理用自生成前缀 → 分布不一致 → 误差累积‘的因果链,并说明’长序列更严重’与’三阶段缓解方案’;能类比 off-policy/on-policy 是深度理解的标志。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
RL / DPO Alternative: In modern LLM training, models undergo Supervised Fine-Tuning (SFT) with teacher forcing, followed by Reinforcement Learning from Human Feedback (RLHF / DPO) where models generate full unconstrained trajectories, directly eliminating exposure bias.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 以为 teacher forcing 的并行性是免费的(代价是暴露偏差)
- ⚠️ 把 exposure bias 与’过拟合’混为一谈
English Pitfalls:
– Attempting to use Scheduled Sampling in standard Transformer training; feeding self-generated tokens breaks causal mask parallelization and forces slow sequential unrolling
– Equating exposure bias with simple overfitting; exposure bias is a distribution shift between training conditioning and inference conditioning
六、高频深度面试追问与预测 (Follow-Up Questions)
- 为什么 exposure bias 在长序列上更严重?
- Why is Scheduled Sampling computationally incompatible with parallel Transformer training?
- scheduled sampling 的退火策略怎么设计?
- How does Direct Preference Optimization (DPO) alleviate exposure bias during LLM post-training?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
从 Seq2Seq 到 Bahdanau 注意力:信息瓶颈与加性/乘性对齐(Seq2Seq to Bahdanau Attention: Additive & Dot-Product Alignment) - 🗺️ 知识图谱模块:
大语言模型全景图谱
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。