【AI 核心深度 M4-010】解释 teacher forcing 与 exposure bias(Teacher Forcing and Exposure Bias in Autoregressive Sequence Generation)深度数理推导与工程落地解析

所属模块:M4 · 序列与 Transformer (Sequences & Transformers) | 专题分类:Seq2Seq 与注意力起源 (Seq2Seq & Attention Origins) | 难度等级:Medium

一、核心一句话结论 (One-Sentence Summary)

teacher forcing 训练时用真实前缀,推理时用模型自己的输出;训练-推理输入分布不一致导致误差累积(exposure bias)。

ADVERTISEMENT · 赞助推荐

Teacher forcing trains models by feeding ground-truth historical prefixes; exposure bias arises because inference feeds self-generated tokens, causing errors to compound rapidly.

二、核心考点要义 (Key Insights)

  • 📌 训练用真值前缀,推理用自生成前缀(分布不同)
  • 📌 一步错误会改变后续输入分布,导致误差累积
  • 📌 缓解:scheduled sampling、序列级训练、RL/最小风险训练

English Insights:
– Teacher Forcing: training step $t$ receives ground-truth target $y_{t-1}^$ as input; enables parallel training across all time steps
–
Exposure Bias: during inference, the model receives its own potentially corrupted prediction $hat{y}_{t-1}$; out-of-distribution drift
–
Compounding errors: a single erroneous token early in generation throws subsequent hidden states into unexplored regions*

三、核心数学原理与机理推导 (Mathematical Principles & Derivation)

$$text{train}: p(y_t|y^*{<t},x);qquad text{infer}: p(y_t|hat y$$},x);qquad text{dist mismatch

数学机理:teacher forcing 是自回归模型的标准训练方式:第 t 步的输入是真实的前缀 y{<t}(来自训练数据),而非模型自己的预测;这使得每一步的训练信号都是’干净’的、且所有时间步可并行计算(因为输入不依赖模型输出)。exposure bias(暴露偏差) 指训练与推理的输入分布不一致:训练时模型只见过’正确前缀’,推理时(第 t 步)输入是模型自己生成的 ŷ,一旦某步生成错误,后续步骤的输入就落入了训练时从未见过的分布区域——模型在这个分布外区域的预测质量无保证,可能继续出错,形成误差累积(error accumulation)。形式上,训练优化的是 Σt log p(yt|y*),两者条件分布不同。序列长度的影响——误差累积的概率随长度增长(每步有 ε 的概率偏离,T 步后偏离概率约 1−(1−ε)^T),故长序列上 exposure bias 更严重。缓解手段:(a) scheduled sampling(Bengio 等 2015)——训练时以概率 p 用模型自己的输出替代真值,p 随训练从 0 退火到某上界;(b) 序列级训练(最小风险训练 MRT、BLEU 作为目标);(c) 强化学习(用序列级奖励做 policy gradient,如 SCST);(d) RLHF/DPO(用偏好数据直接优化生成质量)。}),而推理时的目标是 Σ_t log p(ŷ_t|ŷ{<t

📖 查看英文严格数学推导 (English Mathematical Derivation)

Mathematical Formulations and Dynamics:
① Teacher Forcing Training Regime:
Target likelihood: $max_theta sum_{t=1}^T log P(y_t^* mid y_{1:t-1}^*, X; theta)$.
– Advantage: Clean gradients, rapid convergence, and full temporal parallelization (in Transformers, causal masking computes all $T$ positions in a single forward pass).
– Flaw: The network is never trained on its own mistakes; it always assumes past history is 100% grammatically and semantically perfect.
② Inference Reality (Free Running):
Generated sequence: $hat{y}_t sim P(cdot mid hat{y}_{1:t-1}, X; theta)$.
– Distributional Shift: Let $P_{text{data}}$ be the distribution of ground-truth prefixes, and $P_{text{model}}$ be the prefix distribution induced by sampling. In inference, inputs are drawn from $P_{text{model}} ne P_{text{data}}$. If the model generates an awkward word at step 5, it enters an unfamiliar state not seen during training, triggering cascading hallucination.
③ Mitigation: Scheduled Sampling (Bengio et al., 2015):
At step $t$, input token is selected via coin flip: $x_t = begin{cases} y_{t-1}^* & text{with probability } epsilon_i \ hat{y}_{t-1} & text{with probability } 1 – epsilon_i end{cases}$.
Decay schedule gradually anneals $epsilon_i$ from $1.0$ (pure teacher forcing) to $0$ (free running).

四、工业级落地权衡与工程考量 (Industrial Trade-offs)

深度剖析与工程权衡:① 为什么不能完全用自回归训练——若一开始就用模型自己的输出,早期模型输出质量极差、梯度信号混乱,训练无法启动;故需要’从 teacher forcing 逐步过渡’(scheduled sampling 的退火)或’两阶段’(先 teacher forcing 预训练、再序列级微调)。② scheduled sampling 的副作用——它引入了训练/推理的另一种不一致(随机混合),且梯度估计有偏(每步依赖采样);实践中效果不稳定,故更常用’两阶段 + 序列级微调’。③ 现代 LLM 的立场——LLM 的预训练与 SFT 几乎全用 teacher forcing(因为规模足够大、数据足够多样,模型对’自己生成的分布’也见得多);exposure bias 主要通过 (a) 海量数据、(b) RLHF/DPO 的偏好对齐、(c) 推理时的解码策略(beam search、温度)缓解。这说明’exposure bias 的重要性随数据规模下降’。④ 与 teacher forcing 并行的效率优势——teacher forcing 使所有时间步可并行(输入已知),这是 Transformer 训练高效的前提;纯自回归训练(每步依赖上一步输出)无法并行。⑤ 与’off-policy/on-policy’的类比——teacher forcing 相当于 off-policy(用行为策略的数据),自回归推理是 on-policy(用自身策略);RLHF 的核心正是’在 on-policy 数据上优化’。⑥ 面试要点——被问’exposure bias’,应给出’训练用真值前缀、推理用自生成前缀 → 分布不一致 → 误差累积‘的因果链,并说明’长序列更严重’与’三阶段缓解方案’;能类比 off-policy/on-policy 是深度理解的标志。

⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)

RL / DPO Alternative: In modern LLM training, models undergo Supervised Fine-Tuning (SFT) with teacher forcing, followed by Reinforcement Learning from Human Feedback (RLHF / DPO) where models generate full unconstrained trajectories, directly eliminating exposure bias.

五、常见面试避坑陷阱 (Common Pitfalls & Traps)

  • ⚠️ 以为 teacher forcing 的并行性是免费的(代价是暴露偏差)
  • ⚠️ 把 exposure bias 与’过拟合’混为一谈

English Pitfalls:
– Attempting to use Scheduled Sampling in standard Transformer training; feeding self-generated tokens breaks causal mask parallelization and forces slow sequential unrolling
– Equating exposure bias with simple overfitting; exposure bias is a distribution shift between training conditioning and inference conditioning

六、高频深度面试追问与预测 (Follow-Up Questions)

  1. 为什么 exposure bias 在长序列上更严重?
  2. Why is Scheduled Sampling computationally incompatible with parallel Transformer training?
  3. scheduled sampling 的退火策略怎么设计?
  4. How does Direct Preference Optimization (DPO) alleviate exposure bias during LLM post-training?

七、知识图谱对齐 (Knowledge Graph Anchor)

  • 🔗 关联底层卡片:从 Seq2Seq 到 Bahdanau 注意力:信息瓶颈与加性/乘性对齐 (Seq2Seq to Bahdanau Attention: Additive & Dot-Product Alignment)
  • 🗺️ 知识图谱模块:大语言模型全景图谱

🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)

本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。

👉 前往 TalentMe 交互式研读本题 (M4-010) →


Discover more from AirSOTA – Air School Of Thoughts AtoZ

Subscribe to get the latest posts sent to your email.