所属模块:
M5 · NLP 与大语言模型 (NLP & Large Language Models)| 专题分类:预训练目标与数据 (Pretraining Objectives & Data Curation)| 难度等级:Easy
一、核心一句话结论 (One-Sentence Summary)
CLM 预测下一个 token(因果 mask,可生成);MLM 随机掩码预测被掩位置(双向,不能直接生成)。
CLM trains autoregressive decoders to predict the next token via unidirectional causal masking for text generation, whereas MLM trains bidirectional encoders to predict randomly masked tokens for discriminative representation learning.
二、核心考点要义 (Key Insights)
- 📌 CLM:因果注意力,每 token 一个监督信号,天然可生成
- 📌 MLM:双向注意力,只监督被掩位置(约 15%),不能直接生成
- 📌 CLM 是当前 LLM 主流;MLM 主要用于编码器/表示学习
English Insights:
– CLM (Causal LM / GPT / LLaMA): objective $sum log P(x_t mid x_{<t})$; causal lower-triangular attention mask; natural alignment with autoregressive sequence generation
– MLM (Masked LM / BERT / RoBERTa): objective $sum_{m in M} log P(x_m mid x_{setminus M})$; bidirectional full attention; superior for discriminative representation and classification
– Why CLM won the LLM revolution: CLM models unify generation, in-context few-shot learning, and reasoning into a single scalable next-token prediction objective
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$text{CLM}: mathcal{L}=-sum_tlog p(x_t|x_{<t});qquad text{MLM}: mathcal{L}=-sum_{iinmathcal{M}}log p(x_i|x_{setminusmathcal{M}})$$
数学机理:CLM(Causal Language Modeling)——用因果(下三角)mask 的注意力,目标是最小化’预测下一个 token’的负对数似然:L=−Σt log p(x_t|x{<t})。特点:(a) 每个位置都是监督信号(N 个 token 提供 N 个预测任务),训练信号密集;(b) 天然可生成(自回归采样与训练目标一致);(c) 只需单向上下文(与生成场景一致);(d) 可复用 KV cache 做增量推理。MLM(Masked Language Modeling,BERT)——随机把约 15% 的 token 替换为 [MASK](其中 80% 替换、10% 随机替换、10% 保持,以缓解’预训练-微调不一致’),用双向注意力预测被掩位置的原始 token:L=−Σ{i∈M} log p(x_i|x{setminus M})。特点:(a) 双向上下文(每个位置可看左右),故表示更适合理解类任务(分类、NER、抽取);(b) 训练信号稀疏(只监督 15% 的位置);(c) 不能直接生成(因为训练时看到未来、且目标是’填空’而非’续写’);(d) 存在 预训练-微调不一致(预训练有 [MASK]、下游任务没有)。关键差异总结——(a) 注意力方向:CLM 因果、MLM 双向;(b) 监督密度:CLM 100%、MLM 15%;(c) 可用性:CLM 可生成、MLM 擅理解;(d) 规模效应:CLM 随规模持续提升(且可统一所有任务为生成),MLM 的收益在规模化上不如 CLM——这是 Decoder-only + CLM 成为主流的原因。其他目标——(a) Prefix-LM(前缀双向 + 后缀因果);(b) Span corruption(T5:掩码连续片段并预测,兼顾两者);(c) FIM(Fill-in-the-Middle)(代码场景,把中间段移到末尾预测,使 CLM 具备’填空’能力);(d) 去噪目标(BART)。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Mathematical Mechanism: 1. Causal Language Modeling (Autoregressive): Maximum likelihood over sequence $X = [x_1, dots, x_T]$: $$mathcal{L}_{text{CLM}}(theta) = -sum_{t=1}^T log P_theta(x_t mid x_1, x_2, dots, x_{t-1})$$ Implemented via lower-triangular causal attention mask $M_{ij} = -infty$ for $j > i$. Every token position contributes a supervised training signal, achieving $100%$ token supervision efficiency per forward pass. 2. Masked Language Modeling (Bidirectional): A random subset of tokens $M subset {1, dots, T}$ (typically $15%$) is replaced with `[MASK]` (80%), random token (10%), or unchanged (10%): $$mathcal{L}_{text{MLM}}(theta) = -sum_{m in M} log P_theta(x_m mid X_{setminus M})$$ The attention mask is fully bidirectional ($M_{ij} = 0$). Only $15%$ of tokens contribute to the loss function, requiring $approx 6times$ more tokens to achieve equivalent gradient updates compared to CLM.
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
深度剖析与工程权衡:① ‘监督密度’的量化价值——CLM 每 token 都是监督,MLM 只有 15%;若以’每个 token 的监督次数’衡量,CLM 的数据效率(按此口径)更高。但 MLM 的双向性带来更强的表示质量——这是’效率 vs 表示质量’的权衡。② MLM 的 80/10/10 技巧——直接用 [MASK] 会让模型学到’只在看到 [MASK] 时预测’,造成预训练-微调不一致;故 BERT 用 80% [MASK]、10% 随机词、10% 原词的混合。这是’缓解分布不匹配’的经典手法。③ CLM 的统一性优势——CLM 可把所有任务表达为’条件续写’(分类=续写标签、抽取=续写答案、翻译=续写译文),故一个模型可覆盖多任务,且 in-context learning 天然涌现。这是 Decoder-only 胜出的结构性原因。④ 与长上下文的关系——CLM 的因果性使 KV cache 增量推理可行(生成第 t 个 token 只需新 token 的 Q 与历史的 K/V);MLM 的双向性则无法做增量生成。⑤ FIM 的工程价值——代码补全需要’在中间插入’(而非只在末尾续写);FIM 把’前缀 + 后缀 + 中间’重排为’前缀 + 后缀 → 中间’,使 CLM 学会填空;这是’用重排让 CLM 支持新任务’的巧妙手法。⑥ 面试要点——被问’CLM vs MLM’,应给出’注意力方向 + 监督密度 + 是否可生成 + 规模效应‘四维对比,并说明’CLM 因统一性与可生成性成为主流’;能提到’FIM 让 CLM 支持填空’与’MLM 的 80/10/10’是深度理解的标志。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
Deep Dive & Engineering Trade-offs: ① Supervision Efficiency: CLM trains on every single token ($100%$ token efficiency); MLM trains on only $15%$ of tokens per sequence. MLM requires significantly more pre-training tokens to achieve the same total gradient update volume. ② In-Context Learning (ICL) Emergence: Autoregressive CLM models naturally learn to continue prompts conditionally, enabling zero-shot and few-shot in-context learning ($P(text{Answer} mid text{Context}, text{Question})$). MLM models lack autoregressive factorization, rendering them incapable of natural conversational generation or multi-step reasoning. ③ Prefix-LM / Encoder-Decoder Middle Ground: Models like T5 or GLM use prefix-LM (bidirectional attention on the prompt, causal attention on the completion). While strong at summarization, decoder-only CLM scales more cleanly in multi-turn chat and KV-cache serving. ④ Pre-training vs Downstream Disparity in MLM: The `[MASK]` token never appears in downstream real-world text, introducing a fundamental pre-training/inference distribution mismatch (partially mitigated by the 80/10/10 rule). ⑤ Interview Strategy: Contrast the loss objectives, compare token supervision efficiency ($100%$ vs $15%$), and explain why CLM’s autoregressive factorization is the foundational prerequisite for in-context learning and reasoning.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 以为 MLM 也能直接用于生成
- ⚠️ 忽略 MLM 的预训练-微调不一致问题
English Pitfalls:
– Assuming MLM is better for generation (bidirectional attention cannot generate coherent long text autoregressively without iterative non-causal Gibbs sampling)
– Overlooking that MLM only computes loss on the $15%$ masked tokens (much lower data efficiency per forward pass)
– Failing to explain how the causal mask in CLM enables full parallel training of all positions simultaneously
六、高频深度面试追问与预测 (Follow-Up Questions)
- 为什么 MLM 需要 mask token 而 CLM 不需要?
- Why does the autoregressive factorization of CLM naturally yield in-context few-shot learning capabilities?
- MLM 的训练信号密度如何?
- What is Prefix-LM, and why did the industry converge on pure decoder-only CLM rather than Encoder-Decoder architectures?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
大模型预训练:自回归因果语言建模 (CLM)、掩码建模与高质量数据配比(Pretraining Objectives: Causal LM & High-Quality Data Recipes) - 🗺️ 知识图谱模块:
大语言模型全景图谱
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。