所属模块:
M3 · 深度学习基础 (Deep Learning Foundations)| 专题分类:正则化与训练技巧 (Regularization & Training Tricks)| 难度等级:Hard
一、核心一句话结论 (One-Sentence Summary)
小模型/CV/微调用 dropout(0.1);大 LLM 预训练几乎不用(数据充足、显存开销大、与 LN 功能重叠)。
Dropout prevents co-adaptation in small-data regimes, but is universally set to 0 in large language model pre-training to preserve model capacity and maximize token throughput.
二、核心考点要义 (Key Insights)
- 📌 dropout 等价于子网络集成 + 隐式 L2 正则
- 📌 LLM 预训练通常只保留 attention dropout=0
- 📌 大模型靠数据规模与 weight decay 正则,而非 dropout
English Insights:
– Inverted Dropout: $y = frac{1}{1-p} m odot x$ where $m sim text{Bernoulli}(1-p)$; leaves inference untouched
– Standard Transformers (BERT/T5): uses $p = 0.1$ in attention probabilities, hidden projections, and embeddings
– Modern Foundation LLMs (LLaMA/GPT-3): sets dropout to $0.0$; massive web corpora make overfitting nonexistent during single-epoch pretraining
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$text{dropout}: h’=hodot m/(1-p), msimmathrm{Bernoulli}(1-p)$$
数学机理:dropout 在训练时以概率 p 随机置零激活、并以 1/(1−p) 缩放(inverted dropout,保证期望不变);推理时不丢弃。理论解释有二:(1) 集成视角——训练时相当于采样 2ⁿ 个子网络(n 为神经元数),推理时用全部神经元(近似子网络集成);(2) 正则视角——阻止神经元间’共适应’(co-adaptation),迫使每个单元独立有用。为什么大 LLM 不用:(a) 数据充足——LLM 预训练数据量远超参数量,过拟合不是主要矛盾,正则收益低;(b) 成本——dropout 需要额外随机数与掩码、破坏 kernel 融合、降低 GPU 利用率,在超大规模训练中代价显著;(c) 功能重叠——LayerNorm/残差连接已经提供了稳定训练的作用,dropout 的边际收益下降;(d) 与 MoE/长序列的冲突——随机丢弃会加剧训练-推理不一致。实证上,GPT-3/LLaMA 等仅保留极少量 dropout(常为 0)。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Formulation and System Impact:
① Inverted Dropout (Srivastava et al., 2014):
– Training: $y = frac{1}{1 – p} (m odot x)$, where $m_i sim text{Bernoulli}(1 – p)$. The scaling factor $frac{1}{1-p}$ ensures $mathbb{E}[y] = x$, eliminating the need to scale activations during inference.
– Ensemble Interpretation: Acts as an implicit geometric ensemble of $2^d$ sub-networks sharing weights.
② Why Foundation LLMs Abandon Dropout ($p=0$):
– Capacity Under-utilization: Pretraining models on 15+ trillion tokens is strictly under-fitting regime (single pass over data). Adding dropout artificially reduces effective model parameter capacity by $p%$.
– Compute & Bandwidth Overhead: Generating random PRNG masks and performing element-wise masking increases GPU memory traffic and prevents operator fusion in fused attention/MLP kernels.
– Fine-tuning: In low-resource supervised fine-tuning (SFT) or LoRA with small sample counts ($N < 5000$), dropout is reintroduced ($p=0.05-0.1$) to prevent memorization.
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
深度剖析与工程权衡:① 微调场景反转——在小数据微调(如几千条指令)时 dropout 仍有效,故很多 SFT 配置用 dropout=0.1;但 QLoRA 等高效微调常设 0(因为 LoRA 本身的低秩结构已提供正则)。② 注意力 dropout——部分模型保留 attention 权重的 dropout(防止注意力过度集中),但主流大模型也设为 0。③ Stochastic depth——ViT 等用随机跳过整层(drop path),是 dropout 的层级版本,在大视觉模型(ViT-H)中仍有效;LLM 中较少用。④ 与 weight decay 的分工——大模型的正则主要交给 weight decay 与’早停/数据规模’;dropout 退居次要。⑤ 历史教训——Transformer 原论文用 dropout=0.1;随着数据与规模增长,dropout 逐渐被移除(BERT 0.1 → GPT-3 0.0),这本身是’正则强度随数据规模下降’的经典案例。⑥ 面试要点——被问’为什么大模型不用 dropout’,应从数据规模、显存/算力成本、与 LN 的功能重叠三个角度回答,而非’因为大模型不需要正则’这种笼统说法。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
Sub-layer placement: When dropout is used in Transformers, it is applied at: (1) token embeddings, (2) attention weights after softmax, and (3) residual output projections. It is never applied inside LayerNorm/RMSNorm blocks.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 在大规模预训练上保留高 dropout(浪费算力且无收益)
- ⚠️ 忽略 dropout 会破坏 kernel 融合、降低吞吐
English Pitfalls:
– Leaving dropout enabled ($p=0.1$) during large-scale foundation pretraining, wasting millions of dollars in compute capacity
– Forgetting to call model.eval(), which leaves random dropout active during inference evaluation
六、高频深度面试追问与预测 (Follow-Up Questions)
- 为什么大模型不需要 dropout?
- Why is Inverted Dropout preferred over standard original dropout in modern deep learning libraries?
- dropout 与 LayerNorm 的功能重叠在哪里?
- Under what specific fine-tuning conditions should dropout be reintroduced in LLMs?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
深度学习正则化:Dropout、Weight Decay、DropPath 与EMA(DL Regularization: Dropout, Weight Decay, DropPath & EMA) - 🗺️ 知识图谱模块:
深度学习架构导图
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。