【AI 核心深度 M3-032】解释 Adam 的 ε 参数作用,以及它在大模型中的常见取值(Role of the Epsilon Parameter in Adam and Common Settings in Large Models)深度数理推导与工程落地解析

所属模块:M3 · 深度学习基础 (Deep Learning Foundations) | 专题分类:优化器 (Optimizers & Second-Order Methods) | 难度等级:Medium

一、核心一句话结论 (One-Sentence Summary)

ε 防止 v̂ 为 0 时除零,同时在下界约束步长上界(η/ε);LLM 常用 1e-8,某些低精度训练用 1e-6。

ADVERTISEMENT · 赞助推荐

Epsilon prevents division by zero and bounds the maximum effective step size ($eta/epsilon$); LLMs typically use $10^{-8}$, or $10^{-6}$ in lower-precision (FP16/BF16) training.

二、核心考点要义 (Key Insights)

  • 📌 ε 决定’梯度极小时的最大步长’,不是纯数值保护
  • 📌 ε 太大 → 退化为 SGD;ε 太小 → 小梯度参数步长爆炸
  • 📌 BF16 训练常放大 ε(如 1e-6)以吸收精度噪声

English Insights:
– Division protection: avoids division by zero when gradient second moment $hat{v}_t to 0$
– Maximum step-size bound: when $sqrt{hat{v}_t} ll epsilon$, effective step is bounded by $frac{eta}{epsilon} hat{m}_t$
– Precision compatibility: $10^{-8}$ in FP32/BF16; increase to $10^{-6}$ or $10^{-5}$ in FP16 to avoid underflow into unnormalized steps

三、核心数学原理与机理推导 (Mathematical Principles & Derivation)

$$frac{eta}{sqrt{hat v}+epsilon} Rightarrow text{step}lefrac{eta}{epsilon}quad(hat vto0)$$

数学机理:Adam 的更新为 η·m̂/(√v̂+ε)。当 √v̂≫ε 时,ε 可忽略、更新量级 ≈ η;当 √v̂≪ε(该参数梯度长期极小,如 embedding 中极少出现的 token)时,分母被 ε 主导,更新量级 ≈ η·m̂/ε——可能远大于 η,造成该参数步长失控。故 ε 的作用是给分母设下界 ε,从而给步长设上界 η/ε(在 |m̂|≤1 的近似下)。这也说明 ε 不是’防除零’这么简单,它实际定义了’梯度多小时的步长上限’。工程实现中常见两种写法:√(v̂)+ε(PyTorch 默认)与 √(v̂+ε)(原始论文),后者在 v̂ 极小时更平滑但等价性有细微差别;大模型中因 v̂ 通常远大于 ε,两者差异可忽略。

📖 查看英文严格数学推导 (English Mathematical Derivation)

Mathematical Mechanics:
The update term in Adam is: $Delta theta_t = – frac{eta}{sqrt{hat{v}_t} + epsilon} hat{m}_t$.
– When $sqrt{hat{v}_t} gg epsilon$, $epsilon$ has negligible influence, and update size is scale-invariant: $Delta theta_t approx – eta frac{hat{m}_t}{sqrt{hat{v}_t}} approx pm eta$.
– When $sqrt{hat{v}_t} ll epsilon$ (e.g., rarely activated embedding tokens or zero-gradient features), the denominator collapses to $epsilon$: $Delta theta_t approx – frac{eta}{epsilon} hat{m}_t$.
If $epsilon = 10^{-16}$, a small non-zero gradient could trigger an explosive update step of magnitude $10^{16} cdot eta$. Thus, $epsilon$ functions as an implicit cap bounding maximum parameter displacement.
Positioning of $epsilon$: $frac{1}{sqrt{hat{v}_t} + epsilon}$ vs $frac{1}{sqrt{hat{v}_t + epsilon^2}}$. PyTorch places $epsilon$ outside the square root (`eps=1e-8`).

四、工业级落地权衡与工程考量 (Industrial Trade-offs)

深度剖析与工程权衡:① ε 与有效学习率——在梯度尺度整体很小的层(如某些初始化下),ε 主导分母会让有效 lr 变成 η/ε,可能比预期大几个数量级;这就是为什么 μP 理论强调 ε 应与宽度解耦、并在大模型中用较大 ε。② BF16 下的必要性——BF16 的梯度噪声(约 1e-2 相对误差)远大于 FP32,若 ε=1e-8 则噪声可能主导分母;故 BF16/FP16 训练常把 ε 放大到 1e-6~1e-8 甚至更大,让噪声被 ε 吸收。③ 与 lr 的耦合——真正重要的是 ε/η 的比值:PaLM 等报告用 ε=1e-8 配 η 变化时,若 η 变小则 ε 相对变大、训练变慢;故某些实现固定 ε 与 lr 的比例。④ 8-bit Adam——量化 v 会引入额外误差,故 8-bit Adam 通常要求 ε 略大,且用’块级’缩放来降低量化误差。⑤ 实践建议——从 1e-8 起步;若训练不稳定(loss spike、参数爆炸)先试放大 ε 到 1e-6;若收敛过慢且确认梯度尺度正常,再试缩小 ε。⑥ 面试延伸——追问常落到’Adam 更新量级为何与梯度大小无关’:答案是 m̂/√v̂ 的比值性质,而 ε 正是这个归一化在小梯度区的’截断’。

⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)

Hyperparameter guidance: In large language model pre-training, setting $epsilon = 10^{-6}$ or $10^{-5}$ stabilizes training when using FP16 or BF16 mixed-precision by preventing sudden gradient spikes on rare tokens.

五、常见面试避坑陷阱 (Common Pitfalls & Traps)

  • ⚠️ 把 ε 当成无关紧要的数值保护(它实质定义了步长上界)
  • ⚠️ 在 BF16 下沿用 FP32 的小 ε(噪声被放大、训练不稳)

English Pitfalls:
– Treating $epsilon$ as an irrelevant mathematical artifact and setting it to machine epsilon ($10^{-16}$), triggering training explosions
– Setting $epsilon$ excessively large ($> 10^{-3}$), which degenerates Adam into standard unnormalized SGD

六、高频深度面试追问与预测 (Follow-Up Questions)

  1. 为什么 ε 与 lr 的比值会影响训练稳定性?
  2. How does the ratio $eta / epsilon$ determine the theoretical upper bound on parameter step size in Adam?
  3. ε 放在平方根内外有区别吗?
  4. Why is $epsilon = 10^{-6}$ preferred over $10^{-8}$ when training with FP16 mixed precision?

七、知识图谱对齐 (Knowledge Graph Anchor)

  • 🔗 关联底层卡片:一阶优化器家族:SGD 动量、AdamW、AdaFactor 与 Lion (First-Order Optimizers: Momentum, AdamW & Lion)
  • 🗺️ 知识图谱模块:深度学习架构导图

🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)

本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。

👉 前往 TalentMe 交互式研读本题 (M3-032) →


Discover more from AirSOTA – Air School Of Thoughts AtoZ

Subscribe to get the latest posts sent to your email.