【AI 核心深度 M2-096】解释 L1 与 L2 正则化的贝叶斯解释(Bayesian Interpretation of L1 and L2 Regularization)深度数理推导与工程落地解析

所属模块:M2 · 经典机器学习 (Classical Machine Learning) | 专题分类:正则化 (Regularization (L1 / L2)) | 难度等级:Medium

一、核心一句话结论 (One-Sentence Summary)

L2 = 高斯先验下的 MAP;L1 = Laplace 先验下的 MAP;正则强度 = 先验方差的倒数。

ADVERTISEMENT · 赞助推荐

L2 regularization corresponds to a Gaussian prior (shrinkage); L1 regularization corresponds to an independent Laplace prior (sparsity via its non-differentiable peak at zero).

二、核心考点要义 (Key Insights)

  • 📌 λ = 1/(2τ²)(高斯)或 λ = 1/b(Laplace)
  • 📌 Laplace 在 0 处的尖峰 ⇒ 稀疏

English Insights:
– MAP framework: $argmax_w [log P(D|w) + log P(w)]$
– L2 / Ridge: Gaussian prior $w_j sim mathcal{N}(0, tau^2) implies -frac{1}{2tau^2} |w|_2^2$
– L1 / Lasso: Laplace prior $w_j sim text{Laplace}(0, b) implies -frac{1}{b} |w|_1$, whose sharp peak at zero forces coefficients exactly to zero

三、核心数学原理与机理推导 (Mathematical Principles & Derivation)

$$p(w)=mathcal N(0,tau^2I)Rightarrow text{L2};qquad p(w)propto e^{-|w|/b}Rightarrow text{L1}$$

推导:MAP 估计最大化 log p(D|w)+log p(w)。① 高斯先验 p(w)=N(0,τ²I) → log p(w)=−‖w‖²/(2τ²)+const → 目标变为 log 似然 − λ‖w‖²,其中 λ=1/(2τ²),即 L2 正则(岭回归/权重衰减);② Laplace 先验 p(w)∝exp(−|w|/b) → log p(w)=−‖w‖₁/b → 目标变为 log 似然 − λ‖w‖₁,即 L1 正则(LASSO)。为什么 Laplace 导致稀疏:Laplace 分布在 w=0 处有尖峰(不可导的角点),其对数先验在该点的’梯度’有跳变;当似然项对某系数的’推力’不够强时,后验众数恰好落在角点上(w=0)。相比之下高斯先验在 0 处平滑(导数为 0),后验众数只会被’收缩’但不会恰好为零。这给出了 L1 稀疏性的概率解释(与几何解释——菱形可行域的顶点——相互印证)。

📖 查看英文严格数学推导 (English Mathematical Derivation)

Maximum A Posteriori (MAP) Derivation: $hat{w}_{text{MAP}} = argmax_w left[ sum_{i=1}^n log P(y_i | x_i, w) + log P(w) right]$.
① Gaussian Prior $implies$ L2: If $w_j sim mathcal{N}(0, sigma_0^2)$, $log P(w) = -frac{1}{2sigma_0^2} sum_j w_j^2 – C$. The MAP objective is: $min_w mathcal{L}(w) + frac{sigma^2}{sigma_0^2} |w|_2^2$. Regularization weight $lambda = sigma^2 / sigma_0^2$.
② Laplace Prior $implies$ L1: If $w_j sim text{Laplace}(0, b)$, $P(w_j) = frac{1}{2b} expleft(-frac{|w_j|}{b}right)$, and $log P(w) = -frac{1}{b} |w|_1 – C$. The MAP objective is: $min_w mathcal{L}(w) + frac{sigma^2}{b} |w|_1$.
Why Laplace induces exact sparsity: The log-prior density $-frac{1}{b}|w_j|$ has a non-zero derivative $pm 1/b$ everywhere including as $w_j to 0$. When data likelihood gradient at zero $|nabla mathcal{L}| < 1/b$, the posterior mode is locked precisely at the non-differentiable corner $w_j = 0$.

四、工业级落地权衡与工程考量 (Industrial Trade-offs)

这一视角的实践价值:① 正则强度 = 先验信念强度——λ 大对应先验方差小(强信念’系数应接近 0’),λ 小对应弱先验趋近 MLE;这解释了为什么数据量增加时最优 λ 应减小(先验被数据稀释)。② 选择先验的依据——若认为’多数特征确实无关’(稀疏真值),用 Laplace(L1);若认为’所有特征都有小贡献’,用高斯(L2);若不确定,用 ElasticNet(混合先验)或层次先验(Horseshoe 先验,自适应稀疏)。③ 完整贝叶斯 vs MAP——MAP 只取后验众数(丢弃不确定性);完整贝叶斯需对后验积分(如贝叶斯线性回归的后验预测分布),能得到预测的不确定性估计,但计算更贵。④ 对超参的先验——λ 本身也可加先验(层次贝叶斯),用经验贝叶斯或 MCMC 自动确定,避免交叉验证的成本。⑤ 与深度学习的关系——weight decay 就是高斯先验(这解释了为什么它对所有权重同等收缩),而稀疏先验(L1、Horseshoe)在需要剪枝的场景更有用;此外 LoRA 的零初始化可理解为’强先验:增量应为 0’。

⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)

Posterior inference distinction: MAP gives a single point estimate (the posterior mode). In full Bayesian regression, the posterior mean under an L1 prior is actually non-sparse; true sparsity is a property of the mode (MAP), not the mean.

五、常见面试避坑陷阱 (Common Pitfalls & Traps)

  • ⚠️ 认为正则化只是’防止过拟合的技巧’(它有明确的先验解释)
  • ⚠️ 在高斯先验下期待稀疏解

English Pitfalls:
– Claiming that a full Bayesian posterior under a Laplace prior produces sparse sample paths, when only the MAP point estimate is sparse
– Assuming L1 and L2 priors imply different data likelihood models, rather than purely different prior distributions on $w$

六、高频深度面试追问与预测 (Follow-Up Questions)

  1. 为什么 Laplace 先验导致稀疏?
  2. Why is the posterior mean under a Laplace prior non-sparse while the MAP estimate is sparse?
  3. 这与’正则化就是加先验’如何统一?
  4. What prior distribution corresponds to ElasticNet regularization?

七、知识图谱对齐 (Knowledge Graph Anchor)

  • 🔗 关联底层卡片:L1 Lasso 与 L2 Ridge 正则化几何与拉普拉斯/高斯先验 (L1 Lasso & L2 Ridge Regularization Geometry & Priors)
  • 🗺️ 知识图谱模块:经典机器学习思维导图

🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)

本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。

👉 前往 TalentMe 交互式研读本题 (M2-096) →


Discover more from AirSOTA – Air School Of Thoughts AtoZ

Subscribe to get the latest posts sent to your email.