【AI 核心深度 M1-029】解释 L1 为什么产生稀疏解,而 L2 不会。(Explain Geometrically and Algebraically Why L1 Regularization Promotes Sparsity While L2 Does Not)深度数理推导与工程落地解析

所属模块:M1 · 数学与统计基础 (Mathematics & Statistics Fundamentals) | 专题分类:凸优化与 KKT (Convex Optimization & KKT) | 难度等级:Medium

一、核心一句话结论 (One-Sentence Summary)

L1 的约束区域是菱形,其顶点在坐标轴上;最优解易落在顶点 → 部分系数恰为 0。L2 是圆,无顶点。

ADVERTISEMENT · 赞助推荐

Geometrically, L1 contour balls have sharp non-smooth corners on coordinate axes where loss contours touch first; algebraically, L1 subgradient produces a constant shrinkage threshold, setting small weights exactly to zero.

二、核心考点要义 (Key Insights)

  • 📌 L1 在 0 处不可导 → 次梯度/坐标下降/近端算法
  • 📌 L2 有闭式解(岭回归)
  • 📌 ElasticNet 兼顾稀疏与稳定性

English Insights:
– Geometric perspective: The L1 norm ball $|w|_1 le C$ is a diamond/cross-polytope with vertices on coordinate axes where $w_i=0$.
– Algebraic perspective: Soft-thresholding operator $S_lambda(w) = text{sign}(w)max(|w|-lambda, 0)$ truncates weights with magnitude $< lambda$ to exact zero.
– L2 (Ridge) shrinks weights proportionally by factor $frac{1}{1+lambda}$, asymptotically approaching zero without ever reaching it.

三、核心数学原理与机理推导 (Mathematical Principles & Derivation)

$$min|y-Xw|^2+lambda|w|_1quadtext{vs}quad min|y-Xw|^2+lambda|w|_2^2$$

两种理解方式:几何视角——约束形式 min‖y−Xw‖² s.t. ‖w‖₁≤t 的可行域是菱形(二维时四个顶点在坐标轴上),而 L2 约束的可行域是圆。损失函数的等高线(椭圆)与可行域首次相切处即最优解;由于菱形有尖角且尖角恰在坐标轴上,切点极可能落在尖角上,此时某些 wⱼ=0。解析视角——对 L1 项求次梯度,最优性条件为 0∈∇(loss)+λ∂‖w‖₁,其中 ∂|wⱼ|=[−1,1](当 wⱼ=0)或 {sign(wⱼ)}(当 wⱼ≠0)。这给出软阈值形式:wⱼ=sign(zⱼ)·max(|zⱼ|−λ,0),即只要原始梯度对应的解 |zⱼ|≤λ,该系数就被精确置零。L2 的导数 2λwⱼ 在 wⱼ=0 处为 0,无法把系数’推’到零。

📖 查看英文严格数学推导 (English Mathematical Derivation)

For an orthogonal design matrix, the 1D objective is $min_w frac{1}{2}(w – hat{w})^2 + lambda |w|$. The subdifferential is $w – hat{w} + lambda partial |w| = 0$, where $partial |w| = text{sign}(w)$ if $w ne 0$ and $[-1, 1]$ if $w=0$. Solving gives the soft-thresholding solution: $w^* = hat{w} – lambda$ if $hat{w} > lambda$; $w^* = hat{w} + lambda$ if $hat{w} < -lambda$; and $w^* = 0$ if $|hat{w}| le lambda$. In contrast, for L2 objective $min_w frac{1}{2}(w – hat{w})^2 + frac{lambda}{2}w^2$, setting derivative to zero gives $w^* = frac{hat{w}}{1+lambda}$, which is a smooth linear scaling that equals zero only when empirical estimate $hat{w}=0$.

四、工业级落地权衡与工程考量 (Industrial Trade-offs)

工程权衡:① L1 的稀疏性带来特征选择与可解释性,适合高维稀疏场景(文本、基因);但代价是在强相关特征组中只随机保留一个,选择不稳定(换个样本可能选中另一个),且解路径不连续。② L2 收缩但不置零,对共线特征组做’均摊收缩’,更稳定,且有闭式解(岭回归)计算更快。③ ElasticNet 结合两者:λ₁‖w‖₁+λ₂‖w‖²,既稀疏又对相关特征组稳定(’分组效应’)。④ 算法上 L1 不可微,需用近端梯度法(ISTA/FISTA)、坐标下降或 ADMM;L2 可直接求导用梯度法或闭式解。实践中,若只关心预测精度而不关心稀疏性,L2 通常够用;若需特征选择,L1 或 ElasticNet 更合适。

⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)

L1 (Lasso) performs automatic feature selection, producing highly interpretable models and drastically reducing memory during inference by pruning dead features. However, under high multicollinearity, L1 arbitrarily selects one feature from a correlated group and discards the rest. Elastic Net blends L1 and L2 penalties: $lambda_1 |w|_1 + lambda_2 |w|_2^2$, combining sparsity with group selection stability.

五、常见面试避坑陷阱 (Common Pitfalls & Traps)

  • ⚠️ 认为 L1 在相关特征上稳定(实际是随机选一个)
  • ⚠️ 以为 L2 也能产生精确零(它只收缩)

English Pitfalls:
– Using standard gradient descent on L1 objectives (the gradient at $w=0$ is undefined, causing oscillations around zero; proximal gradient must be used).
– Assuming L2 weight decay sets small weights to absolute zero.

六、高频深度面试追问与预测 (Follow-Up Questions)

  1. 为什么 L1 有闭式解吗?
  2. How does the Bayesian perspective interpret L1 as a Laplace prior and L2 as a Gaussian prior?
  3. 坐标下降为什么适合 L1?
  4. Why does L0 regularization produce the ultimate sparse solution, and why is it NP-hard to optimize?

七、知识图谱对齐 (Knowledge Graph Anchor)

  • 🔗 关联底层卡片:凸优化理论、对偶问题与 KKT 互补松弛条件 (Convex Optimization, Duality & KKT Conditions)
  • 🗺️ 知识图谱模块:数理基础思维导图

🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)

本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。

👉 前往 TalentMe 交互式研读本题 (M1-029) →


Discover more from AirSOTA – Air School Of Thoughts AtoZ

Subscribe to get the latest posts sent to your email.