【AI 核心深度 M2-011】比较 L1、L2 与 ElasticNet 正则化的几何与效果差异。(Compare the Geometric and Functional Differences Between L1, L2, and ElasticNet Regularization)深度数理推导与工程落地解析

所属模块:M2 · 经典机器学习 (Classical Machine Learning) | 专题分类:正则化 (Regularization (L1 / L2)) | 难度等级:Easy

一、核心一句话结论 (One-Sentence Summary)

L1 稀疏(特征选择),L2 收缩(抗共线),ElasticNet 兼顾两者。

ADVERTISEMENT · 赞助推荐

L1 (Lasso) promotes sparsity by shrinking coefficients to exact zero on diamond corners; L2 (Ridge) shrinks coefficients smoothly without sparsity; ElasticNet combines both, grouping correlated features while maintaining sparsity.

二、核心考点要义 (Key Insights)

  • 📌 L1 在 0 处不可导 → 坐标下降/近端法
  • 📌 L2 有闭式解且可微
  • 📌 EN 在强相关特征组上更稳定

English Insights:
– Constraint Ball Geometry: L1 ball $|w|_1 le C$ has sharp vertices on axes where $w_i=0$; L2 ball $|w|_2^2 le C$ is smooth and spherical.
– L1 (Lasso): Automatic feature selection, sparse models; under collinearity, arbitrarily selects one feature from a correlated cluster.
– L2 (Ridge): Solves multicollinearity, distributes weights evenly among correlated features; no sparsity.
– ElasticNet: $L = text{Loss} + lambda_1 |w|_1 + frac{lambda_2}{2}|w|_2^2$; strictly convex, retains group selection property.

三、核心数学原理与机理推导 (Mathematical Principles & Derivation)

$$L1:lambda|w|_1,quad L2:lambda|w|_2^2,quad EN:lambda_1|w|_1+lambda_2|w|_2^2$$

几何视角最直观:约束形式 min‖y−Xw‖² s.t. ‖w‖₁≤t 的可行域是菱形(顶点在坐标轴上),L2 约束的可行域是圆(无顶点)。损失等高线(椭圆)与可行域首次相切处即最优解;菱形的尖角使切点极可能落在坐标轴上 → 某些 wⱼ=0(稀疏)。解析视角:L1 的最优性条件含次梯度 ∂|wⱼ|=[−1,1](当 wⱼ=0),这给出软阈值 wⱼ=sign(zⱼ)max(|zⱼ|−λ,0),即 |zⱼ|≤λ 的系数被精确置零;而 L2 的导数 2λwⱼ 在 0 处为 0,无法把系数推到零。

📖 查看英文严格数学推导 (English Mathematical Derivation)

Analysis of ElasticNet grouping effect: Let features $x_i$ and $x_j$ be highly correlated ($r = x_i^T x_j$). For strictly convex ElasticNet with $lambda_2 > 0$, Zou & Hastie (2005) proved that the difference in estimated coefficients satisfies: $|hat{beta}_i – hat{beta}_j| le frac{|y|_2}{lambda_2} sqrt{2(1 – r)}$. As correlation $r to 1$, $|hat{beta}_i – hat{beta}_j| to 0$, forcing correlated features to receive nearly identical coefficients. In pure Lasso ($lambda_2 = 0$), this bound does not exist, causing one coefficient to take the entire weight while the other is set to zero arbitrarily.

四、工业级落地权衡与工程考量 (Industrial Trade-offs)

实践选择依据:① L1(LASSO)——需要特征选择、高维稀疏(文本、基因)、要求可解释性;缺点是相关特征组中只随机保留一个(选择不稳定)、解路径不连续、p>n 时最多选出 n 个非零系数。② L2(Ridge)——特征都相关且有贡献、共线性严重、p>n 场景;缺点是不产生稀疏解,模型包含全部特征。③ ElasticNet——两者结合,既稀疏又在相关特征组上稳定(’分组效应’:相关特征倾向于同进同出);代价是需调两个超参(λ₁、λ₂),可用 CV 在二维网格搜索。④ 算法差异——L1 不可微,用坐标下降(最常用,因每步有解析解)、近端梯度(ISTA/FISTA)、或 ADMM;L2 有闭式解或可直接用梯度法。⑤ 实践中:若只关心预测精度,L2 通常足够;若需特征选择,L1 或 ElasticNet;若特征间相关性强且需稀疏,优先 ElasticNet。

⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)

When to choose which: (1) L1: High-dimensional sparse feature regimes (e.g. 1 million text n-grams or gene expressions) where 99% of features are noise. (2) L2: Dense numerical signals where all features contribute small predictive value and collinearity must be controlled. (3) ElasticNet: High-dimensional data with multicollinear feature clusters ($p gg n$), achieving both sparsity and group stability.

五、常见面试避坑陷阱 (Common Pitfalls & Traps)

  • ⚠️ 认为 L1 在相关特征上稳定(实际随机选一个)
  • ⚠️ 以为 L2 也能产生精确零(只收缩不置零)

English Pitfalls:
– Using pure Lasso when $p > n$ (Lasso can select at most $n$ non-zero features before saturating; ElasticNet can select all $p$ features).
– Forgetting to scale/normalize features prior to applying L1/L2 penalties.

六、高频深度面试追问与预测 (Follow-Up Questions)

  1. 高维稀疏场景选哪个?
  2. Why can Lasso select at most $n$ features when $p > n$, and how does ElasticNet overcome this limit?
  3. 为什么 L1 能产生精确 0?
  4. How does the L1/2 norm compare to L1, and why is it non-convex?

七、知识图谱对齐 (Knowledge Graph Anchor)

  • 🔗 关联底层卡片:L1 Lasso 与 L2 Ridge 正则化几何与拉普拉斯/高斯先验 (L1 Lasso & L2 Ridge Regularization Geometry & Priors)
  • 🗺️ 知识图谱模块:经典机器学习思维导图

🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)

本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。

👉 前往 TalentMe 交互式研读本题 (M2-011) →


Discover more from AirSOTA – Air School Of Thoughts AtoZ

Subscribe to get the latest posts sent to your email.