Normalization & Regularization Taxonomy: BatchNorm, LayerNorm, RMSNorm, L0/L1/L2 Weight Decay & Inverted Dropout Guide
Summary: Normalization and Regularization stabilize training dynamics and prevent overfitting. This 100% exhaustive guide covers feature scaling (Standardization, MinMax, RobustScaler), L0/L1/L2 convex relaxation geometry, Label Smoothing, Max-Norm constraints, Inverted Dropout mechanisms, and 5 Normalization paradigms (BatchNorm, LayerNorm, InstanceNorm, GroupNorm, RMSNorm) with rich SEO explanatory text and Pure Numpy implementations.
🧭 Knowledge Map & Architecture Graph
graph TD
subgraph A["1. Feature Scaling"]
A1["Scale Sensitive: KNN, K-Means, SVM, PCA, Neural Nets"]
A2["Scale Invariant: Tree Models (Decision Trees, XGBoost)"]
A3["StandardScaler vs MinMaxScaler vs RobustScaler"]
A1 --> A3
A2 --> A3
end
subgraph B["2. Loss Regularization"]
B1["L0 Norm: Non-convex NP-Hard count"]
B2["L1 Norm: Convex relaxation diamond contour → Sparsity"]
B3["L2 Norm: Weight Shrinkage"]
B4["Label Smoothing: Prevents over-confidence"]
B1 --> B2 --> B3 --> B4
end
subgraph C["3. Inverted Dropout"]
C1["Standard Dropout vs Inverted Dropout"]
C2["Max-Norm Weight Constraints"]
C1 --> C2
end
subgraph D["4. Normalization Paradigms"]
D1["BatchNorm: Mini-batch training vs Inference EMA"]
D2["LayerNorm vs RMSNorm in LLMs"]
D1 --> D2
end
A --> B --> C --> D
💡 Intuition: Normalization = “where do you compute mean/variance?”. BatchNorm aggregates across samples (CV-friendly, but batch < 8 makes statistics noisy and inference must switch to EMA — forgetting
model.eval()is a classic bug); LayerNorm aggregates within each sample (batch-independent — that’s why Transformers use it); RMSNorm = LayerNorm minus mean-centering (LLaMA/DeepSeek use it; the mean is empirically unnecessary, saving ~7–10% per layer). Regularization is geometry: L1’s diamond constraint hits coordinate axes (sparse, feature selection), L2’s circle shrinks weights uniformly but never to zero. Inverted Dropout moves the $(1-p)$ compensation into training (divide by $1-p$), so inference is zero-overhead.🎤 Quick Answer: “BN output is NOT always standard normal — γ, β are learnable and can undo the normalization; its real job is smoothing the loss surface. Example: ε=0.1, K=10 label smoothing turns 1.0 into 0.91 + 0.01 per class. Batch 2 + BN = noisy stats → use GroupNorm. L1 pushes $w=0.01$ by a constant step, L2 by $lambda w$ — only L1 reaches exactly zero.”
📚 Chapter 1: Pure Numpy Normalization Engine
Plain-language reading (full implementations in the zh version): all three norms share the template “(x − mean)/std × γ + β”; they differ only in the statistics axis — BN uses axis=(0,2,3), LN uses axis=-1, RMSNorm skips the mean and scales by sqrt(mean(x²)). The training branch also updates running_mean via EMA — that’s the inference-time statistic.
import numpy as np
class PureNumpyNormEngine:
@staticmethod
def batch_norm_forward(x: np.ndarray, gamma: np.ndarray, beta: np.ndarray, running_mean: np.ndarray, running_var: np.ndarray, is_training: bool = True):
pass
💡 Intuition: BN is a two-mode machine: train on batch statistics (and update EMA), inference on frozen EMA. The
is_trainingflag is exactly themodel.train()/model.eval()distinction interviewers love.🎤 Quick Answer: “BN inference uses EMA statistics (momentum 0.9) so single-sample predictions are deterministic; training uses batch stats — mixing the two (forgetting
model.eval()) silently degrades results. RMSNorm’s saving is literally one fewermeanreduction per layer.”
🧠 深入探索 TalentMe 全景技术图谱与备考路线
本文选自 TalentMe AI 技术专栏与高维职业罗盘。支持双模态 Obsidian 本地私域同步、艾宾浩斯智能复习与 IDE 内嵌 AI 导师模拟面试。