所属模块:
M1 · 数学与统计基础 (Mathematics & Statistics Fundamentals)| 专题分类:估计理论 (MLE/MAP) (估计理论 (MLE/MAP))| 难度等级:Medium
一、核心一句话结论 (One-Sentence Summary)
Fisher 信息是对数似然关于参数的曲率(二阶导期望);CRB 给出无偏估计方差的下界。
Fisher Information measures the amount of information that an observable random variable carries about an unknown parameter; the Cramér-Rao Lower Bound proves that no unbiased estimator can have variance smaller than the inverse Fisher Information: $text{Var}(hat{theta}) ge I(theta)^{-1}$.
二、核心考点要义 (Key Insights)
- 📌 Fisher 信息可加(独立样本相加)
- 📌 达到 CRB 的估计量称为有效估计量(MLE 渐近有效)
English Insights:
– Score function: $S(theta) = nabla_theta log p(xmid theta)$; expectation of score is always zero: $E[S(theta)] = 0$.
– Fisher Information: $I(theta) = E[(S(theta))(S(theta))^T] = -E[nabla^2_theta log p(xmid theta)]$; expectation of negative Hessian.
– Cramér-Rao Lower Bound: For any unbiased estimator $hat{theta}$, $text{Cov}(hat{theta}) succeq I(theta)^{-1}$.
– Asymptotic Efficiency: MLE asymptotically achieves the CRLB: $sqrt{N}(hat{theta}_{text{MLE}} – theta_0) xrightarrow{d} mathcal{N}(0, I(theta_0)^{-1})$.
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$I(theta)=mathbb EBig[Big(frac{partiallog p}{partialtheta}Big)^2Big]=-mathbb EBig[frac{partial^2log p}{partialtheta^2}Big],qquad mathrm{Var}(hattheta)gefrac{1}{nI(theta)}$$
Fisher 信息量的两种等价定义:① 得分函数平方的期望 I(θ)=E[(∂log p/∂θ)²];② 对数似然二阶导的负期望 I(θ)=−E[∂²log p/∂θ²]。第二种形式给出直观解释:Fisher 信息是对数似然在真值附近的曲率——曲率越大(似然越尖锐),参数越容易被数据确定,信息量越大。可加性:对独立样本,I_n(θ)=n·I(θ)(信息相加),这直接导致标准误 ∝1/√n。Cramér-Rao 下界(CRB):对任何无偏估计量,Var(θ̂)≥1/(nI(θ))。它是无偏估计方差的理论下界,达到它的估计量称为有效估计量(efficient)。MLE 在正则条件下渐近达到 CRB(渐近有效),这是 MLE 最优性的核心依据。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Proof via Cauchy-Schwarz inequality: For 1D unbiased estimator $hat{theta}$ ($E[hat{theta}] = theta$). Differentiating with respect to $theta$: $frac{d}{dtheta} int hat{theta}(x) p(xmid theta)dx = int hat{theta}(x) frac{partial log p(xmid theta)}{partial theta} p(xmid theta)dx = E[hat{theta} cdot S(theta)] = 1$. Since $E[S(theta)]=0$, $text{Cov}(hat{theta}, S(theta)) = E[(hat{theta} – theta)S(theta)] = E[hat{theta} S] – theta E[S] = 1$. By Cauchy-Schwarz inequality: $text{Cov}(hat{theta}, S(theta))^2 le text{Var}(hat{theta}) text{Var}(S(theta))$. Since $text{Var}(S(theta)) = E[S^2] = I(theta)$, substituting gives: $1 le text{Var}(hat{theta}) I(theta) implies text{Var}(hat{theta}) ge frac{1}{I(theta)}$.
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
实践应用:① 标准误与置信区间——渐近地 Var(θ̂_MLE)≈1/(nI(θ̂)),故标准误可用观测信息矩阵(海森的负逆)估计:SE=√[(−H)⁻¹]ᵢᵢ;这是逻辑回归、GLM、混合模型输出标准误的标准做法(statsmodels 即此)。② 实验设计的指导——Fisher 信息给出’哪种实验设计能最大化参数信息’的判据:D-最优设计最大化 det(I(θ))、A-最优最小化 tr(I⁻¹),这是最优实验设计的理论基础。③ 与海森矩阵的关系——在 MLE 处,观测信息 −H(θ̂) 与 Fisher 信息 I(θ̂) 在大样本下趋于一致;故’海森 = 信息’,这解释了为什么牛顿法(用海森)在统计上与 Fisher 得分法(IRLS)一致。④ 与 KL 散度的关系——两个邻近分布的 KL 散度局部近似为 ½·ΔθᵀI(θ)Δθ,即 Fisher 信息是统计流形上的度量张量(Riemannian metric),这是信息几何的基础,也用于自然梯度(用 I⁻¹ 调整梯度方向,实现参数空间的’最速下降’)。⑤ 局限——CRB 只适用于无偏估计;有偏估计(如岭回归、收缩估计)可以突破 CRB(用偏差换方差)。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
Fisher Information is foundational to modern ML: (1) Natural Gradient Descent (Amari, 1998): Parameter updates $Delta theta = -I(theta)^{-1} nabla_theta mathcal{L}$ take steepest descent steps with respect to the Riemannian manifold of probability distributions (invariant to reparameterization). (2) Elastic Weight Consolidation (EWC): Mitigates catastrophic forgetting in continual learning by regularizing parameter drift using the diagonal Fisher information $F_i (theta_i – theta_{A, i}^*)^2$.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 把 CRB 当作所有估计量的下界(只适用于无偏估计)
- ⚠️ 忽略 Fisher 信息的可加性(导致对 √n 规律的困惑)
English Pitfalls:
– Applying CRLB to biased estimators without the derivative correction factor $(1 + b'(theta))^2 / I(theta)$.
– Confusing the observed Fisher Information (negative Hessian evaluated at $hat{theta}$) with expected Fisher Information.
六、高频深度面试追问与预测 (Follow-Up Questions)
- Fisher 信息与海森矩阵的关系?
- How does Natural Gradient Descent relate to TRPO (Trust Region Policy Optimization) and K-FAC?
- 哪些估计量能达到 CRB?
- Why does the Fisher Information metric tensor define a Riemannian geometry via the second-order expansion of KL divergence?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
极大似然估计 (MLE) 与极大后验估计 (MAP)(MLE, MAP & Bayesian Parameter Estimation) - 🗺️ 知识图谱模块:
经典机器学习思维导图
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。