所属模块:
M2 · 经典机器学习 (Classical Machine Learning)| 专题分类:偏差-方差与模型选择 (Bias-Variance Tradeoff & Model Selection)| 难度等级:Medium
一、核心一句话结论 (One-Sentence Summary)
信息准则 = 拟合优度 − 复杂度惩罚;AIC/BIC 用解析惩罚近似 CV,BIC 惩罚更强且有一致性。
AIC and BIC penalize model complexity analytically based on likelihood; AIC asymptotically equates to Leave-One-Out CV, while BIC enforces selection consistency.
二、核心考点要义 (Key Insights)
- 📌 AIC 惩罚 2k,BIC 惩罚 k·log n(n>7 时更严)
- 📌 BIC 具有模型选择一致性(n→∞ 选出真模型)
English Insights:
– AIC: $2k – 2ln hat{L}$; optimizes predictive KL-divergence, asymptotically equivalent to LOO-CV
– BIC: $k ln n – 2ln hat{L}$; penalizes parameters more heavily for $n > 7$, consistent model identification
– Adjusted $R^2$: $1 – frac{(1-R^2)(n-1)}{n-k-1}$; degree-of-freedom correction restricted to linear regression
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$text{AIC}=2k-2loghat L,qquad text{BIC}=klog n-2loghat L$$
三者的构造与含义:① AIC = 2k − 2log L̂——衡量’拟合优度(对数似然)减去参数个数的 2 倍惩罚’;理论依据是 AIC 渐近无偏地估计了模型的KL 风险(预测分布与真实分布的差异),故目标偏向预测最优(可能选出比真模型更复杂的模型)。② BIC = k·log n − 2log L̂——惩罚更重(当 n>7 时 log n>2),且具有一致性:n→∞ 时以概率 1 选出真模型(若真模型在候选集中);故目标偏向识别真模型(更保守、更简洁)。③ 调整 R² = 1 − (1−R²)(n−1)/(n−k−1)——对 R² 做自由度校正,惩罚增加特征;但它只适用于线性回归且不适用于比较非嵌套模型(不同响应变量时不可比)。与交叉验证的关系:AIC/BIC 是 CV 的解析近似——AIC 渐近等价于留一交叉验证(LOO-CV),BIC 则近似于某种加权的 CV;它们的优势是只需一次拟合(无需 K 次重训),在大数据上成本极低;劣势是依赖渐近假设(大样本、真模型在候选集中、似然正确指定)。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Mathematical Formulations:
① Akaike Information Criterion (AIC): $text{AIC} = 2k – 2ln hat{L}$. Asymptotically estimates the expected Kullback-Leibler (KL) divergence between true data generating distribution $g(x)$ and candidate model $f(x|hat{theta})$. Stone (1977) proved that AIC is asymptotically equivalent to Leave-One-Out Cross-Validation (LOO-CV).
② Bayesian Information Criterion (BIC): $text{BIC} = k ln n – 2ln hat{L}$. Derived from Laplace approximation of the marginal log-likelihood $ln P(D|M)$. BIC is consistent: as $n to infty$, the probability of selecting the true model converges to 1 (if true model is in the candidate set).
③ Adjusted $R^2$: $bar{R}^2 = 1 – frac{text{SS}_{text{res}} / (n – k – 1)}{text{SS}_{text{tot}} / (n – 1)}$. Penalizes adding non-significant predictors but does not apply across non-nested or non-linear models.
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
实践选择与要点:① AIC vs BIC——若目标是预测(不假设真模型在候选集中),用 AIC;若目标是识别真模型(如变量选择、科学推断),用 BIC。实践中 BIC 更常用(倾向简洁模型)。② 何时用 CV 而非信息准则——(a) 小样本(渐近近似不准);(b) 模型非嵌套或包含正则化(似然不标准);(c) 目标指标不是似然(如 AUC、F1)——此时 CV 更直接可靠。③ 信息准则的适用条件——需极大似然估计、样本量足够、模型正确指定(若模型误设,AIC/BIC 的解读需谨慎);且只能比较同一响应变量的模型(比较不同变换的 y 无意义)。④ 与正则化的关系——Lasso 等路径上的 λ 选择不能用 AIC/BIC 直接做(因为自由度定义模糊),应改用 BIC 的推广(EBIC)、CV 或 稳定性选择。⑤ 自由度 k 的定义——对正则化模型 k 应是有效自由度(如岭回归的 tr(H)、Lasso 的非零系数个数),而非参数总数。⑥ 报告规范——报告多个准则(AIC/BIC/调整 R²/CV)并说明一致性;若结论冲突,应说明依据(预测 vs 解释)。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
Efficiency vs Assumptions: AIC and BIC require only a single training fit on the full dataset, making them computationally instant compared to $K$-fold cross-validation. However, they rely strictly on parametric likelihood assumptions. $K$-fold CV is non-parametric, robust to model mis-specification, and directly measures test error.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 用 AIC/BIC 比较不同响应变量变换的模型
- ⚠️ 把正则化模型的参数总数当作自由度 k
English Pitfalls:
– Comparing AIC or BIC scores across models fit using different sample sizes $n$ or different response variable transformations
– Assuming lower AIC implies the simpler model is true; AIC prioritizes predictive accuracy and tends to overfit compared to BIC
六、高频深度面试追问与预测 (Follow-Up Questions)
- 为什么 AIC 与 BIC 会给出不同结论?
- Why is AIC asymptotically equivalent to Leave-One-Out Cross-Validation (LOO-CV)?
- 信息准则 vs 交叉验证的取舍?
- Under what conditions will BIC consistently recover the true data-generating model while AIC overfits?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
偏差-方差分解权衡 (Bias-Variance Tradeoff) 与交叉验证(Bias-Variance Tradeoff & Cross-Validation Strategy) - 🗺️ 知识图谱模块:
经典机器学习思维导图
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。