【AI 核心深度 M2-029】如何得到随机森林的特征重要度?它有什么缺陷。(Detail Mean Decrease Impurity (MDI) vs. Permutation Importance (MDA) in Random Forests and Their Known Biases)深度数理推导与工程落地解析

所属模块:M2 · 经典机器学习 (Classical Machine Learning) | 专题分类:集成方法 (Bagging/RF) (集成方法 (Bagging/RF)) | 难度等级:Medium

一、核心一句话结论 (One-Sentence Summary)

基于不纯度下降(MDI)或置换重要度(permutation);MDI 偏向高基数特征且受相关特征稀释。

ADVERTISEMENT · 赞助推荐

Feature importance is measured via Mean Decrease Impurity (MDI, summing tree split gains) and Permutation Feature Importance (MDA, measuring score drops after shuffling a column); MDI is severely biased toward high-cardinality features, while MDA is robust.

二、核心考点要义 (Key Insights)

  • 📌 相关特征会互相分走重要度
  • 📌 置换重要度更可靠但计算贵
  • 📌 更严谨可用 SHAP

English Insights:
– Mean Decrease Impurity (MDI / Gini Importance): Sums impurity reductions $Delta I$ achieved by feature $j$ across all split nodes in all trees.
– Permutation Feature Importance (MDA / Mean Decrease Accuracy): Shuffles feature $j$ values on out-of-bag/test data and measures validation accuracy drop.
– Flaws of MDI: Severely inflated for continuous or high-cardinality categorical variables; evaluates training fit rather than generalization.

三、核心数学原理与机理推导 (Mathematical Principles & Derivation)

$$text{MDI}j=sum$$}j}Deltatext{impurity

三种重要度:① MDI(Mean Decrease in Impurity)——累加该特征在所有树中作为分裂特征带来的不纯度下降(按样本数加权)。优点是免费(训练时顺便算);缺点是:(a) 偏向高基数/多取值特征(更多候选阈值 → 更可能被选中且增益看似更大);(b) 在训练集上计算故对过拟合特征给高分;(c) 相关特征互相稀释——若两个特征强相关,树会在它们之间随机选,导致两者重要度都被低估。② 置换重要度(Permutation Importance)——在验证集上随机打乱某特征的值,看性能下降多少;优点是不偏向高基数、反映真实预测贡献;缺点是 (a) 计算需 O(p) 次预测;(b) 相关特征仍被稀释(打乱一个,另一个仍提供信息故性能不降);(c) 若特征分布与真实推理分布不同会失真。③ SHAP——基于博弈论的 Shapley 值,理论上满足一致性、局部准确性等公理,能给出逐样本的贡献分解,且能处理相关特征(需指定特征依赖);代价是计算成本高(TreeSHAP 对树模型有高效精确算法)。

📖 查看英文严格数学推导 (English Mathematical Derivation)

Mathematical formulation: (1) MDI: $text{MDI}(x_j) = frac{1}{B}sum_{b=1}^B sum_{t in T_b : v(t)=j} p(t) Delta I(t)$, where $p(t) = N_t / N$ is node sample fraction. Because high-cardinality features provide many candidate thresholds, they have vastly more opportunities to reduce training impurity by chance, even if pure noise. (2) Permutation Importance (MDA): Let baseline OOB score be $S_{text{OOB}}$. Randomly shuffle column $j$ in the OOB set to form $tilde{X}^{(j)}$, breaking its relationship with target $y$. Evaluate shuffled score $tilde{S}_{text{OOB}}^{(j)}$. The importance is: $text{MDA}(x_j) = S_{text{OOB}} – tilde{S}_{text{OOB}}^{(j)}$. If $x_j$ is critical, shuffling drops score massively; if noise, score drop is $approx 0$.

四、工业级落地权衡与工程考量 (Industrial Trade-offs)

实践建议:① 优先用置换重要度 + SHAP——MDI 只适合快速粗筛;报告重要度时应在测试集上算置换重要度并多次重复取平均(减少随机性)。② 处理相关特征——用分组置换(把相关特征组一起打乱)、或先做聚类再把同组特征合并为一个特征;SHAP 的 feature_perturbation 参数取 interventional 与 tree_path_dependent 时,在相关特征下给出不同答案,需明确选择。③ 重要度 ≠ 因果——重要度只反映模型内的预测贡献,不能解释为因果效应;若需因果推断应用专门的因果方法(如 DML、因果森林)。④ 重要度稳定性——用多个随机种子重复训练并比较重要度排序,若不稳定说明特征间存在强相关或数据不足。

⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)

Handling Collinear Features: If two features $x_1$ and $x_2$ are strongly collinear (e.g. $rho = 0.99$), when $x_1$ is shuffled, the model can still access the identical information via unshuffled $x_2$. Thus, Permutation Importance shows both features as having low importance! In production attribution, SHAP (Shapley Additive Explanations) based on cooperative game theory or grouped permutation importance resolves this collinear masking.

五、常见面试避坑陷阱 (Common Pitfalls & Traps)

  • ⚠️ 用 MDI 的重要度做特征选择(偏向高基数 + 训练集上算)
  • ⚠️ 把重要度当作因果效应解读

English Pitfalls:
– Relying on default feature_importances_ (MDI) in scikit-learn for high-cardinality features (e.g. Random numerical IDs appear at the top of importance rankings).
– Computing Permutation Feature Importance on training data instead of test/OOB data.

六、高频深度面试追问与预测 (Follow-Up Questions)

  1. 如何正确评估相关特征的重要度?(分组置换/SHAP)
  2. Why does Strobl et al. (2007) prove that MDI is systematically biased toward variables with many categories?
  3. 为什么 MDI 偏向高基数特征?
  4. How do TreeSHAP algorithms compute exact Shapley values in polynomial time $O(T L D^2)$?

七、知识图谱对齐 (Knowledge Graph Anchor)

  • 🔗 关联底层卡片:Bagging 随机森林 (Random Forest) 与 Out-of-Bag (OOB) 评估 (Bagging, Random Forests & Out-of-Bag Evaluation)
  • 🗺️ 知识图谱模块:经典机器学习思维导图

🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)

本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。

👉 前往 TalentMe 交互式研读本题 (M2-029) →


Discover more from AirSOTA – Air School Of Thoughts AtoZ

Subscribe to get the latest posts sent to your email.