所属模块:
M2 · 经典机器学习 (Classical Machine Learning)| 专题分类:模型校准 (Model Calibration)| 难度等级:Medium
一、核心一句话结论 (One-Sentence Summary)
在 logits 上除以温度 T,用验证集拟合 T;不改预测类别,只调整置信度。
A single scalar parameter $T > 0$ divides pre-softmax logits, softening or sharpening probabilities without changing the argmax prediction order or classification accuracy.
二、核心考点要义 (Key Insights)
- 📌 现代深度网络普遍过度自信(T>1)
- 📌 单参数、几乎不损精度
English Insights:
– Mechanism: $hat{p}_i = frac{exp(z_i / T)}{sum_j exp(z_j / T)}$ with learned scalar $T > 0$
– Accuracy preservation: strictly preserves logit ranking, ensuring zero change in top-1 classification accuracy
– Overconfidence alleviation: modern deep networks typically converge to $T > 1$, which softens overconfident extremes
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$hat p=mathrm{softmax}(z/T)$$
温度缩放的机制:深度网络的 softmax 输出为 p̂=softmax(z),其中 z 是 logits。温度缩放引入参数 T>0:p̂’=softmax(z/T)。当 T>1 时,分布变得更平坦(降低最大概率、抬高其他概率)→ 缓解过度自信;T<1 则更尖锐。关键性质:① 不改变预测类别——因为 argmax(z/T)=argmax(z)(除以正数不改变排序),故准确率不变(或几乎不变);② 单参数——只需在验证集上最小化 NLL 拟合一个 T,计算极便宜、不易过拟合;③ 改善 NLL 与 ECE——实验表明温度缩放能显著降低深度网络的 ECE(如从 10% 降到 2%)。为什么深网过度自信:与训练目标(交叉熵)在过参数化模型上的行为有关——模型会把 logits 的尺度推得很大以降低训练损失,而 NLL 对’自信但错误’的惩罚在训练集上被最小化后,泛化到测试集时表现为过度自信;此外现代训练技巧(BN、大 batch、强数据)都会加剧过度自信。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Formulation: Let $z = (z_1, dots, z_K)$ be the unnormalized logit vector produced by the penultimate layer. Temperature scaling modifies the softmax operator with temperature parameter $T > 0$: $hat{p}_k(x; T) = frac{exp(z_k / T)}{sum_{j=1}^K exp(z_j / T)}$.
– Entropy & Softness: When $T > 1$, logits are compressed toward 0, pushing softmax outputs toward uniform distribution (increasing entropy, mitigating overconfidence). When $T < 1$, logits expand, sharpening confident predictions.
– Preservation of Argmax: Because division by $T > 0$ is a strictly monotonic transformation, $argmax_k (z_k / T) = argmax_k z_k$. Top-1 accuracy and top-$k$ classification predictions remain completely unchanged.
– Optimization: All neural network weights are frozen; $T$ is optimized via Negative Log-Likelihood (NLL) on a held-out validation set using simple 1D gradient descent or L-BFGS: $min_T – sum_{i=1}^{N_{text{val}}} log hat{p}_{y_i}(x_i; T)$.
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
实践要点:① T 的拟合——在独立验证集上最小化 NLL(等价于最大化校准后的似然);可用 LBFGS 或简单网格搜索(T 通常落在 1–3)。② 与知识蒸馏的区别——蒸馏中的温度 T 用于软化教师的输出(让学生学到类间相似性),是训练时用的;校准的温度用于修正已训练模型的置信度,是后处理;两者数学形式相同但目的不同(蒸馏的 T 常取 3–20,校准的 T 通常 <3)。③ 与其他校准方法的对比——Platt scaling 在校准 sigmoid 后可能改变排序(多参数),温度缩放严格保持排序(单参数);Isotonic 更灵活但过拟合风险大;实践中温度缩放是深度网络的默认首选。④ 局限——温度缩放只做全局缩放(所有样本同一 T),若模型的过度自信程度随样本变化(如难样本更过度自信),需更精细的方法(如输入相关的温度、直方图分箱、或 Mixup 等训练时正则化,后者同时改善泛化与校准)。⑤ 分布漂移下失效——线上分布变化后 T 需重新拟合。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
Engineering advantages: Requires optimizing only 1 scalar parameter, making it virtually immune to overfitting even on small validation datasets. Extremely lightweight to deploy in production inference servers (simple logit division). Limitation: Cannot correct non-monotonic miscalibration.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 用训练集拟合温度(应使用独立验证集)
- ⚠️ 认为温度缩放能提升准确率(它只改校准)
English Pitfalls:
– Optimizing temperature $T$ on the training set, which drives $T to 0$ due to cross-entropy overfitting
– Expecting temperature scaling to fix misclassifications, when it only alters probability calibration without changing top-1 error
六、高频深度面试追问与预测 (Follow-Up Questions)
- 为什么深网会过度自信?
- Why is temperature scaling strictly incapable of altering a model’s top-1 accuracy or ROC-AUC?
- T 与知识蒸馏中的温度关系?
- How does vector scaling or matrix scaling extend temperature scaling, and what overfitting risks do they introduce?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
概率模型校准:Platt Scaling、保序回归与 ECE 指标(Probability Calibration: Platt Scaling, Isotonic & ECE) - 🗺️ 知识图谱模块:
经典机器学习思维导图
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。