所属模块:
M2 · 经典机器学习 (Classical Machine Learning)| 专题分类:线性回归 (Linear Regression)| 难度等级:Medium
一、核心一句话结论 (One-Sentence Summary)
给每个样本乘权重再最小化加权残差平方和;权重取方差的倒数时达到最小方差。
WLS weights each residual inversely by the error variance ($w_i = 1/sigma_i^2$), transforming heteroscedastic data into homoscedastic space to restore BLUE optimality.
二、核心考点要义 (Key Insights)
- 📌 权重 ωᵢ 通常取 1/σᵢ²(方差倒数)
- 📌 异方差下 OLS 仍无偏但非有效
English Insights:
– Heteroscedasticity problem: OLS estimates remain unbiased, but standard errors are biased and efficiency is lost
– WLS solution: solves $min_beta (y – Xbeta)^T W (y – Xbeta)$ with weight matrix $W = text{diag}(1/sigma_i^2)$
– Feasible GLS (FGLS): estimates variance function $hat{sigma}_i^2 = exp(z_i^T gamma)$ from OLS residuals when variance is unknown
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$min_wsum_i omega_i,(y_i-x_i^top w)^2,qquad hat w_{WLS}=(X^top WX)^{-1}X^top Wy$$
问题的背景:OLS 假设同方差(Var(εᵢ)=σ²),若违反则 OLS 仍无偏但不再有效(不是 BLUE),且标准误公式失效。WLS 的做法:对每个样本赋予权重 ωᵢ,最小化 Σωᵢ(yᵢ−xᵢᵀw)²,解为 ŵ_WLS=(XᵀWX)⁻¹XᵀWy。最优权重的推导:由广义最小二乘(GLS)理论,当 Var(ε)=σ²W⁻¹ 时最优权重 ωᵢ=1/σᵢ²——即方差越大的观测权重越小(因为它携带的信息更少)。这符合直觉:噪声大的观测应被’打折’。加权后的估计量达到 Cramér-Rao 下界(在已知方差结构的条件下有效)。实践中的两阶段法:σᵢ² 通常未知,可用 (a) 用 OLS 残差建模方差(如假设 Var(εᵢ)∝xᵢ 或 ∝xᵢ²,用残差平方回归估计);(b) 迭代重加权最小二乘(IRLS)。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Mathematical formulation: In ordinary regression $y = Xbeta + epsilon$, homoscedasticity assumes $text{Cov}(epsilon) = sigma^2 I$. When heteroscedasticity is present, $text{Cov}(epsilon) = Omega = text{diag}(sigma_1^2, dots, sigma_n^2)$.
Under Gauss-Markov, OLS is no longer the Best Linear Unbiased Estimator (BLUE). Multiply the system by $Omega^{-1/2}$: $Omega^{-1/2} y = Omega^{-1/2} Xbeta + Omega^{-1/2}epsilon$. The transformed disturbance has covariance $text{Cov}(Omega^{-1/2}epsilon) = Omega^{-1/2}OmegaOmega^{-1/2} = I$.
Applying OLS to the transformed system yields the WLS Estimator: $hat{beta}_{text{WLS}} = (X^T W X)^{-1} X^T W y$, where $W = Omega^{-1} = text{diag}(1/sigma_1^2, dots, 1/sigma_n^2)$.
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
实践要点与对比:① WLS vs 稳健标准误——两者解决不同问题:WLS 改变点估计(更有效),稳健标准误(Huber-White)只修正推断(不改变点估计);若只关心正确的 p 值与置信区间,稳健标准误更简单(无需建模方差结构);若关心估计效率(小样本下方差更小),WLS 更好。② 权重结构的假设——WLS 的收益依赖权重正确;若权重误设,反而可能比 OLS 更差。常见假设:分组数据的组方差 ∝1/组大小、泊松数据的方差 ∝均值(此时 WLS 等价于 Poisson 回归的 IRLS)。③ 与 GLM 的关系——WLS 是 GLM 的一个特例(Gaussian 分布 + identity 链接 + 已知方差结构);GLM 的 IRLS 算法每步就是一次 WLS。④ 应用场景——实验数据中不同测量的精度不同(如仪器精度)、聚合数据的组大小不同(组均值回归时权重 ∝ 组大小)、异方差明显的数据。⑤ 陷阱——用估计的权重后,标准误需调整(否则低估);且权重不能为 0 或负(会丢弃/翻转样本)。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
Practical application: If variance $sigma_i^2$ is unknown, estimate it via auxiliary regression of squared residuals $log(e_i^2)$ on covariates (Feasible Generalized Least Squares, FGLS). Alternatively, keep OLS point estimates and use Huber-White robust standard errors (sandwich estimator) to obtain valid confidence intervals without modifying coefficients.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 异方差下只用 OLS 点估计而不修正推断
- ⚠️ 权重误设时仍认为 WLS 必然优于 OLS
English Pitfalls:
– Believing heteroscedasticity biases OLS point estimates $hat{beta}$; it only destroys statistical efficiency and invalidates standard errors
– Mis-specifying the variance weighting model in FGLS, which can yield worse estimators than unweighted OLS
六、高频深度面试追问与预测 (Follow-Up Questions)
- 如何估计 σᵢ²?(残差平方的两阶段法)
- How does the Huber-White sandwich estimator correct standard errors without requiring WLS?
- WLS 与稳健标准误的取舍?
- What statistical tests can detect heteroscedasticity (e.g., Breusch-Pagan, White test)?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
线性回归 OLS 闭式解与 Gauss-Markov 定理(Linear Regression: OLS Normal Equation & Gauss-Markov) - 🗺️ 知识图谱模块:
经典机器学习思维导图
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。