所属模块:
M1 · 数学与统计基础 (Mathematics & Statistics Fundamentals)| 专题分类:假设检验 (Hypothesis Testing)| 难度等级:Easy
一、核心一句话结论 (One-Sentence Summary)
z 检验用已知总体方差(或大样本近似);t 检验用样本方差(小样本、方差未知)。
Use a z-test when population variance $sigma^2$ is known or sample size $N$ is very large ($N > 30$) by CLT; use a t-test when variance is unknown and estimated via sample variance $s^2$ in small-to-moderate samples.
二、核心考点要义 (Key Insights)
- 📌 t 分布比正态更重尾(反映方差估计的不确定性)
- 📌 n→∞ 时 t 分布收敛到正态
English Insights:
– z-test: Assumes known variance or invokes CLT for large samples; statistic $Z = frac{bar{X} – mu_0}{sigma / sqrt{n}} sim mathcal{N}(0, 1)$.
– t-test: Accounts for additional variance from estimating $s^2$; statistic $T = frac{bar{X} – mu_0}{s / sqrt{n}} sim t_{n-1}$.
– Heavier tails: Student’s t-distribution has fatter tails than Gaussian, producing more conservative critical values for small $N$.
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$t=frac{bar X-mu}{s/sqrt n}sim t_{n-1},qquad z=frac{bar X-mu}{sigma/sqrt n}simmathcal N(0,1)$$
选择的判据是’方差是否已知‘而非单纯的样本量:① z 检验——总体标准差 σ 已知(罕见),或样本量很大(n>30–50)时用 CLT 近似;② t 检验——σ 未知且用样本标准差 s 替代时,必须用 t 分布。为什么必须用 t:用 s 替代 σ 引入了额外的不确定性(s 本身是随机变量),若仍用正态分布会低估尾部概率(导致假阳性膨胀);t 分布正是为补偿这一不确定性而设计的——它的尾部比正态更厚,自由度 ν=n−1 越小尾部越厚,n→∞ 时收敛到正态。配对 vs 非配对:配对 t 检验用于同一批对象的两次测量(或用相同配对条件),检验差值的均值是否为零,能消除个体差异、大幅降低方差;非配对(独立样本)t 检验用于两组独立样本,需假设方差齐性(否则用 Welch 校正)。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Derivation of Student’s t-distribution: Let $X_1, dots, X_n sim mathcal{N}(mu, sigma^2)$. Then $bar{X} sim mathcal{N}(mu, sigma^2/n)$, and $(n-1)s^2/sigma^2 sim chi^2_{n-1}$, with $bar{X}$ and $s^2$ statistically independent by Cochran’s theorem. A Student’s t-variable is defined as the ratio of a standard normal to the square root of an independent normalized chi-square variable: $T = frac{(bar{X} – mu)/(sigma/sqrt{n})}{sqrt{((n-1)s^2/sigma^2)/(n-1)}} = frac{bar{X} – mu}{s/sqrt{n}} sim t_{n-1}$. As degrees of freedom $nu = n-1 to infty$, $t_nu xrightarrow{d} mathcal{N}(0, 1)$ by Slutsky’s theorem.
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
实践要点:① Welch t 检验——不假设两组方差相等(默认推荐),自由度用 Welch-Satterthwaite 近似;经典的学生 t 检验假设方差齐性,实践中该假设常不成立(尤其两组样本量差异大时),故 Welch 更稳健。② 样本量的经验阈值——’n>30 用 z’ 是常见经验,但若数据重偏或重尾,30 远不够;更稳妥的做法是统一用 t 检验(大样本下 t 与 z 几乎相同,无损失)。③ 配对设计的前提——配对要求两组观测一一对应(同一对象的前后测、同一批次的处理/对照);错用配对会低估方差导致假阳性。④ A/B 测试中的应用——两组用户独立,用非配对(Welch)t 检验;但若指标重尾,t 检验的 CLT 近似在中等样本下不可靠,应改用 bootstrap 或 Mann-Whitney U(秩和)检验。⑤ 多重比较——同时做多个 t 检验需 Bonferroni/BH 校正。⑥ 与 ANOVA 的关系——比较 3 组以上时不应做多次 t 检验(膨胀假阳性),应用 ANOVA + 事后检验。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
In industrial A/B testing where sample size $N$ is in the millions per variant, the $t$-distribution and normal distribution are numerically indistinguishable ($t_{1000000} approx mathcal{N}(0, 1)$). However, Welch’s t-test (which relaxes the equal variance assumption: $text{se} = sqrt{s_A^2/N_A + s_B^2/N_B}$ with Satterthwaite degrees of freedom) is the industry standard default because equal group variance is rarely guaranteed.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 用 z 检验而方差未知且样本小(应用 t)
- ⚠️ 比较三组以上时做多次 t 检验(应用 ANOVA)
English Pitfalls:
– Using Student’s t-test with pooled variance when group variances are unequal (violates homoscedasticity; Welch’s t-test must be used).
– Assuming t-tests are fully robust to extreme outliers in small sample sizes ($N < 20$).
六、高频深度面试追问与预测 (Follow-Up Questions)
- 为什么小样本必须用 t 检验?
- How does Welch’s t-test approximate degrees of freedom via the Welch-Satterthwaite equation?
- 配对 t 检验与非配对的区别?
- Why is Welch’s t-test preferred over Student’s pooled t-test even when sample sizes are equal?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
数理统计假说检验、P 值、I/II 类错误与统计功效(Hypothesis Testing, P-Values, Power & Type I/II Error) - 🗺️ 知识图谱模块:
数理基础思维导图
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。