【AI 核心深度 M1-042】定义第一类错误、第二类错误、显著性水平与功效。(Define Type I Error, Type II Error, Significance Level (Alpha), and Statistical Power (Beta))深度数理推导与工程落地解析

所属模块:M1 · 数学与统计基础 (Mathematics & Statistics Fundamentals) | 专题分类:假设检验 (Hypothesis Testing) | 难度等级:Easy

一、核心一句话结论 (One-Sentence Summary)

α = 弃真概率;β = 取伪概率;功效 = 1-β = 正确检出真实效应的概率。

ADVERTISEMENT · 赞助推荐

Type I error $alpha$ is false positive (rejecting true $H_0$); Type II error $beta$ is false negative (failing to reject false $H_0$); Power $1-beta$ is the probability of correctly detecting a real effect.

二、核心考点要义 (Key Insights)

  • 📌 功效随效应量、样本量、α 上升
  • 📌 常规要求 power ≥ 0.8

English Insights:
– Type I Error ($alpha$): $P(text{Reject } H_0 mid H_0 text{ is True})$; standard industrial threshold is $alpha = 0.05$.
– Type II Error ($beta$): $P(text{Fail to Reject } H_0 mid H_1 text{ is True})$; standard industrial threshold is $beta = 0.20$.
– Statistical Power ($1 – beta$): Probability of detecting a true effect; standard target is $80%$ or $90%$.
– Fundamental tradeoff: Lowering $alpha$ shifts critical threshold outward, automatically increasing $beta$ unless sample size $N$ is increased.

三、核心数学原理与机理推导 (Mathematical Principles & Derivation)

$$alpha=Pr(text{reject}mid H_0),quad text{power}=1-beta=Pr(text{reject}mid H_1)$$

两类错误的完整对照:第一类错误(弃真)——H₀ 为真却拒绝了它,概率为 α(显著性水平),这是’误报’;第二类错误(取伪)——H₁ 为真却未拒绝 H₀,概率为 β,这是’漏报’。功效 power=1−β 是正确检出真实效应的概率。四者的关系可写成 2×2 表:判决/真相的组合中,正确拒绝的概率是 power,正确不拒绝的概率是 1−α。关键性质:固定 n 时 α 与 β 反向变动(降低 α 会抬高 β);要同时降低两者只能增大样本量或效应量。

📖 查看英文严格数学推导 (English Mathematical Derivation)

For testing $H_0: mu = 0$ vs $H_1: mu = delta > 0$ with test statistic $Z = frac{bar{X}}{sigma/sqrt{N}}$: Rejection region is $Z > z_{1-alpha}$. Under $H_0$, $Z sim mathcal{N}(0, 1)$, giving Type I error $P(Z > z_{1-alpha} mid H_0) = alpha$. Under $H_1$, $Z sim mathcal{N}left(frac{delta}{sigma/sqrt{N}}, 1right)$. The Type II error is $beta = P(Z le z_{1-alpha} mid H_1) = Phileft(z_{1-alpha} – frac{delta}{sigma/sqrt{N}}right)$. Setting this to $z_beta$ yields the fundamental relationship: $z_{1-alpha} + z_{1-beta} = frac{delta}{sigma/sqrt{N}}$, linking significance level, power, effect size, and sample size.

四、工业级落地权衡与工程考量 (Industrial Trade-offs)

实践要点:① 功效分析的四个输入——显著性水平 α、期望功效(通常 0.8)、效应量(MDE,最小可检测效应)、方差 σ²;给定其中三个可解出第四个(通常是样本量 n)。A/B 测试上线前必须做功效分析,否则可能跑了一个注定检不出效果的实验。② 功效不足的严重后果——低功效实验不仅容易漏掉真实效应,还会使已检出的显著结果被高估(’winner’s curse’):在低功效下,只有效应量被随机高估的样本才能越过显著性阈值,故观测到的效应系统性偏大,这解释了为什么很多小样本研究的效应无法被大样本复现。③ 多重比较的放大——同时检验 m 个假设时,至少一次假阳性的概率为 1−(1−α)^m,m=20 时达 64%,故需 Bonferroni(控制 FWER)或 BH(控制 FDR)校正。

⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)

In online experimentation: (1) If launching a risky feature (e.g. Redesigning core checkout UI), Type I error is dangerous, so one might tighten $alpha = 0.01$. (2) If screening hundreds of candidate algorithms in early funnel stages, Type II error is costlier (missing a breakthrough), so one might tolerate $alpha = 0.10$ to boost power. The only way to decrease both $alpha$ and $beta$ simultaneously is to increase sample size $N$ or reduce metric variance via CUPED.

五、常见面试避坑陷阱 (Common Pitfalls & Traps)

  • ⚠️ 认为 α 与 β 可以同时任意降低(需增大样本量)
  • ⚠️ 忽略低功效导致的效应量高估(winner’s curse)

English Pitfalls:
– Concluding an experiment was successful simply because it had a high statistical power target (power is a pre-experiment planning metric, not a post-hoc evaluation).
– Calculating ‘post-hoc observed power’ using the observed sample effect size (which is mathematically redundant with the p-value and provides zero new information).

六、高频深度面试追问与预测 (Follow-Up Questions)

  1. 功效分析的四个输入是什么?
  2. Why is post-hoc power considered mathematically invalid and misleading in experimental analysis?
  3. 为什么功效不足会导致’显著的结果不可靠’?
  4. How does the Neyman-Pearson Lemma prove that likelihood ratio tests maximize power for simple hypotheses?

七、知识图谱对齐 (Knowledge Graph Anchor)

  • 🔗 关联底层卡片:数理统计假说检验、P 值、I/II 类错误与统计功效 (Hypothesis Testing, P-Values, Power & Type I/II Error)
  • 🗺️ 知识图谱模块:数理基础思维导图

🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)

本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。

👉 前往 TalentMe 交互式研读本题 (M1-042) →


Discover more from AirSOTA – Air School Of Thoughts AtoZ

Subscribe to get the latest posts sent to your email.