【AI 核心深度 M7-099】解释 A/B 实验的样本量计算与最小可检测效应(MDE)(Explain Sample Size Determination, Statistical Power, and Minimum Detectable Effect (MDE) in A/B Testing)深度数理推导与工程落地解析

所属模块:M7 · 检索、排序与推荐系统 (Retrieval, Ranking & RecSys) | 专题分类:在线指标与实验 (Online Metrics & Guardrails) | 难度等级:Hard

一、核心一句话结论 (One-Sentence Summary)

样本量 ∝ 方差 / MDE²;MDE 是’能检出的最小效应’,需与业务意义平衡(太小则样本量爆炸)。

ADVERTISEMENT · 赞助推荐

Sample size scales proportionally with metric variance and inversely with the square of the Minimum Detectable Effect (n ~ sigma^2 / MDE^2); balancing MDE against business value prevents under-powered experiments and wasteful multi-month test durations.

二、核心考点要义 (Key Insights)

  • 📌 MDE:能检出的最小效应(如 +0.5% CTR)
  • 📌 样本量 ∝ σ²/MDE²(MDE 减半则样本量 ×4)
  • 📌 MDE 需与’业务意义’平衡(太小则样本量不可行)

English Insights:
– The sample size scaling law: Halving the target MDE quadruples required sample size (1 / MDE^2 scaling), demanding massive traffic investments for marginal gains.
– Statistical parameters: Type I error alpha (false positive rate, typically 5%) and Statistical Power 1 – beta (probability of detecting a true effect, typically 80%).
– Continuous vs. binary metric variance: Continuous revenue metrics have high skew and variance (requiring larger N); binary CTR metrics scale with p(1-p).
– CUPED sample size reduction: Reducing metric variance sigma^2 by 50% reduces required sample size by exactly 50% for the same target MDE.

三、核心数学原理与机理推导 (Mathematical Principles & Derivation)

$$napproxfrac{2(z_{1-alpha/2}+z_{1-beta})^2sigma^2}{text{MDE}^2};qquad text{MDE}downarrowRightarrow nuparrow text{(quadratic)}$$

数学机理:样本量计算的原理——(1) 目标——在给定’显著性水平 α’(如 0.05)与’统计功效 1−β’(如 0.8)下,检出’最小可检测效应(MDE)’所需的最少样本量。(2) 公式(两样本均值比较)——n ≈ 2(z_{1−α/2}+z_{1−β})²·σ²/MDE²,其中 (a) σ²——指标的方差(CTR 的方差 = p(1−p));(b) MDE——要检出的效应(绝对值,如 0.005 = 0.5%);(c) z——正态分位数(α=0.05 时 z≈1.96;功效 0.8 时 z≈0.84)。(3) 关键关系——样本量 ∝ σ²/MDE²:(a) MDE 减半 → 样本量 ×4(平方关系);(b) 方差减半 → 样本量减半(线性);(c) 功效提高(如 0.8→0.9)→ 样本量 ×1.3;(d) α 减小(如 0.05→0.01)→ 样本量 ×1.6。(4) MDE 的选择(关键权衡)——(a) 太小——样本量爆炸(如 MDE 0.1% 需要数百万样本、跑数周);(b) 太大——错过’小但重要’的改进(如 +0.5% 的 CTR 提升可能值数百万);(c) 实用做法——(i) 按’业务价值’定(’多大的提升值得上线’);(ii) 按’可行的实验时长’反推(如’一周能收集多少样本 → 能检出多大 MDE’);(iii) 参考历史实验的典型效应;(d) 注意——MDE 是’相对’还是’绝对’(CTR +0.5% 是相对还是绝对?需明确)。(5) 降方差以提高’有效样本量’——(a) CUPED(用实验前数据;常降 30%~50% 方差 → 相当于样本量 ×2~4);(b) 分层;(c) 协变量调整;(d) 更长的基线期(更稳的协变量);(e) 效果——降方差’等效于’增加样本量(但成本更低)。(6) 实际考量——(a) 流量约束(每天能分配多少用户);(b) 实验时长(= 所需样本量 / 日流量);(c) 周期效应(需覆盖完整周期,如’一周’以包含周末);(d) ‘新奇效应’(前几天的效应可能不代表长期);(e) ‘多指标’(每个指标都需要样本量;多重比较需校正)。(7) 与’最小可检测效应’的关系——(a) MDE 是’设计参数’(事前定);(b) 实际效应可能小于 MDE(则检不出);(c) ‘不显著’不等于’没效果’(可能是样本量不足)——这是重要的解释纪律。与其他问题的关系——(a) 与’统计显著性’(M5 的评估题);(b) 与’多重比较’(下一题);(c) 与’CUPED’(实验平台题)。实践建议——(a) 事前算样本量(用公式或平台工具);(b) MDE 按业务价值定(并考虑可行性);(c) 用 CUPED 降方差(等效增样本);(d) 覆盖完整周期(含周末);(e) ‘不显著’的解释要谨慎(可能是样本不足);(f) 多指标时校正。度量——(a) 所需样本量与实验时长;(b) 实际 MDE(事后);(c) CUPED 的方差降低比例;(d) 实验的检出力。

📖 查看英文严格数学推导 (English Mathematical Derivation)

Mathematical & Statistical Formulation: Sample Size Mechanics.

(1) Two-Sample Hypothesis Testing Framework:
Let null hypothesis be $H_0: mu_T – mu_C = 0$, and alternative be $H_1: mu_T – mu_C = delta = text{MDE}$.
– Significance Level $alpha$: $P(text{Reject } H_0 mid H_0 text{ is true}) = alpha$ (Standard: $alpha = 0.05 implies z_{1-alpha/2} = 1.96$).
– Statistical Power $1 – beta$: $P(text{Reject } H_0 mid H_1 text{ is true}) = 1 – beta$ (Standard: $1 – beta = 0.80 implies z_{1-beta} = 0.84$).

(2) Sample Size Formula per Variant ($n$):
For a continuous metric with standard deviation $sigma$ across both variants:
$$n approx frac{2 cdot (z_{1-alpha/2} + z_{1-beta})^2 cdot sigma^2}{delta^2} = frac{2 cdot (1.96 + 0.84)^2 cdot sigma^2}{text{MDE}^2} approx frac{15.68 cdot sigma^2}{text{MDE}^2}$$
For a binary proportion metric (e.g., Conversion Rate $p$):
$$sigma^2 = p(1 – p) implies n approx frac{15.68 cdot p(1 – p)}{text{MDE}^2}$$

(3) The Quadratic Penalty ($1 / text{MDE}^2$):
Suppose baseline conversion rate $p = 0.05$ ($5%$):
– To detect a relative $+10%$ lift ($text{MDE} = 0.005$):
$$n approx frac{15.68 times 0.05 times 0.95}{(0.005)^2} = frac{0.7448}{0.000025} approx 29,792 text{ users per variant}$$$$text{Total } N approx 60,000 text{ users (Feasible in 1 day)}.$$
– To detect a relative $+1%$ lift ($text{MDE} = 0.0005$):
$$n approx frac{0.7448}{(0.0005)^2} = frac{0.7448}{0.00000025} approx 2,979,200 text{ users per variant}$$$$text{Total } N approx 6,000,000 text{ users (Requires 100x more traffic!)}$$

四、工业级落地权衡与工程考量 (Industrial Trade-offs)

深度剖析与工程权衡:① ‘MDE 减半则样本量 ×4’是平方关系——这是最实用的量化直觉;面试中能给出是深度理解的标志。② ‘MDE 需与业务价值平衡’——太小则不可行、太大则错过改进。③ ‘CUPED 等效于增样本’——降方差是’免费的样本量’。④ ‘不显著 ≠ 没效果’——可能是样本不足;这是重要的解释纪律。⑤ ‘覆盖完整周期’——避免’周末效应’的偏置。⑥ 面试要点——被问’A/B 需要多少样本’,应给出’n ∝ σ²/MDE² + MDE 的业务权衡 + CUPED 降方差 + 完整周期 + 不显著的解释‘;能给出’平方关系’是深度理解的标志。

⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)

In-Depth Analysis & Engineering Trade-offs: ① The Under-Powered Experiment Trap ($1 – beta < 0.50$)—running an A/B test with insufficient sample size does not mean ‘it will only catch large wins’; it means true $+1%$ algorithmic improvements fail to reach statistical significance ($p > 0.05$) and are discarded, while random noise outliers that happen to breach $p < 0.05$ are shipped (the 'Winner's Curse'); power calculations must be executed before launching any test. ② CUPED as a sample size multiplier—because $n propto sigma^2$, applying CUPED to reduce variance by 40% drops required sample size from 100k to 60k users, shortening experiment duration from 14 days to 8 days. ③ Metric variance capping and log transformation—continuous revenue distributions contain extreme whale spenders ($>$10,000$), causing $sigma^2$ to explode; capping revenue at the 99th percentile or taking $ln(1 + text{spend})$ cuts $sigma^2$ by 3x, vastly reducing sample size requirements. ④ Minimum Detectable Effect (MDE) alignment with business value—setting MDE too small wastes testing capacity on trivial gains; setting MDE too large misses meaningful cumulative optimizations; MDE should be calibrated to the smallest lift that justifies engineering maintenance costs. ⑤ Fixed-horizon testing vs. Continuous peeking—evaluating $p$-values daily during an experiment inflates false positive rates; sample size formulas assume evaluation occurs strictly at pre-calculated horizon $N$. ⑥ Interview takeaway—write out the sample size formula $n approx frac{16 sigma^2}{text{MDE}^2}$, explain the quadratic $1/text{MDE}^2$ scaling law with a concrete calculation, define Type I/II errors, and explain how CUPED reduces sample size requirements.

五、常见面试避坑陷阱 (Common Pitfalls & Traps)

  • ⚠️ 事后解释’不显著’为’没效果’(可能样本不足)
  • ⚠️ MDE 设得过小导致实验无法完成

English Pitfalls:
– Halving the target MDE without recognizing that required traffic volume increases by a factor of 4, causing tests to run for months.
– Running under-powered experiments on small traffic slices, discarding genuine algorithmic improvements that fail to reach statistical significance.
– Calculating sample size on skewed continuous revenue metrics without outlier capping, underestimating required traffic by a factor of 10.

六、高频深度面试追问与预测 (Follow-Up Questions)

  1. 为什么 MDE 减半样本量要 ×4?
  2. Why does the ‘Winner’s Curse’ cause statistically significant results from under-powered experiments to drastically overestimate true effect sizes?
  3. 如何选 MDE?
  4. How does outlier percentile capping on skewed revenue metrics reduce required sample size without introducing unacceptable bias?

七、知识图谱对齐 (Knowledge Graph Anchor)

  • 🔗 关联底层卡片:在线推荐实验与业务指标:CTR、CVR、留存时长、网络溢出效应与 CUPED (Online Metrics & A/B Testing: CTR, CVR, CUPED & Spillover)
  • 🗺️ 知识图谱模块:数据科学与因果实验导图

🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)

本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。

👉 前往 TalentMe 交互式研读本题 (M7-099) →


Discover more from AirSOTA – Air School Of Thoughts AtoZ

Subscribe to get the latest posts sent to your email.