所属模块:
M7 · 检索、排序与推荐系统 (Retrieval, Ranking & RecSys)| 专题分类:在线指标与实验 (Online Metrics & Guardrails)| 难度等级:Hard
一、核心一句话结论 (One-Sentence Summary)
样本量 ∝ 方差 / MDE²;MDE 是’能检出的最小效应’,需与业务意义平衡(太小则样本量爆炸)。
Sample size scales proportionally with metric variance and inversely with the square of the Minimum Detectable Effect (n ~ sigma^2 / MDE^2); balancing MDE against business value prevents under-powered experiments and wasteful multi-month test durations.
二、核心考点要义 (Key Insights)
- 📌 MDE:能检出的最小效应(如 +0.5% CTR)
- 📌 样本量 ∝ σ²/MDE²(MDE 减半则样本量 ×4)
- 📌 MDE 需与’业务意义’平衡(太小则样本量不可行)
English Insights:
– The sample size scaling law: Halving the target MDE quadruples required sample size (1 / MDE^2 scaling), demanding massive traffic investments for marginal gains.
– Statistical parameters: Type I error alpha (false positive rate, typically 5%) and Statistical Power 1 – beta (probability of detecting a true effect, typically 80%).
– Continuous vs. binary metric variance: Continuous revenue metrics have high skew and variance (requiring larger N); binary CTR metrics scale with p(1-p).
– CUPED sample size reduction: Reducing metric variance sigma^2 by 50% reduces required sample size by exactly 50% for the same target MDE.
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$napproxfrac{2(z_{1-alpha/2}+z_{1-beta})^2sigma^2}{text{MDE}^2};qquad text{MDE}downarrowRightarrow nuparrow text{(quadratic)}$$
数学机理:样本量计算的原理——(1) 目标——在给定’显著性水平 α’(如 0.05)与’统计功效 1−β’(如 0.8)下,检出’最小可检测效应(MDE)’所需的最少样本量。(2) 公式(两样本均值比较)——n ≈ 2(z_{1−α/2}+z_{1−β})²·σ²/MDE²,其中 (a) σ²——指标的方差(CTR 的方差 = p(1−p));(b) MDE——要检出的效应(绝对值,如 0.005 = 0.5%);(c) z——正态分位数(α=0.05 时 z≈1.96;功效 0.8 时 z≈0.84)。(3) 关键关系——样本量 ∝ σ²/MDE²:(a) MDE 减半 → 样本量 ×4(平方关系);(b) 方差减半 → 样本量减半(线性);(c) 功效提高(如 0.8→0.9)→ 样本量 ×1.3;(d) α 减小(如 0.05→0.01)→ 样本量 ×1.6。(4) MDE 的选择(关键权衡)——(a) 太小——样本量爆炸(如 MDE 0.1% 需要数百万样本、跑数周);(b) 太大——错过’小但重要’的改进(如 +0.5% 的 CTR 提升可能值数百万);(c) 实用做法——(i) 按’业务价值’定(’多大的提升值得上线’);(ii) 按’可行的实验时长’反推(如’一周能收集多少样本 → 能检出多大 MDE’);(iii) 参考历史实验的典型效应;(d) 注意——MDE 是’相对’还是’绝对’(CTR +0.5% 是相对还是绝对?需明确)。(5) 降方差以提高’有效样本量’——(a) CUPED(用实验前数据;常降 30%~50% 方差 → 相当于样本量 ×2~4);(b) 分层;(c) 协变量调整;(d) 更长的基线期(更稳的协变量);(e) 效果——降方差’等效于’增加样本量(但成本更低)。(6) 实际考量——(a) 流量约束(每天能分配多少用户);(b) 实验时长(= 所需样本量 / 日流量);(c) 周期效应(需覆盖完整周期,如’一周’以包含周末);(d) ‘新奇效应’(前几天的效应可能不代表长期);(e) ‘多指标’(每个指标都需要样本量;多重比较需校正)。(7) 与’最小可检测效应’的关系——(a) MDE 是’设计参数’(事前定);(b) 实际效应可能小于 MDE(则检不出);(c) ‘不显著’不等于’没效果’(可能是样本量不足)——这是重要的解释纪律。与其他问题的关系——(a) 与’统计显著性’(M5 的评估题);(b) 与’多重比较’(下一题);(c) 与’CUPED’(实验平台题)。实践建议——(a) 事前算样本量(用公式或平台工具);(b) MDE 按业务价值定(并考虑可行性);(c) 用 CUPED 降方差(等效增样本);(d) 覆盖完整周期(含周末);(e) ‘不显著’的解释要谨慎(可能是样本不足);(f) 多指标时校正。度量——(a) 所需样本量与实验时长;(b) 实际 MDE(事后);(c) CUPED 的方差降低比例;(d) 实验的检出力。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Mathematical & Statistical Formulation: Sample Size Mechanics.
(1) Two-Sample Hypothesis Testing Framework:
Let null hypothesis be $H_0: mu_T – mu_C = 0$, and alternative be $H_1: mu_T – mu_C = delta = text{MDE}$.
– Significance Level $alpha$: $P(text{Reject } H_0 mid H_0 text{ is true}) = alpha$ (Standard: $alpha = 0.05 implies z_{1-alpha/2} = 1.96$).
– Statistical Power $1 – beta$: $P(text{Reject } H_0 mid H_1 text{ is true}) = 1 – beta$ (Standard: $1 – beta = 0.80 implies z_{1-beta} = 0.84$).
(2) Sample Size Formula per Variant ($n$):
For a continuous metric with standard deviation $sigma$ across both variants:
$$n approx frac{2 cdot (z_{1-alpha/2} + z_{1-beta})^2 cdot sigma^2}{delta^2} = frac{2 cdot (1.96 + 0.84)^2 cdot sigma^2}{text{MDE}^2} approx frac{15.68 cdot sigma^2}{text{MDE}^2}$$
For a binary proportion metric (e.g., Conversion Rate $p$):
$$sigma^2 = p(1 – p) implies n approx frac{15.68 cdot p(1 – p)}{text{MDE}^2}$$
(3) The Quadratic Penalty ($1 / text{MDE}^2$):
Suppose baseline conversion rate $p = 0.05$ ($5%$):
– To detect a relative $+10%$ lift ($text{MDE} = 0.005$):
$$n approx frac{15.68 times 0.05 times 0.95}{(0.005)^2} = frac{0.7448}{0.000025} approx 29,792 text{ users per variant}$$$$text{Total } N approx 60,000 text{ users (Feasible in 1 day)}.$$
– To detect a relative $+1%$ lift ($text{MDE} = 0.0005$):
$$n approx frac{0.7448}{(0.0005)^2} = frac{0.7448}{0.00000025} approx 2,979,200 text{ users per variant}$$$$text{Total } N approx 6,000,000 text{ users (Requires 100x more traffic!)}$$
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
深度剖析与工程权衡:① ‘MDE 减半则样本量 ×4’是平方关系——这是最实用的量化直觉;面试中能给出是深度理解的标志。② ‘MDE 需与业务价值平衡’——太小则不可行、太大则错过改进。③ ‘CUPED 等效于增样本’——降方差是’免费的样本量’。④ ‘不显著 ≠ 没效果’——可能是样本不足;这是重要的解释纪律。⑤ ‘覆盖完整周期’——避免’周末效应’的偏置。⑥ 面试要点——被问’A/B 需要多少样本’,应给出’n ∝ σ²/MDE² + MDE 的业务权衡 + CUPED 降方差 + 完整周期 + 不显著的解释‘;能给出’平方关系’是深度理解的标志。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
In-Depth Analysis & Engineering Trade-offs: ① The Under-Powered Experiment Trap ($1 – beta < 0.50$)—running an A/B test with insufficient sample size does not mean ‘it will only catch large wins’; it means true $+1%$ algorithmic improvements fail to reach statistical significance ($p > 0.05$) and are discarded, while random noise outliers that happen to breach $p < 0.05$ are shipped (the 'Winner's Curse'); power calculations must be executed before launching any test. ② CUPED as a sample size multiplier—because $n propto sigma^2$, applying CUPED to reduce variance by 40% drops required sample size from 100k to 60k users, shortening experiment duration from 14 days to 8 days. ③ Metric variance capping and log transformation—continuous revenue distributions contain extreme whale spenders ($>$10,000$), causing $sigma^2$ to explode; capping revenue at the 99th percentile or taking $ln(1 + text{spend})$ cuts $sigma^2$ by 3x, vastly reducing sample size requirements. ④ Minimum Detectable Effect (MDE) alignment with business value—setting MDE too small wastes testing capacity on trivial gains; setting MDE too large misses meaningful cumulative optimizations; MDE should be calibrated to the smallest lift that justifies engineering maintenance costs. ⑤ Fixed-horizon testing vs. Continuous peeking—evaluating $p$-values daily during an experiment inflates false positive rates; sample size formulas assume evaluation occurs strictly at pre-calculated horizon $N$. ⑥ Interview takeaway—write out the sample size formula $n approx frac{16 sigma^2}{text{MDE}^2}$, explain the quadratic $1/text{MDE}^2$ scaling law with a concrete calculation, define Type I/II errors, and explain how CUPED reduces sample size requirements.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 事后解释’不显著’为’没效果’(可能样本不足)
- ⚠️ MDE 设得过小导致实验无法完成
English Pitfalls:
– Halving the target MDE without recognizing that required traffic volume increases by a factor of 4, causing tests to run for months.
– Running under-powered experiments on small traffic slices, discarding genuine algorithmic improvements that fail to reach statistical significance.
– Calculating sample size on skewed continuous revenue metrics without outlier capping, underestimating required traffic by a factor of 10.
六、高频深度面试追问与预测 (Follow-Up Questions)
- 为什么 MDE 减半样本量要 ×4?
- Why does the ‘Winner’s Curse’ cause statistically significant results from under-powered experiments to drastically overestimate true effect sizes?
- 如何选 MDE?
- How does outlier percentile capping on skewed revenue metrics reduce required sample size without introducing unacceptable bias?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
在线推荐实验与业务指标:CTR、CVR、留存时长、网络溢出效应与 CUPED(Online Metrics & A/B Testing: CTR, CVR, CUPED & Spillover) - 🗺️ 知识图谱模块:
数据科学与因果实验导图
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。