所属模块:
M1 · 数学与统计基础 (Mathematics & Statistics Fundamentals)| 专题分类:假设检验 (Hypothesis Testing)| 难度等级:Hard
一、核心一句话结论 (One-Sentence Summary)
反复查看并随时停止会极大抬高假阳性;需用 alpha-spending 或 always-valid 方法。
The Peeking Problem occurs when continuously checking p-values and stopping once $p<0.05$, which inflates false positive rates from 5% to over 30%; resolved via alpha spending or sequential probability ratio tests.
二、核心考点要义 (Key Insights)
- 📌 固定时长实验是简单解法
- 📌 Sequential test / group sequential / mSPRT 允许安全 peeking
English Insights:
– The Peeking Problem: Repeatedly evaluating fixed-horizon tests over time turns the decision into $sup_{t le T} Z_t$, guaranteeing eventual false positive crossing by the Law of Iterated Logarithm.
– Alpha Spending Function (Lan-DeMets / Pocock / O’Brien-Fleming): Allocates portions of total significance budget $alpha$ across intermediate evaluation points.
– Always-Valid Inference (mSPRT / e-values): Constructs nonnegative supermartingales allowing arbitrary, continuous real-time peeking without inflating Type I error.
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$text{alpha spending}: alpha(t) text{随累计信息量分配,}sumalpha(t)=alpha$$
peeking 问题的数学本质:经典假设检验的 α 是一次性检验的假阳性率,其推导假设检验只做一次。若在数据累积过程中反复检验并在任一次显著时停止(optional stopping),假阳性率会随检验次数单调上升——对布朗运动式的随机游走,连续监控下的最终假阳性率趋向 1(因为随机游走几乎必然会在某个时刻越过任意有限阈值)。直觉上,这相当于把’一次抽样’变成’多次抽样取最大’,而最大值的分布远比单次抽样极端。这是 A/B 测试中’跑三天看到显著就上线’的经典陷阱。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Law of the Iterated Logarithm (LIL): For standardized random walk $S_n = sum_{i=1}^n X_i$, $limsup_{ntoinfty} frac{S_n}{sqrt{2n log log n}} = 1$ almost surely under $H_0$. Because $sqrt{2n log log n} to infty$, the sample path will cross any fixed threshold $z_{1-alpha/2} sqrt{n}$ with probability 1 as $ntoinfty$. Hence, peeking daily for 30 days inflates effective $alpha$ to $approx 30%$. The Mixture Sequential Probability Ratio Test (mSPRT, Johari et al., 2017) computes test statistic $Lambda_n = int prod_{i=1}^n frac{p(x_imid theta)}{p(x_imid 0)} dH(theta)$. By Ville’s inequality for supermartingales, $P(exists n : Lambda_n ge 1/alpha mid H_0) le alpha$, guaranteeing valid inference at all stopping times.
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
三类正确做法:① 固定时长/固定样本量——最简单的解法,预先用功效分析定 n,跑满再分析一次;缺点是无法提前止损(若效应明显为负也只能跑完)。② Group Sequential(分组序贯)——预先指定 k 次中期分析,用 alpha-spending 函数(如 O’Brien-Fleming)分配 α 预算,早期检验用更严阈值(因为早期数据少、波动大),总 α 仍为 0.05。这是临床试验的标准做法。③ Always-valid 推断——用 mSPRT(混合序贯概率比检验)或 confidence sequences 构造在任何停止时刻都有效的置信区间,可随时查看而不失真;代价是区间略宽(功效略低)。补充说明:贝叶斯 A/B 在一定程度上规避该问题——后验随数据更新是合法的,P(p_A>p_B) 不会因反复查看而系统性失真(但若用后验做频繁的’上线/不上线’决策仍需考虑预期损失)。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
Standard fixed-horizon A/B testing forbids inspecting results before reaching the planned sample size $N$. However, product teams want to stop losing experiments early to minimize revenue destruction. Modern platforms (Optimizely, Netflix, Uber) implement mSPRT or O’Brien-Fleming boundaries, enabling continuous monitoring dashboards with rigorous statistical guarantees at the cost of slightly higher average sample size.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 跑几天看到显著就停止(peeking 导致假阳性膨胀)
- ⚠️ 认为贝叶斯方法完全无需考虑停止规则
English Pitfalls:
– Checking the A/B dashboard every morning and turning off the experiment the moment $p < 0.05$.
– Treating standard confidence intervals as valid when early stopping rules were applied.
六、高频深度面试追问与预测 (Follow-Up Questions)
- 为什么 peeking 会抬高假阳性?
- How does Ville’s inequality for nonnegative supermartingales provide time-uniform confidence boundaries?
- 贝叶斯 A/B 能否规避这个问题?
- What is the difference between Pocock boundaries (constant threshold) and O’Brien-Fleming boundaries (conservative early on)?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
数理统计假说检验、P 值、I/II 类错误与统计功效(Hypothesis Testing, P-Values, Power & Type I/II Error) - 🗺️ 知识图谱模块:
数理基础思维导图
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。