所属模块:
M1 · 数学与统计基础 (Mathematics & Statistics Fundamentals)| 专题分类:实验设计 (A/B) (实验设计 (A/B))| 难度等级:Easy
一、核心一句话结论 (One-Sentence Summary)
定假设与指标 → 功效分析与样本量 → 随机分流 → 运行固定时长 → 校验分流 → 统计检验 → 决策与复盘。
A standard A/B experiment progresses through seven rigorous phases: business hypothesis formulation, metric taxonomy definition, power-based sample sizing, randomization hashing, SRM integrity sanity checks, significance evaluation, and post-launch guardrail monitoring.
二、核心考点要义 (Key Insights)
- 📌 分流单元要与干扰边界一致(用户级 vs 会话级)
- 📌 必须预先注册主指标,防 p-hacking
English Insights:
– Lifecycle: 1. Hypothesis & Metrics -> 2. Sample Size (MDE) -> 3. Hash Assignment -> 4. A/A & SRM Checks -> 5. Statistical Inference -> 6. Decision & Rollout.
– Metric Hierarchy: Primary decision metric (e.g. conversion rate), Secondary behavioral diagnostics (click depth), and Guardrail safety metrics (latency, crash rate, unsubscribe rate).
– Unit of Randomization: Must align strictly with the treatment experience (e.g. User ID for personalized UI, Request ID for stateless backend optimizations).
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$nproptofrac{sigma^2}{Delta^2},qquad text{SUTVA}: text{无干扰}$$
七个阶段的要点:① 定假设与指标——明确可检验的假设(H₁ 是什么)、指定唯一 primary metric 与若干护栏指标,并预注册分析计划(防 p-hacking 与 HARKing);② 功效分析——按 MDE、α=0.05、power=0.8 反解样本量与实验时长(注意日活限制:若日活 10 万、需 50 万样本,则至少 5 天);③ 随机分流——用哈希(如 hash(user_id+salt) % 100)保证确定性与均匀性,分流单元需与干扰边界一致(有社交/库存干扰时用集群随机化);④ 运行——固定时长(避免 peeking),覆盖完整周周期(消除工作日/周末差异);⑤ 校验——做 A/A 检验与 SRM 检验(样本比例失配),确认分流无偏;⑥ 统计检验——按预注册方法(t 检验/序贯检验/CUPED)计算效应量与 CI;⑦ 决策与复盘——结合业务显著性、护栏指标与长期影响做决策,记录结论与后续实验。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Rigorous statistical protocol: (1) Sizing: Compute minimum sample size per variant $N = frac{2(z_{1-alpha/2} + z_{1-beta})^2 sigma^2}{delta^2}$ for targeted MDE $delta$. (2) Assignment: Hash user identifier using salt: $text{Bucket} = text{MurmurHash3}(text{user_id} + text{salt}) pmod{100}$. (3) Integrity Check: Evaluate Sample Ratio Mismatch via Chi-Square goodness-of-fit: $chi^2 = sum_{i} frac{(O_i – E_i)^2}{E_i} sim chi^2_{k-1}$. If $p_{text{SRM}} < 0.001$, the experiment is invalid. (4) Treatment Effect: Compute lift $hat{Delta} = bar{Y}_B – bar{Y}_A$ with standard error $text{se}(hat{Delta}) = sqrt{frac{s_A^2}{N_A} + frac{s_B^2}{N_B}}$, applying CUPED covariate adjustment.
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
最易被忽视的三点:① SUTVA(无干扰假设)——标准 A/B 要求一个个体的结果不受其他个体处理分配的影响。但在社交网络(好友互相影响)、双边市场(买卖双方)、共享资源(库存/运力)中该假设被违反,处理组会’污染’对照组,导致效应被低估或高估。解法是集群随机化(按地理/社区分组)、switchback(时间轮换)、或专门的干扰感知设计。② 指标的口径与延迟——转化/退款类指标有延迟,若实验期太短会低估效应;应确保观测窗口覆盖完整的转化周期,或使用代理指标并做延迟校正。③ 多重比较与选择性报告——若同时看 20 个指标并只报告显著的,假阳性率极高;正确做法是预注册唯一主指标,其余作为探索性分析并明确标注。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
Common industrial trade-offs: (1) Duration vs. Novelty effect: Running for at least 1-2 full weeks captures weekly cyclical patterns (weekend vs weekday) while letting novelty effects (users clicking due to unfamiliar UI) burn off. (2) Ramp-up strategy: Start at 1% -> 5% -> 50% traffic allocation to protect production revenue against catastrophic bugs before launching full statistical evaluation.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 未做 SRM/A/A 校验就解读结果
- ⚠️ 在存在网络效应的场景用个体级随机化
English Pitfalls:
– Changing the allocation ratio mid-experiment without stratifying analysis (causes severe Simpson’s paradox and artificial SRM).
– Concluding significance based on secondary metrics when the pre-registered primary metric was flat (cherry-picking).
六、高频深度面试追问与预测 (Follow-Up Questions)
- 为什么要在实验前注册指标?
- How do multiple layers of concurrent experiments run without interference using orthogonal hashing salts?
- 网络效应/溢出效应会怎样破坏 A/B?
- What is the difference between intent-to-treat (ITT) and average treatment effect on the treated (ATT) in online experiments?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
工业级 A/B 实验设计、分流正交、SRM 卡方排查与方差缩减(Industrial A/B Testing: Split, SRM & Variance Reduction) - 🗺️ 知识图谱模块:
数据科学与因果实验导图
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。