所属模块:
M1 · 数学与统计基础 (Mathematics & Statistics Fundamentals)| 专题分类:实验设计 (A/B) (实验设计 (A/B))| 难度等级:Easy
一、核心一句话结论 (One-Sentence Summary)
两组都跑原策略;用于验证分流系统与统计管线,应得到无显著差异。
An A/A test assigns identical experiences to both groups to validate experimental infrastructure integrity, verifying that the empirical false positive rate equals $alpha$ and p-values follow a uniform distribution.
二、核心考点要义 (Key Insights)
- 📌 上线前跑 A/A 校验分流均匀性与方差估计
- 📌 若 A/A 频繁显著 → 分流或统计有问题
English Insights:
– Primary Purpose: Validates traffic splitting randomization, pipeline data logging, and variance estimation algorithms without real treatment effects.
– Diagnostic Criterion 1: P-values across 1000 simulated A/A runs must be uniformly distributed $mathcal{U}(0, 1)$ (verified via Kolmogorov-Smirnov test).
– Diagnostic Criterion 2: Exactly $alpha = 5%$ of A/A tests should reject $H_0$ at the $0.05$ significance level; any higher rate signals pipeline bias.
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$text{A/A}: Pr(text{false positive})approxalpha$$
A/A 测试把两组都分配到相同的策略(原策略),因此理论上任何观测到的差异都来自随机噪声。它的用途是验证实验基础设施的正确性,具体检查四项:① 分流是否均匀(组间样本量、用户特征分布);② 指标计算管线是否正确(同样的数据两条路径应给出一致结果);③ 方差估计是否准确(A/A 的假阳性率应接近名义 α=5%);④ 是否存在系统性偏差(如缓存、日志丢失、埋点差异导致的不对称)。判读标准:单个 A/A 测试的 p 值应服从均匀分布 U(0,1);若在多个 A/A 测试中显著比例远高于 5%,说明基础设施有问题。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Under the sharp null hypothesis of identical variants ($Y_A, Y_B sim F$), the test statistic $Z = frac{bar{Y}_B – bar{Y}_A}{sqrt{sigma_A^2/N_A + sigma_B^2/N_B}}$ converges to standard normal $mathcal{N}(0, 1)$. The two-sided p-value is $P = 2(1 – Phi(|Z|))$. For any $u in [0, 1]$, $P(p le u) = P(2(1 – Phi(|Z|)) le u) = P(|Z| ge Phi^{-1}(1 – u/2)) = 2(1 – Phi(Phi^{-1}(1 – u/2))) = u$, proving uniform distribution $mathcal{U}(0, 1)$. If the QQ-plot of empirical p-values deviates from the 45-degree diagonal line, or if the false positive rate is inflated, systematic pipeline errors exist.
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
实践要点:① 上线前必做——新实验平台/新指标/新分流逻辑上线前应跑 A/A 验证,这是防止’实验结果不可信’的第一道防线;② A/A 显著的可能原因——分流不均匀(哈希碰撞、过滤逻辑不对称)、埋点丢失(某组日志缺失)、缓存污染(CDN 缓存跨组)、时序偏差(分组上线时间不同)、方差低估(未考虑聚类结构);③ A/A 的局限——A/A 只能检测对称的系统性偏差,若两组都有同样的偏差(如整体日志丢失),A/A 无法发现;④ 实践中常用 A/A/A(三组)以更好地估计方差,或使用样本分割验证(把同一处理的数据随机分成两半,检查是否一致)。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
Modes of running A/A tests: (1) Dedicated A/A test: Run for 1-2 weeks before launching major new platform algorithms. (2) Continuous offline synthetic A/A testing: Re-split control group data into two random pseudo-buckets in historical logs and evaluate false positive rates across thousands of iterations, catching metric definition bugs with zero opportunity cost.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 跳过 A/A 直接跑 A/B(基础设施问题会被误读为策略效果)
- ⚠️ 认为 A/A 不显著就证明基础设施完全正确
English Pitfalls:
– Discarding an A/B platform setup because a single A/A test yielded $p = 0.03$ (by definition, 5% of all valid A/A tests yield $p < 0.05$).
– Failing to check for Sample Ratio Mismatch (SRM) inside the A/A test.
六、高频深度面试追问与预测 (Follow-Up Questions)
- A/A 显著说明什么?
- How does the Kolmogorov-Smirnov (KS) test formally test whether empirical A/A p-values deviate from $mathcal{U}(0, 1)$?
- 如何诊断分流不均?(SRM 检验)
- Why does un-modeled intra-user correlation in session-level metrics cause severe A/A p-value inflation?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
工业级 A/B 实验设计、分流正交、SRM 卡方排查与方差缩减(Industrial A/B Testing: Split, SRM & Variance Reduction) - 🗺️ 知识图谱模块:
数据科学与因果实验导图
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。