所属模块:
M1 · 数学与统计基础 (Mathematics & Statistics Fundamentals)| 专题分类:实验设计 (A/B) (实验设计 (A/B))| 难度等级:Medium
一、核心一句话结论 (One-Sentence Summary)
在已知的重要协变量层内独立随机化,保证各层内处理/对照平衡;小样本或强预测协变量时必要。
Stratified randomization divides the population into homogeneous strata based on key covariates (e.g. Operating system, spend tier) and randomizes within each stratum, guaranteeing perfect covariate balance and reducing treatment effect variance.
二、核心考点要义 (Key Insights)
- 📌 保证每层的组间平衡(降低方差)
- 📌 层数过多会导致层内样本过少
English Insights:
– Core Mechanism: Partitions population into $K$ mutually exclusive strata $S_1, dots, S_K$; within each stratum, exactly $50%$ of units are allocated to Treatment and $50%$ to Control.
– Guaranteed Balance: Eliminates the possibility of accidental pre-experiment covariate imbalance (e.g. 60% iOS in treatment vs 40% in control).
– Variance Reduction: Removes between-strata variance from the standard error of treatment effect: $text{Var}(hat{tau}{text{strat}}) le text{Var}(hat{tau})$.}
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$text{stratified}: text{randomize within each stratum } s$$
分层随机化的做法:按一个或多个实验前协变量(如地区、设备类型、用户活跃度分位)把样本划分为若干层(stratum),在每层内独立进行随机化(保证各层内处理/对照比例符合预期)。为什么有效:① 保证平衡——若某协变量与指标强相关,完全随机化在小样本下可能产生组间不平衡(如处理组恰好包含更多高活跃用户),分层消除了这种风险;② 降低方差——分层估计量 = Σ(层权重 × 层内效应),其方差 ≤ 完全随机化的方差(因为消除了层间差异的贡献),这与 CUPED 的目标一致(都是用协变量降方差)。何时必要:① 样本量小——随机化本身无法保证平衡时;② 协变量与指标强相关——不平衡会显著偏倚结果;③ 需要按层分析——如关心不同地区的效应(分层保证每层都有足够样本);④ 分层是业务流程要求——如按地区分批上线。
📖 查看英文严格数学推导 (English Mathematical Derivation)
By the Law of Total Variance, population metric variance decomposes into: $sigma^2 = text{Var}(Y) = sum_{k=1}^K p_k sigma_k^2 + sum_{k=1}^K p_k (mu_k – mu)^2 = sigma_{text{within}}^2 + sigma_{text{between}}^2$, where $p_k$ is stratum weight and $sigma_k^2$ is within-stratum variance. Under simple random assignment, variance of estimated treatment effect is $text{Var}(hat{tau}_{text{simple}}) = frac{2sigma^2}{N} = frac{2(sigma_{text{within}}^2 + sigma_{text{between}}^2)}{N}$. Under stratified randomization, because each stratum is balanced exactly, the between-strata component $sigma_{text{between}}^2$ is completely purged from estimation noise: $text{Var}(hat{tau}_{text{strat}}) = sum_{k=1}^K p_k^2 left(frac{2sigma_k^2}{N_k}right) = frac{2sigma_{text{within}}^2}{N}$.
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
实践要点:① 分层 vs 后分层(Post-stratification)——分层在随机化时使用协变量(需在实验前可知);后分层在分析时按协变量加权调整(可用于实验后才获得的协变量,如用户历史行为);两者都能降方差,后分层更灵活但需保证各层有足够样本。② 分层 vs CUPED——两者都降方差但机制不同:分层用于离散协变量(层内完全平衡),CUPED 用于连续协变量(回归调整);实践中可叠加(先分层再 CUPED)。③ 层数的权衡——层数多则每层样本少(随机化不平衡风险再现)、分析复杂;通常 2–5 个分层变量、每层至少数百样本。④ 分层的正确性——分层变量必须是实验前确定的(不能被处理影响),否则会引入偏差;常见错误是用’实验期行为’分层。⑤ 分析——分层实验应用分层估计量(层内效应按层权重加权),而非简单合并所有样本(后者会引入 Simpson 悖论风险);也可用带层固定效应的回归。⑥ 与分组随机化的区别——分层是对个体按协变量分组后随机化;集群随机化是对整组随机化(处理干扰);两者常结合(分层集群随机化)。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
When stratified randomization is mandatory: (1) Small sample experiments ($N < 5000$, such as enterprise B2B A/B tests or hospital trials) where simple random sampling frequently produces bad imbalances. (2) Presence of heavy whales: In gaming or ad monetization, stratified bucketing based on historical spend tiers prevents top-spending ‘whales’ from clustering by chance into one bucket. In large-scale consumer apps ($N > 10^7$), simple randomization achieves balance by the Law of Large Numbers, making post-experiment CUPED adjustment mathematically equivalent and operationally simpler.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 用实验期行为做分层变量(引入偏差)
- ⚠️ 分层实验用简单合并分析(丢失分层收益)
English Pitfalls:
– Stratifying on post-treatment variables (must strictly stratify on pre-treatment baseline covariates).
– Over-stratifying across too many crossed variables, leading to empty or single-unit strata.
六、高频深度面试追问与预测 (Follow-Up Questions)
- 分层随机化与后分层的区别?
- Why is post-stratification (or CUPED ANCOVA) asymptotically equivalent to pre-experiment stratified randomization in large samples?
- 分层与 CUPED 如何配合?
- How does re-randomization (Morgan & Rubin, 2012) reject unbalanced assignments before deploying treatments?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
工业级 A/B 实验设计、分流正交、SRM 卡方排查与方差缩减(Industrial A/B Testing: Split, SRM & Variance Reduction) - 🗺️ 知识图谱模块:
数据科学与因果实验导图
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。