所属模块:
M1 · 数学与统计基础 (Mathematics & Statistics Fundamentals)| 专题分类:因果推断 (Causal Inference)| 难度等级:Easy
一、核心一句话结论 (One-Sentence Summary)
因果效应是个体两种潜在结果的差;观察到的只有一种,故需假设才能识别。
The Rubin Causal Model defines causal effect as the difference between potential outcomes under treatment and control: $Y_i(1) – Y_i(0)$; the Fundamental Problem of Causal Inference is that we can only ever observe one potential outcome for any individual.
二、核心考点要义 (Key Insights)
- 📌 随机化使处理独立于潜在结果 → ATE 可识别
- 📌 观察数据需条件独立/可忽略性假设
English Insights:
– Potential Outcomes: $Y_i(1)$ (outcome if treated) and $Y_i(0)$ (outcome if control).
– Observed outcome: $Y_i = T_i Y_i(1) + (1 – T_i) Y_i(0)$, where $T_i in {0, 1}$.
– Average Treatment Effect (ATE): $tau_{text{ATE}} = E[Y(1) – Y(0)]$.
– SUTVA (Stable Unit Treatment Value Assumption): No interference between units and no hidden variations of treatment.
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$tau_i=Y_i(1)-Y_i(0),qquad mathrm{ATE}=mathbb E[Y(1)-Y(0)]$$
Rubin 潜在结果框架的核心:对每个个体 i 定义两个潜在结果 Y_i(1)(接受处理时)与 Y_i(0)(未接受时),个体因果效应 τ_i=Y_i(1)−Y_i(0)。根本困难(因果推断的基本问题):我们只能观测到其中一个——若个体接受了处理,观测 Y_i=Y_i(1),而 Y_i(0) 是反事实,永远不可观测。因此个体因果效应不可识别,只能识别群体平均效应。识别条件:随机化使处理分配 T 独立于潜在结果 (Y(0),Y(1)),即 T⊥(Y(0),Y(1)),于是 E[Y|T=1]−E[Y|T=0]=E[Y(1)]−E[Y(0)]=ATE——随机化通过’构造可比组’解决了反事实缺失问题。观察性数据需条件可忽略性(T⊥(Y(0),Y(1))|X,即控制了 X 后处理与潜在结果独立)。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Observed difference in group means decomposes into treatment effect plus selection bias: $E[Ymid T=1] – E[Ymid T=0] = E[Y(1)mid T=1] – E[Y(0)mid T=0] = E[Y(1)mid T=1] – E[Y(0)mid T=1] + E[Y(0)mid T=1] – E[Y(0)mid T=0] = tau_{text{ATT}} + text{Selection Bias}$, where $text{Selection Bias} = E[Y(0)mid T=1] – E[Y(0)mid T=0]$. In observational data, treated units typically have higher baseline outcomes even without treatment (e.g. Users who adopt a premium feature have higher organic activity), so selection bias $ne 0$. Randomized Controlled Trials (RCT) enforce $T perp (Y(0), Y(1))$, eliminating selection bias so $E[Y(0)mid T=1] = E[Y(0)mid T=0] = E[Y(0)]$, guaranteeing that observed difference equals true ATE.
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
关键概念辨析:① ATE vs ATT vs ATE——ATE 是全体平均效应,ATT 是处理组的平均效应 E[Y(1)−Y(0)|T=1],ATU 是未处理组的。三者不同,且当处理效应异质时(如只有高意愿用户受益),用 ATT 外推到全体是错误的外推;实验测的是 ATE(或按分流比例混合),而观察性研究常识别 ATT。② SUTVA——潜在结果框架隐含假设’个体 i 的结果只依赖自己的处理’(无干扰)且处理版本唯一;违反时需干扰感知设计。③ 与 Pearl 的 do-calculus 的关系——两套框架数学等价但语言不同:Rubin 框架强调潜在结果与随机化,Pearl 框架强调 DAG 与后门/前门准则。实践中 DAG 用于判断该控制哪些变量(后门准则),Rubin 框架用于估计与推断,两者互补。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
SUTVA requirements: (1) No Interference: User $i$’s outcome is unaffected by whether user $j$ receives treatment. In ride-sharing (Uber/Lyft) and social networks, treating drivers or users spills over to competitors, violating SUTVA and requiring switchback or cluster randomization. (2) Consistency: The treatment must be a single, uniform intervention without ambiguous variations.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 认为可以估计个体因果效应(只能识别群体平均)
- ⚠️ 把 ATT 当作 ATE 外推到全体
English Pitfalls:
– Confusing association $E[Ymid T=1] – E[Ymid T=0]$ with causation $E[Y(1) – Y(0)]$ in observational datasets.
– Ignoring network spillover interference in two-sided marketplaces, violating SUTVA.
六、高频深度面试追问与预测 (Follow-Up Questions)
- 为什么随机化能识别因果效应?
- What is the difference between Average Treatment Effect (ATE), Average Treatment Effect on the Treated (ATT), and Conditional Average Treatment Effect (CATE)?
- ATE 与 ATT 的区别?
- How do spillover effects violate SUTVA in marketplace experiments?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
因果推断框架:潜在结果模型、倾向评分匹配与双重差分(Causal Inference: Potential Outcomes, PSM & DiD) - 🗺️ 知识图谱模块:
数据科学与因果实验导图
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。