【AI 核心深度 M7-089】解释反事实评估的因果图与假设(Explain Causal Directed Acyclic Graphs (DAGs), Do-Calculus, and Backdoor Paths in Counterfactual Evaluation)深度数理推导与工程落地解析

所属模块:M7 · 检索、排序与推荐系统 (Retrieval, Ranking & RecSys) | 专题分类:离线评估与 OPE (Offline Evaluation & OPE) | 难度等级:Hard

一、核心一句话结论 (One-Sentence Summary)

用因果图(DAG)表达’混淆、动作、奖励’的关系;OPE 的目标是估计’干预(do)’的效果,需切断混淆路径。

ADVERTISEMENT · 赞助推荐

Counterfactual evaluation utilizes causal Directed Acyclic Graphs (DAGs) to isolate causal policy effects from observational correlations; closing backdoor confounding paths via Pearl’s do-calculus transforms observational logs into interventional distributions.

二、核心考点要义 (Key Insights)

  • 📌 因果图(DAG):特征 X、动作 A、奖励 R
  • 📌 混淆:X 同时影响 A 与 R(后门路径)
  • 📌 OPE 目标:估计 P(R|do(A))(干预分布)而非 P(R|A)(观测分布)

English Insights:
– Observational vs. Interventional distribution: P(R | A=a) reflects correlation confounded by context; P(R | do(A=a)) reflects true causal outcome under algorithmic intervention.
– The Backdoor Criterion: Confounding occurs when context X simultaneously influences action selection A and reward R; conditioning on X blocks the backdoor path.
– Causal DAG topology: Nodes represent User Context (X), Logging Policy Action (A), Candidate Action (A), and Observable Reward (R).*

三、核心数学原理与机理推导 (Mathematical Principles & Derivation)

$$text{causal graph}: Xto A, Xto R, Ato R;qquad text{goal}: P(R|do(A))$$

数学机理:因果图与反事实——(1) 因果图(DAG)——用有向图表示变量间的因果关系:(a) X(上下文特征:用户/场景);(b) A(动作:展示哪个物品);(c) R(奖励:点击/转化);(d) 边——X→A(特征影响策略的选择)、X→R(特征影响奖励)、A→R(动作影响奖励)。(2) 混淆(confounding)——X 同时影响 A 与 R → 形成’后门路径’ X→A 与 X→R;(a) 后果——观测到的 P(R|A) 不等于因果的 P(R|do(A))(因为 X 的分布在不同 A 下不同);(b) 例子——’高价值用户更常被展示某广告’ → 该广告的观测点击率高(但可能是用户本身爱点,而非广告效果好)。(3) OPE 的目标——估计干预分布(interventional distribution) P(R|do(A))(’如果强制对所有用户展示动作 A,奖励的分布’),而非观测分布 P(R|A)(’历史中 A 被展示时的奖励’)。(4) ‘do(A)’与’条件 A’的差异(核心)——(a) P(R|A=a)——’在历史中 A=a 的样本里,R 的分布’(受混淆影响);(b) P(R|do(A=a))——’干预设置 A=a(切断 X→A 的边)后,R 的分布’(因果效果);(c) 两者相等当且仅当’无混淆’(X 不影响 A,或所有混淆都被观测并调整)。(5) 后门准则(backdoor criterion)——若一组变量 Z 满足’阻断所有从 A 到 R 的后门路径’且’不含 A 的后代’,则 P(R|do(A))=Σ_z P(R|A,Z=z)P(Z=z)(调整 Z 后可识别);意义——它给出’需要调整哪些变量’的判据。与 OPE 方法的对应——(a) IPS——本质是’用倾向 1/P(A|X) 加权’来’切断 X→A 的边’(重新加权使 A 与 X 独立);(b) DM——用模型预测’给定 X 与 A 的 R’(隐含’无混淆’假设);(c) DR——两者结合;(d) 随机化——物理上切断 X→A 的边(最可靠)。其他因果概念——(a) 反事实(counterfactual)——’如果当时选另一个动作会怎样’(不可观测);(b) 中介(mediation)——A→M→R 的间接路径;(c) 工具变量(IV)——有隐藏混淆时的识别策略;(d) 敏感性分析——’若混淆强度为 X,估计偏差多大’。与其他问题的关系——(a) 与’为什么不能用历史数据’(混淆);(b) 与’位置偏置’(位置是混淆);(c) 与’样本选择偏差(SSB)’(同一类问题)。实践建议——(a) 画因果图(明确混淆与路径);(b) 明确目标是 P(R|do(A))(而非 P(R|A));(c) 用后门准则判断’需要调整哪些变量’;(d) 倾向模型要包含所有混淆(否则 IPS 有偏);(e) 敏感性分析(评估’未观测混淆’的影响);(f) 随机化(最可靠)。度量——(a) OPE 估计 vs A/B(偏差);(b) 敏感性分析(未观测混淆的影响);(c) 后门调整的完备性。

📖 查看英文严格数学推导 (English Mathematical Derivation)

Causal Graph Theory & Do-Calculus: Pearl’s Causal Framework for RecSys.

(1) The Observational Causal DAG:
In a production recommendation logging system, the causal relationships are represented by Directed Acyclic Graph $mathcal{G}$:
$$begin{matrix} & X text{ (User Context: Demographics, History, Time)} & \ & swarrow qquadqquadqquadqquad searrow & \ A text{ (Logged Action)} & xrightarrow{qquadqquadqquadqquad} & R text{ (Observed Reward: Click, CVR)} end{matrix}$$
– Edge $X to A$: The logging policy chooses action $A$ conditioning on context $X$: $A sim pi_0(a mid X)$.
– Edge $X to R$: User context directly influences intrinsic click/purchase affinity (e.g., wealthy users buy more).
– Edge $A to R$: The causal impact of presenting item $A$ on user reward $R$.

(2) The Backdoor Confounding Path:
The path $A leftarrow X to R$ is a backdoor path from $A$ to $R$.
Because of this path, the observational conditional distribution mixes true causation with spurious context correlation:
$$P(R mid A = a) = sum_x P(R mid A=a, X=x) P(X=x mid A=a) neq P(R mid do(A = a))$$
Example: If policy $pi_0$ only recommends luxury watches ($A$) to wealthy users ($X$), then luxury watches exhibit high purchase rates $P(R mid A)$ not because the watch is appealing, but because wealthy users buy everything.

(3) The Backdoor Adjustment Formula (Pearl’s Do-Calculus):
To evaluate the causal effect of an intervention $do(A = pi_1(X))$, we sever all incoming edges into $A$ ($X notto A$), forcing $A$ to be chosen by candidate policy $pi_1$:
$$P(R mid do(A = a)) = sum_{x in mathcal{X}} P(R mid A = a, X = x) P(X = x)$$
The expected policy value under intervention $pi_1$ is:
$$V(pi_1) = mathbb{E}_{X sim P(X)} left[ sum_{a in mathcal{A}} pi_1(a mid X) mathbb{E}[R mid X, A = a] right]$$
This bridges the gap between observational logging and counterfactual interventional reasoning.

四、工业级落地权衡与工程考量 (Industrial Trade-offs)

深度剖析与工程权衡:① ‘do(A) ≠ 条件 A’是因果推断的核心区分——面试中能准确解释这一点是深度理解的标志。② ‘后门准则给出需要调整的变量’——这是实践中的判据。③ ‘IPS 本质是切断 X→A 的边’——用重加权实现’干预’;这是 IPS 的因果解释。④ ‘随机化是物理上切断边’——最可靠(不需假设)。⑤ ‘敏感性分析’评估未观测混淆——实用且重要(因为’无混淆’无法验证)。⑥ 面试要点——被问’反事实评估的假设’,应给出’因果图 + 混淆 + do(A) vs 条件 A + 后门准则 + IPS/随机化的因果解释 + 敏感性分析‘;能准确区分’do’与’条件’是深度理解的标志。

⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)

In-Depth Analysis & Engineering Trade-offs: ① The Backdoor Criterion requirement: Logging all confounding variables—conditioning on $X$ blocks the backdoor path $A leftarrow X to R$ if and only if all variables that simultaneously influence action choice and user reward are observed and logged; if the online service used user location to choose items but failed to record location in Kafka, unobserved confounding reopens the backdoor, making unbiased counterfactual inference impossible. ② Front-door adjustment when context is unobserved—if context $X$ is partially hidden, but an intermediate mediator $M$ exists (e.g., $A to M to R$, such as user reading the product title), Pearl’s front-door criterion can theoretically recover causal effects without observing $X$. ③ Selection bias as collider conditioning—conditioning on clicks ($C=1$) when evaluating post-click conversion (CVR) treats a collider as observed: $A to C leftarrow U$; conditioning on collider $C$ opens a spurious correlation between $A$ and unobserved user intent $U$, mathematically explaining Sample Selection Bias (SSB) and justifying ESMM. ④ Instrumental Variables (IV) for recommendation—when unobserved confounding is pervasive, using random algorithmic assignment or network jitter as an Instrumental Variable $Z$ ($Z to A to R$ with $Z perp U$) isolates true causal treatment effects. ⑤ Causal DAG validation in system design—drawing explicit causal DAGs during architecture reviews forces engineering teams to identify which features must be captured at inference time to enable future counterfactual modeling. ⑥ Interview takeaway—draw the causal DAG ($X to A, X to R, A to R$), explain why $P(R mid A) neq P(R mid do(A))$ due to the backdoor path $A leftarrow X to R$, derive the Backdoor Adjustment formula, and explain collider bias in CVR modeling.

五、常见面试避坑陷阱 (Common Pitfalls & Traps)

  • ⚠️ 把 P(R|A) 当作 P(R|do(A))(混淆未调整)
  • ⚠️ 不做敏感性分析(无法评估未观测混淆)

English Pitfalls:
– Confusing observational correlation P(R | A) with interventional causation P(R | do(A)), deploying policies that replicate spurious historical correlations.
– Failing to log critical contextual features that the online ranking algorithm used, permanently opening unblocked backdoor confounding paths.
– Conditioning on post-treatment colliders (e.g., conditioning only on clicked samples in CVR), inducing spurious correlations between items and user intent.

六、高频深度面试追问与预测 (Follow-Up Questions)

  1. ‘do(A)’与’条件 A’的差异?
  2. How does conditioning on a collider node (e.g., user click C in A -> C <- U) induce selection bias in downstream conversion modeling?
  3. 后门准则是什么?
  4. What properties must an Instrumental Variable (IV) satisfy to recover unbiased causal treatment effects when hidden confounders exist?

七、知识图谱对齐 (Knowledge Graph Anchor)

  • 🔗 关联底层卡片:推荐系统离线评估与离线策略评估 (OPE):逆倾向得分 (IPS) 与重要性采样 (Off-Policy Evaluation (OPE): IPS, Doubly Robust & Calibration)
  • 🗺️ 知识图谱模块:机器学习工程师高频考点导图

🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)

本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。

👉 前往 TalentMe 交互式研读本题 (M7-089) →


Discover more from AirSOTA – Air School Of Thoughts AtoZ

Subscribe to get the latest posts sent to your email.