所属模块:
M7 · 检索、排序与推荐系统 (Retrieval, Ranking & RecSys)| 专题分类:离线评估与 OPE (Offline Evaluation & OPE)| 难度等级:Hard
一、核心一句话结论 (One-Sentence Summary)
用因果图(DAG)表达’混淆、动作、奖励’的关系;OPE 的目标是估计’干预(do)’的效果,需切断混淆路径。
Counterfactual evaluation utilizes causal Directed Acyclic Graphs (DAGs) to isolate causal policy effects from observational correlations; closing backdoor confounding paths via Pearl’s do-calculus transforms observational logs into interventional distributions.
二、核心考点要义 (Key Insights)
- 📌 因果图(DAG):特征 X、动作 A、奖励 R
- 📌 混淆:X 同时影响 A 与 R(后门路径)
- 📌 OPE 目标:估计 P(R|do(A))(干预分布)而非 P(R|A)(观测分布)
English Insights:
– Observational vs. Interventional distribution: P(R | A=a) reflects correlation confounded by context; P(R | do(A=a)) reflects true causal outcome under algorithmic intervention.
– The Backdoor Criterion: Confounding occurs when context X simultaneously influences action selection A and reward R; conditioning on X blocks the backdoor path.
– Causal DAG topology: Nodes represent User Context (X), Logging Policy Action (A), Candidate Action (A), and Observable Reward (R).*
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$text{causal graph}: Xto A, Xto R, Ato R;qquad text{goal}: P(R|do(A))$$
数学机理:因果图与反事实——(1) 因果图(DAG)——用有向图表示变量间的因果关系:(a) X(上下文特征:用户/场景);(b) A(动作:展示哪个物品);(c) R(奖励:点击/转化);(d) 边——X→A(特征影响策略的选择)、X→R(特征影响奖励)、A→R(动作影响奖励)。(2) 混淆(confounding)——X 同时影响 A 与 R → 形成’后门路径’ X→A 与 X→R;(a) 后果——观测到的 P(R|A) 不等于因果的 P(R|do(A))(因为 X 的分布在不同 A 下不同);(b) 例子——’高价值用户更常被展示某广告’ → 该广告的观测点击率高(但可能是用户本身爱点,而非广告效果好)。(3) OPE 的目标——估计干预分布(interventional distribution) P(R|do(A))(’如果强制对所有用户展示动作 A,奖励的分布’),而非观测分布 P(R|A)(’历史中 A 被展示时的奖励’)。(4) ‘do(A)’与’条件 A’的差异(核心)——(a) P(R|A=a)——’在历史中 A=a 的样本里,R 的分布’(受混淆影响);(b) P(R|do(A=a))——’干预设置 A=a(切断 X→A 的边)后,R 的分布’(因果效果);(c) 两者相等当且仅当’无混淆’(X 不影响 A,或所有混淆都被观测并调整)。(5) 后门准则(backdoor criterion)——若一组变量 Z 满足’阻断所有从 A 到 R 的后门路径’且’不含 A 的后代’,则 P(R|do(A))=Σ_z P(R|A,Z=z)P(Z=z)(调整 Z 后可识别);意义——它给出’需要调整哪些变量’的判据。与 OPE 方法的对应——(a) IPS——本质是’用倾向 1/P(A|X) 加权’来’切断 X→A 的边’(重新加权使 A 与 X 独立);(b) DM——用模型预测’给定 X 与 A 的 R’(隐含’无混淆’假设);(c) DR——两者结合;(d) 随机化——物理上切断 X→A 的边(最可靠)。其他因果概念——(a) 反事实(counterfactual)——’如果当时选另一个动作会怎样’(不可观测);(b) 中介(mediation)——A→M→R 的间接路径;(c) 工具变量(IV)——有隐藏混淆时的识别策略;(d) 敏感性分析——’若混淆强度为 X,估计偏差多大’。与其他问题的关系——(a) 与’为什么不能用历史数据’(混淆);(b) 与’位置偏置’(位置是混淆);(c) 与’样本选择偏差(SSB)’(同一类问题)。实践建议——(a) 画因果图(明确混淆与路径);(b) 明确目标是 P(R|do(A))(而非 P(R|A));(c) 用后门准则判断’需要调整哪些变量’;(d) 倾向模型要包含所有混淆(否则 IPS 有偏);(e) 敏感性分析(评估’未观测混淆’的影响);(f) 随机化(最可靠)。度量——(a) OPE 估计 vs A/B(偏差);(b) 敏感性分析(未观测混淆的影响);(c) 后门调整的完备性。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Causal Graph Theory & Do-Calculus: Pearl’s Causal Framework for RecSys.
(1) The Observational Causal DAG:
In a production recommendation logging system, the causal relationships are represented by Directed Acyclic Graph $mathcal{G}$:
$$begin{matrix} & X text{ (User Context: Demographics, History, Time)} & \ & swarrow qquadqquadqquadqquad searrow & \ A text{ (Logged Action)} & xrightarrow{qquadqquadqquadqquad} & R text{ (Observed Reward: Click, CVR)} end{matrix}$$
– Edge $X to A$: The logging policy chooses action $A$ conditioning on context $X$: $A sim pi_0(a mid X)$.
– Edge $X to R$: User context directly influences intrinsic click/purchase affinity (e.g., wealthy users buy more).
– Edge $A to R$: The causal impact of presenting item $A$ on user reward $R$.
(2) The Backdoor Confounding Path:
The path $A leftarrow X to R$ is a backdoor path from $A$ to $R$.
Because of this path, the observational conditional distribution mixes true causation with spurious context correlation:
$$P(R mid A = a) = sum_x P(R mid A=a, X=x) P(X=x mid A=a) neq P(R mid do(A = a))$$
Example: If policy $pi_0$ only recommends luxury watches ($A$) to wealthy users ($X$), then luxury watches exhibit high purchase rates $P(R mid A)$ not because the watch is appealing, but because wealthy users buy everything.
(3) The Backdoor Adjustment Formula (Pearl’s Do-Calculus):
To evaluate the causal effect of an intervention $do(A = pi_1(X))$, we sever all incoming edges into $A$ ($X notto A$), forcing $A$ to be chosen by candidate policy $pi_1$:
$$P(R mid do(A = a)) = sum_{x in mathcal{X}} P(R mid A = a, X = x) P(X = x)$$
The expected policy value under intervention $pi_1$ is:
$$V(pi_1) = mathbb{E}_{X sim P(X)} left[ sum_{a in mathcal{A}} pi_1(a mid X) mathbb{E}[R mid X, A = a] right]$$
This bridges the gap between observational logging and counterfactual interventional reasoning.
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
深度剖析与工程权衡:① ‘do(A) ≠ 条件 A’是因果推断的核心区分——面试中能准确解释这一点是深度理解的标志。② ‘后门准则给出需要调整的变量’——这是实践中的判据。③ ‘IPS 本质是切断 X→A 的边’——用重加权实现’干预’;这是 IPS 的因果解释。④ ‘随机化是物理上切断边’——最可靠(不需假设)。⑤ ‘敏感性分析’评估未观测混淆——实用且重要(因为’无混淆’无法验证)。⑥ 面试要点——被问’反事实评估的假设’,应给出’因果图 + 混淆 + do(A) vs 条件 A + 后门准则 + IPS/随机化的因果解释 + 敏感性分析‘;能准确区分’do’与’条件’是深度理解的标志。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
In-Depth Analysis & Engineering Trade-offs: ① The Backdoor Criterion requirement: Logging all confounding variables—conditioning on $X$ blocks the backdoor path $A leftarrow X to R$ if and only if all variables that simultaneously influence action choice and user reward are observed and logged; if the online service used user location to choose items but failed to record location in Kafka, unobserved confounding reopens the backdoor, making unbiased counterfactual inference impossible. ② Front-door adjustment when context is unobserved—if context $X$ is partially hidden, but an intermediate mediator $M$ exists (e.g., $A to M to R$, such as user reading the product title), Pearl’s front-door criterion can theoretically recover causal effects without observing $X$. ③ Selection bias as collider conditioning—conditioning on clicks ($C=1$) when evaluating post-click conversion (CVR) treats a collider as observed: $A to C leftarrow U$; conditioning on collider $C$ opens a spurious correlation between $A$ and unobserved user intent $U$, mathematically explaining Sample Selection Bias (SSB) and justifying ESMM. ④ Instrumental Variables (IV) for recommendation—when unobserved confounding is pervasive, using random algorithmic assignment or network jitter as an Instrumental Variable $Z$ ($Z to A to R$ with $Z perp U$) isolates true causal treatment effects. ⑤ Causal DAG validation in system design—drawing explicit causal DAGs during architecture reviews forces engineering teams to identify which features must be captured at inference time to enable future counterfactual modeling. ⑥ Interview takeaway—draw the causal DAG ($X to A, X to R, A to R$), explain why $P(R mid A) neq P(R mid do(A))$ due to the backdoor path $A leftarrow X to R$, derive the Backdoor Adjustment formula, and explain collider bias in CVR modeling.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 把 P(R|A) 当作 P(R|do(A))(混淆未调整)
- ⚠️ 不做敏感性分析(无法评估未观测混淆)
English Pitfalls:
– Confusing observational correlation P(R | A) with interventional causation P(R | do(A)), deploying policies that replicate spurious historical correlations.
– Failing to log critical contextual features that the online ranking algorithm used, permanently opening unblocked backdoor confounding paths.
– Conditioning on post-treatment colliders (e.g., conditioning only on clicked samples in CVR), inducing spurious correlations between items and user intent.
六、高频深度面试追问与预测 (Follow-Up Questions)
- ‘do(A)’与’条件 A’的差异?
- How does conditioning on a collider node (e.g., user click C in A -> C <- U) induce selection bias in downstream conversion modeling?
- 后门准则是什么?
- What properties must an Instrumental Variable (IV) satisfy to recover unbiased causal treatment effects when hidden confounders exist?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
推荐系统离线评估与离线策略评估 (OPE):逆倾向得分 (IPS) 与重要性采样(Off-Policy Evaluation (OPE): IPS, Doubly Robust & Calibration) - 🗺️ 知识图谱模块:
机器学习工程师高频考点导图
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。