所属模块:
M7 · 检索、排序与推荐系统 (Retrieval, Ranking & RecSys)| 专题分类:离线评估与 OPE (Offline Evaluation & OPE)| 难度等级:Hard
一、核心一句话结论 (One-Sentence Summary)
OPE 需三大假设(无混淆/正性/SUTVA);常见失败:正性违反(新策略选新物品)、混淆、干扰、非平稳。
Off-Policy Evaluation relies on three foundational causal assumptions—Unconfoundedness, Common Support (Positivity), and SUTVA; violations caused by unobserved features, zero-probability actions, interference, or non-stationarity trigger catastrophic evaluation failures.
二、核心考点要义 (Key Insights)
- 📌 无混淆:展示只依赖可观测特征
- 📌 正性/重叠:新策略的动作在历史中有非零展示概率
- 📌 SUTVA:无干扰;另有’非平稳’(分布随时间变化)
English Insights:
– Unconfoundedness (No Hidden Confounders): Action selection depends strictly on observed features logged in data: r indep a | x.
– Common Support / Positivity: Every action the candidate policy can propose must have had a strictly positive probability of being chosen by the logging policy: pi_0(a | x) > 0.
– SUTVA (No Interference): One user’s recommendation treatment does not alter or interfere with the reward observed by another user.
– Non-Stationarity: Temporal drift in user preferences or catalog inventory breaks the assumption that historical response distributions mirror future reality.
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$text{assumptions}: text{unconfoundedness}, text{positivity}, text{SUTVA};qquad text{violation}Rightarrowtext{fail}$$
数学机理:OPE 的三大假设——(1) 无混淆(unconfoundedness / ignorability)——’展示’(动作)只依赖可观测的特征(无隐藏的混淆变量);违反的后果——估计有偏(因为’隐藏因素’同时影响展示与奖励);例子——’高价值用户更可能被展示广告’(若’用户价值’未观测,则估计有偏)。(2) 正性/重叠(positivity / overlap)——新策略会选的每个动作在历史数据中都有非零的展示概率(π_old(a|x)>0);违反的后果——(a) 无法估计(该动作无数据 → 外推);(b) 权重爆炸(π_old→0);这是最易违反的假设——新策略常’推荐历史上没出现过的物品’(如新物品、长尾);检测——检查’新策略的动作在历史中的展示率’(若有动作的展示率为 0 → 违反)。(3) SUTVA(Stable Unit Treatment Value Assumption)——(a) 无干扰(no interference)——一个物品的反馈不受’其他被展示物品’影响;(b) 无隐藏变体——同一动作只有一个版本;违反的后果——估计有偏;例子——推荐列表的’位置’影响点击(同一物品在不同位置点击率不同 → ‘动作’不只是’展示哪个物品’,还包括’位置’)。(4) 非平稳(non-stationarity)——(a) 用户兴趣/物品质量/市场随时间变化;(b) 历史数据的分布与当前不同;(c) 违反的后果——估计不适用(历史不代表现在);(d) 对策——(i) 用’近期数据’;(ii) 加’时间特征’;(iii) 监控漂移。其他失败模式——(a) 倾向模型错误(π_old 估计不准 → IPS 有偏);(b) 模型错误(DR 的模型不准 → 修正项方差大);(c) 奖励的定义问题(奖励延迟、稀疏);(d) 特征缺失(历史日志未记录某些特征 → 无法建模);(e) ‘动作空间大’(推荐的动作空间百万级 → 倾向极难估计)。检测与缓解——(a) 正性检查——统计’新策略动作的历史展示率’;(b) 重叠度——计算’有效样本量(ESS)’;(c) 倾向模型诊断(AUC/校准);(d) 模拟实验(用合成数据验证 OPE 的准确性);(e) 与 A/B 对比(最终验证);(f) 限制动作空间(只在’重叠区’评估);(g) 用’随机化数据’补充(打破混淆与正性问题)。实践建议——(a) 先检查假设(尤其正性);(b) 保留随机化数据(小流量随机是’金标准’);(c) 限制评估范围(只在重叠区);(d) 倾向模型要诊断;(e) 与 A/B 对比(最终验证);(f) 注意非平稳(用近期数据);(g) 离线只做粗筛。度量——(a) 正性违反率;(b) ESS;(c) 倾向模型的 AUC/校准;(d) OPE vs A/B 的偏差。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Mathematical & Causal Formulation: Foundational OPE Assumptions.
(1) Assumption 1: Unconfoundedness (Conditional Ignorability):
$${r(x, a)}_{a in mathcal{A}} perp A mid X$$
Failure Mode: An unobserved variable $U$ (e.g., user seasonal mood, external marketing campaign, sudden celebrity tweet) influenced which item the old policy recommended AND influenced user click probability. Because $U$ is missing from feature log $X$, the backdoor path $A leftarrow U to R$ remains open, inducing severe causal bias.
(2) Assumption 2: Common Support / Overlap (Positivity):
$$forall x in mathcal{X}, , a in mathcal{A}: quad pi_1(a mid x) > 0 implies pi_0(a mid x) ge epsilon > 0$$
Failure Mode: Candidate policy $pi_1$ proposes recommending newly ingested items or novel category crosses that the historical policy $pi_0$ strictly never explored ($pi_0(a mid x) = 0$). The importance ratio $frac{pi_1(a)}{pi_0(a)} = frac{>0}{0}$ is mathematically undefined. OPE cannot evaluate actions never taken in history.
(3) Assumption 3: SUTVA (Stable Unit Treatment Value Assumption):
$$R_i(A_1, A_2, dots, A_N) = R_i(A_i)$$
Failure Mode (Marketplace Interference): If policy $pi_1$ recommends a limited-inventory luxury hotel, User 1 booking it prevents User 2 from booking it. In two-sided ridesharing, dispatching driver to Rider 1 removes the driver from Rider 2. One user’s assignment directly corrupts another user’s potential outcome.
(4) Assumption 4: Stationarity (Distribution Stability):
$$P_{text{future}}(R mid X, A) = P_{text{past}}(R mid X, A)$$
Failure Mode: User preferences drift post-holiday or after competitor product launches.
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
深度剖析与工程权衡:① ‘正性最易违反’是关键——新策略常推荐历史未展示的物品;面试中能指出这一点是深度理解的标志。② ‘位置是动作的一部分’——若’动作’只定义为’展示哪个物品’而忽略位置,则 SUTVA 违反(位置影响点击)。③ ‘随机化数据是金标准’——它能同时解决’混淆’与’正性’问题(因为随机策略对所有动作有正概率)。④ ‘非平稳’常被忽视——历史数据不代表现在;故需用近期数据。⑤ ‘倾向模型诊断’必需——π_old 估计不准则 IPS 有偏。⑥ 面试要点——被问’OPE 什么时候失效’,应给出’三大假设(无混淆/正性/SUTVA)+ 非平稳 + 其他失败(倾向错/模型错/奖励定义/特征缺失)+ 检测与缓解‘;能指出’正性最易违反’是深度理解的标志。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
In-Depth Analysis & Engineering Trade-offs: ① Deterministic logging policies break Positivity—if production runs a deterministic ranker ($a_t = argmax f(x)$), 99.9% of catalog items have $pi_0(a) = 0$ for a given context; Positivity is violated for virtually all candidate policies; reserving a continuous 2% traffic slice for randomized or epsilon-greedy logging is mandatory for OPE. ② Hidden confounders in multi-device sessions—a user browses shoes on desktop at work, then opens the mobile app at home and buys; logging mobile interactions without linking desktop session features introduces severe unobserved confounding; cross-device identity stitching closes the backdoor path. ③ Detecting Positivity violations—evaluating the fraction of candidate policy mass placed on zero-propensity actions: $alpha_{text{uncovered}} = mathbb{E}_{x}[sum_{a: pi_0(a mid x)=0} pi_1(a mid x)]$; if uncovered mass exceeds 5%, OPE must be rejected. ④ Synthetic actions and content generation—in generative recommendation or LLM prompt variations, the new policy outputs novel tokens never generated by the old policy; Positivity is 100% violated; OPE must rely on learned reward models (Direct Method) rather than importance sampling. ⑤ Marketplace cannibalization adjustments—when evaluating inventory-constrained systems, counterfactual simulators must model global inventory state transitions. ⑥ Interview takeaway—state the three core assumptions (Unconfoundedness, Positivity, SUTVA), detail how each assumption breaks in real-world recommender systems, explain why deterministic logging destroys Positivity, and outline diagnostic checks.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 不检查正性(新策略动作无历史数据)
- ⚠️ 忽略’位置是动作的一部分’(SUTVA 违反)
English Pitfalls:
– Attempting to perform importance sampling on candidate policies that propose new items, ignoring severe positivity violations (division by zero).
– Applying standard OPE to ride-sharing or marketplace inventory search without accounting for SUTVA violations caused by shared physical supply.
– Evaluating OPE across seasonal holiday boundaries (e.g., using Black Friday logs to evaluate a January shopping policy), violating stationarity.
六、高频深度面试追问与预测 (Follow-Up Questions)
- 哪个假设最易违反?
- How does randomized exploration traffic guarantee the Positivity (Common Support) condition for downstream offline policy evaluation?
- 违反假设时如何检测?
- What statistical tests detect the presence of unobserved confounders in historical recommendation click logs?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
推荐系统离线评估与离线策略评估 (OPE):逆倾向得分 (IPS) 与重要性采样(Off-Policy Evaluation (OPE): IPS, Doubly Robust & Calibration) - 🗺️ 知识图谱模块:
机器学习工程师高频考点导图
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。