【AI 核心深度 M7-083】解释离线评估为何不能直接用历史数据评估新策略(Explain Why Offline Evaluation Cannot Naively Use Historical Logged Data to Evaluate New Policies)深度数理推导与工程落地解析

所属模块:M7 · 检索、排序与推荐系统 (Retrieval, Ranking & RecSys) | 专题分类:离线评估与 OPE (Offline Evaluation & OPE) | 难度等级:Easy

一、核心一句话结论 (One-Sentence Summary)

历史数据由’旧策略’产生(有偏),故’新策略在历史数据上的表现’不等于’新策略上线后的表现’。

ADVERTISEMENT · 赞助推荐

Historical interaction logs reflect biased observations generated by a past logging policy; naively averaging historical rewards evaluates the old policy rather than the candidate policy, failing to observe how users would respond to novel recommendations.

二、核心考点要义 (Key Insights)

  • 📌 历史数据是’旧策略’选择的(有偏采样)
  • 📌 直接平均历史数据 = 评估旧策略(而非新策略)
  • 📌 新策略可能’推荐历史上没出现过的物品’(无数据)

English Insights:
– Counterfactual unobservability: We only observe user feedback for items the old policy chose to display; user response to items the new policy would choose remains unobserved.
– Selection bias & logged distribution shift: P(Logged | x) reflects the historical algorithm’s biases, not the true global relevance distribution.
– The offline evaluation imperative: Demands Off-Policy Evaluation (OPE) frameworks—such as Inverse Propensity Scoring (IPS) and Doubly Robust estimators—to compute unbiased policy values.

三、核心数学原理与机理推导 (Mathematical Principles & Derivation)

$$text{history}simpi_{text{old}};qquad mathbb{E}{pi)$$}}}nefrac1nsum r_i (text{biased

数学机理:为什么不能直接用历史数据——(1) 历史数据是’旧策略’产生的——历史日志记录了’旧策略 π_old 推荐了什么、用户如何反应’;关键——用户只对’被展示的’物品有反馈;未展示的物品没有数据(’反事实’未知)。(2) 直接平均的偏置——若用’历史数据的平均奖励’评估新策略 π_new:(a) 这实际上是’评估 π_old 的表现’(因为数据由 π_old 产生);(b) 新策略若’推荐了历史上没被展示的物品’ → 那些物品没有数据 → 无法评估;(c) 偏差来源——’选择偏差’(旧策略的选择)+ ‘展示偏差’(只有被展示的有反馈)。(3) 支持集不匹配(support mismatch)——(a) 新策略 π_new 可能给’π_old 从未展示的物品’很高概率 → 但这些物品在历史数据中不存在;(b) 故’无法从历史数据估计新策略的表现’(无数据);(c) 这是 OPE 的根本困难。(4) ‘离线 NDCG’也有偏——(a) 用’历史点击’作为’相关性标签’评估新策略的排序;但点击受位置影响(位置偏置)→ ‘离线 NDCG’衡量的是’新策略与旧策略的一致性’(而非’真实相关性’);(b) 故’离线 NDCG 高’可能只是’新策略像旧策略’。正确的做法(OPE)——(1) 逆倾向加权(IPS)——用’1/倾向’加权历史数据,使其’模拟’随机策略的数据(见 IPS 题);(2) 直接方法(DM)——用模型预测’新策略下的奖励’(依赖模型准确性);(3) 双稳健(DR)——结合 DM 与 IPS(见 DR 题);(4) 随机化数据——用’随机策略’收集数据(无偏,但损害体验);(5) 在线 A/B——最终验证。前提假设(OPE 的三大假设)——(a) 无混淆(unconfoundedness)——’展示’只依赖’可观测的特征’(无隐藏混淆);(b) 正性/重叠(positivity/overlap)——新策略会选的物品在历史中’有非零的展示概率’(否则无法估计);(c) SUTVA(稳定单位处理值假设)——一个物品的反馈不受其他物品影响(无干扰);注意——(b) 是最易违反的(新策略选的新物品从未展示)。与其他问题的关系——(a) 与’位置偏置’(同一问题);(b) 与’离线-在线不一致’(LTR 题);(c) 与’因果推断’(OPE 是因果推断的应用)。实践建议——(a) 不要直接用历史数据平均评估新策略;(b) 用 IPS/DR(配权重截断);(c) 保留随机化数据(金标准);(d) 检查’支持集重叠’(新策略选的物品在历史中有多少概率被展示);(e) 离线只做粗筛、A/B 定胜负。度量——(a) OPE 估计的偏差(vs 在线 A/B);(b) 支持集重叠度;(c) 估计的方差。

📖 查看英文严格数学推导 (English Mathematical Derivation)

Mathematical & Counterfactual Formulation: The Offline Policy Evaluation Gap.

(1) The Logged Bandit Feedback Problem:
Let user/context be $x sim P(x)$, item action be $a in mathcal{A}$, and reward be $r(x, a)$.
Historical logged dataset $mathcal{D} = {(x_i, a_i, r_i, p_i)}_{i=1}^N$ was collected under logging policy $pi_0$:
$$a_i sim pi_0(a mid x_i), quad p_i = pi_0(a_i mid x_i)$$$$r_i text{ is observed ONLY for chosen action } a_i$$
We want to evaluate the expected performance of a novel candidate policy $pi_1$:
$$V(pi_1) = mathbb{E}_{x sim P(x), a sim pi_1(a mid x)} [r(x, a)]$$

(2) Why Direct Empirical Averaging Fails (Naive Evaluation):
Suppose we filter historical logs where the old action matches the new action ($a_i = pi_1(x_i)$) and compute the average reward:
$$hat{V}_{text{naive}}(pi_1) = frac{sum_{i=1}^N r_i cdot mathbb{I}(a_i = pi_1(x_i))}{sum_{i=1}^N mathbb{I}(a_i = pi_1(x_i))}$$
Mathematical Failure: This estimator is heavily biased because the condition $a_i = pi_1(x_i)$ is conditioned on the old policy choosing $a_i$. Actions that $pi_0$ chose frequently are over-represented; actions $pi_0$ rarely explored are under-represented or have zero samples.

(3) The Common Support Condition (Positivity):
Unbiased counterfactual estimation is mathematically possible if and only if the logging policy has non-zero probability for every action the new policy might take:
$$forall x, a: quad pi_1(a mid x) > 0 implies pi_0(a mid x) > 0$$
If $pi_0$ strictly never displayed action $a^*$, historical data contains zero information regarding how users respond to $a^*$.

四、工业级落地权衡与工程考量 (Industrial Trade-offs)

深度剖析与工程权衡:① ‘历史数据由旧策略产生’是根本问题——故’直接平均 = 评估旧策略’;面试中能指出这一点是深度理解的标志。② ‘支持集不匹配’是 OPE 的根本困难——新策略选的新物品无数据;故’正性假设’最易违反。③ ‘离线 NDCG 也有偏’——它衡量’与旧策略的一致性’;这是易被忽视的陷阱。④ ‘随机化数据是金标准’——小流量随机可提供无偏数据;代价是体验损失。⑤ ‘三大假设’需明确——无混淆/正性/SUTVA;违反则 OPE 失效。⑥ 面试要点——被问’为什么不能用历史数据评估新策略’,应给出’历史数据由旧策略产生(选择偏差)+ 支持集不匹配 + 离线 NDCG 也有偏 + 正确做法(IPS/DR/随机化/A-B)+ 三大假设‘;能指出’离线 NDCG 衡量的是与旧策略的一致性’是深度理解的标志。

⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)

In-Depth Analysis & Engineering Trade-offs: ① The necessity of exploration traffic for offline evaluation—if production systems run deterministic greedy policies ($pi_0(a) in {0, 1}$), the positivity condition is violated for 99% of actions, making offline policy evaluation mathematically impossible; reserving 2%–5% of production traffic for randomized exploration logging enables offline counterfactual evaluation. ② The Direct Method (Model-based) vs. Inverse Propensity Scoring (Importance Sampling)—Direct Method trains a regression model $hat{r}(x, a)$ and predicts rewards for $pi_1$; it has low variance but suffers high bias if the reward model is misspecified. IPS is strictly unbiased but suffers high variance. ③ Logging propensity scores ($p_i$) at inference time—to evaluate policies offline using IPS, production systems must log the exact probability $p_i = pi_0(a_i mid x_i)$ into Kafka alongside the interaction event; reconstructing propensity scores retrospectively is impossible. ④ Position bias confounding—in search ranking, actions are slates of items; a relevant item may have received zero clicks because $pi_0$ placed it at rank 20; counterfactual models must account for position propensities alongside item selection propensities. ⑤ Offline replay testing vs. Full A/B testing—unbiased offline policy evaluation acts as a rigorous gatekeeper, eliminating 70% of candidate models before spending cloud budget on live A/B tests. ⑥ Interview takeaway—formalize counterfactual unobservability ($r$ observed only for $pi_0(a)$), prove why naive averaging is biased, state the Common Support condition $pi_1 > 0 implies pi_0 > 0$, and motivate Off-Policy Evaluation (OPE).

五、常见面试避坑陷阱 (Common Pitfalls & Traps)

  • ⚠️ 直接用历史数据的平均奖励评估新策略
  • ⚠️ 忽略’正性假设’(新策略选的物品无历史数据)

English Pitfalls:
– Evaluating a new ranking policy by filtering historical click logs for matching items and calculating empirical click rates, committing severe selection bias.
– Deploying deterministic greedy logging policies that assign zero probability to un-ranked items, permanently preventing counterfactual offline evaluation.
– Failing to log the execution propensity p_i = pi_0(a_i | x_i) at online inference time, rendering historical logs useless for importance sampling.

六、高频深度面试追问与预测 (Follow-Up Questions)

  1. 什么是’支持集不匹配’?
  2. Why is the Common Support (Positivity) assumption mandatory for unbiased Off-Policy Evaluation?
  3. 为什么’离线 NDCG’也有偏?
  4. What architectural logging patterns capture exact action propensities p_i without impacting online query latency?

七、知识图谱对齐 (Knowledge Graph Anchor)

  • 🔗 关联底层卡片:推荐系统离线评估与离线策略评估 (OPE):逆倾向得分 (IPS) 与重要性采样 (Off-Policy Evaluation (OPE): IPS, Doubly Robust & Calibration)
  • 🗺️ 知识图谱模块:机器学习工程师高频考点导图

🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)

本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。

👉 前往 TalentMe 交互式研读本题 (M7-083) →


Discover more from AirSOTA – Air School Of Thoughts AtoZ

Subscribe to get the latest posts sent to your email.