【AI 核心深度 M7-063】解释 ESMM 解决样本选择偏差(SSB)的思路(Explain How the Entire Space Multi-Task Model (ESMM) Eliminates Sample Selection Bias (SSB) and Data Sparsity in CVR)深度数理推导与工程落地解析

所属模块:M7 · 检索、排序与推荐系统 (Retrieval, Ranking & RecSys) | 专题分类:深度推荐模型 (Deep Recommendation Models) | 难度等级:Hard

一、核心一句话结论 (One-Sentence Summary)

CTR 与 CVR 的样本空间不同(CVR 只在点击样本上训练);ESMM 在全样本空间建模 ‘pCTCVR = pCTR × pCVR’,用 CTCVR 的标签训练。

ADVERTISEMENT · 赞助推荐

ESMM models post-click conversion rate (CVR) across the entire impression space by mathematically decomposing Click-Through and Conversion Rate (CTCVR) as pCTCVR = pCTR * pCVR, training on unselected impression logs to eliminate sample selection bias.

二、核心考点要义 (Key Insights)

  • 📌 问题:CVR 只在’点击’样本上训练(样本空间有偏)
  • 📌 SSB:训练分布(点击样本)≠ 推理分布(全样本)
  • 📌 ESMM:建模 pCTCVR = pCTR × pCVR,在全样本上训练,共享嵌入

English Insights:
– Sample Selection Bias (SSB): Traditional CVR models are trained only on clicked samples, but are evaluated across all impressed candidates at serving time.
– Data Sparsity (DS): Conversion events are 100x rarer than clicks; training CVR in isolation leads to severe parameter underfitting.
– Multiplicative probability decomposition: p(Click, Conversion | x) = p(Click | x) * p(Conversion | Click, x); enables joint training over all impression logs.
– Shared embedding table: The CVR network and CTR network share input embedding layers, transferring rich click signals into the sparse conversion domain.

三、核心数学原理与机理推导 (Mathematical Principles & Derivation)

$$underbrace{p(text{CTCVR})}{text{entire space}}=underbrace{p(text{CTR})}$$}}timesunderbrace{p(text{CVR})}_{text{entire space}};qquad text{train on all samples

数学机理:两个问题——(1) 样本选择偏差(Sample Selection Bias,SSB)——(a) CVR(转化率) 的定义是’在点击的条件下转化的概率’:p(conversion | click);(b) 故 CVR 模型只能用’点击’样本训练(因为’未点击’的样本不知道’是否会转化’);(c) 问题——训练分布(点击样本)与推理分布(全样本)不一致;(d) 后果——(i) 模型学到的是’点击样本上的规律’(而这些样本是’有偏’的——被点击的物品与未被点击的物品不同);(ii) 推理时(对全样本预估)表现下降;(e) 本质——这是’分布偏移’问题(训练分布 ≠ 推理分布)。(2) 数据稀疏(DS)——点击样本远少于曝光样本(CTR 常 1%~10%)→ CVR 的训练数据更少 → 更易过拟合。ESMM(Entire Space Multi-task Model,Ma 等 2018,阿里) 的解法——(1) 核心恒等式——利用概率的乘法关系:p(CTCVR) = p(CTR) × p(CVR),其中 CTCVR = ‘点击且转化’(即’曝光后的转化’);(2) 在全样本空间建模——(a) CTR 塔——预测 p(CTR)(用全样本训练,因为’是否点击’对全样本已知);(b) CVR 塔——预测 p(CVR)(但不用 CVR 的标签训练!);(c) CTCVR 塔——把两个塔的输出相乘:p(CTCVR) = p(CTR) × p(CVR),用 CTCVR 的标签(全样本已知) 训练;(d) 关键——CVR 塔虽然只在’点击样本’上有意义,但它的参数通过’CTCVR 塔的损失’在’全样本’上被训练(因为 CTCVR 的标签对全样本已知);(e) 这绕过了’CVR 只能点击样本训练’的限制——CVR 塔在全样本空间上被间接训练。(3) 共享嵌入——两个塔共享特征嵌入(CTR 塔有大量数据,可帮助 CVR 塔学习更好的嵌入);效果——缓解 CVR 的数据稀疏。(4) 推理——预估 CVR 时,直接用 CVR 塔的输出(或 CTR×CVR 得到 CTCVR)。优势——(a) 消除 SSB(CVR 塔在全样本空间训练);(b) 缓解数据稀疏(共享嵌入 + 全样本训练);(c) 端到端(一个模型);(d) 简单有效。实证——ESMM 在阿里的场景显著提升 CVR 预估(AUC 提升数点);成为’CVR 预估’的标准方法。后续发展——(a) ESM2(多任务版本,加入’是否购买’等多个目标);(b) HM3(Hierarchical Multi-task)(建模’点击→加购→购买’的层级);(c) MMoE + ESMM 组合(多任务 + 全样本空间)。与其他问题的关系——(a) 与’样本选择偏差’(推荐/广告的普遍问题);(b) 与’多任务学习’(共享嵌入);(c) 与’因果推断’(SSB 是因果问题——’点击’是一个’后处理变量’)。实践建议——(a) CVR 预估用 ESMM(消除 SSB);(b) 共享嵌入(缓解稀疏);(c) 多级转化(点击→加购→购买) 用 HM3;(d) 监控’训练分布 vs 推理分布’的差异;(e) 与多任务模型组合。度量——(a) CVR 的 AUC;(b) 校准(预估 CVR vs 实际);(c) 在线 GMV/转化率。

📖 查看英文严格数学推导 (English Mathematical Derivation)

Mathematical & Probabilistic Formulation: ESMM Architecture (Ma et al., Alibaba, 2018).

(1) The Two Fundamental Pathologies in CVR Modeling:
– Sample Selection Bias (SSB):
Conversion rate is defined mathematically as the conditional probability of conversion given click:
$$ptext{CVR} = P(y=1 mid z=1, x)$$
where $z in {0, 1}$ is click, $y in {0, 1}$ is conversion, and $x$ is candidate features.
Traditional models train exclusively on clicked samples $mathcal{D}_{text{clicked}} = {(x, y) mid z=1}$. However, during online serving, the ranking model scores all impressed candidates $mathcal{D}_{text{impressed}} = {x}$. Because $P(x mid z=1) neq P(x)$, the training distribution diverges drastically from the serving distribution.
– Data Sparsity (DS): Conversion samples account for only $0.01%text{–}0.1%$ of total traffic, causing standalone CVR embeddings to overfit.

(2) The Multiplicative Probability Decomposition:
Define Click-Through & Conversion Rate ($ptext{CTCVR}$):
$$ptext{CTCVR} = P(z=1, y=1 mid x) = P(z=1 mid x) times P(y=1 mid z=1, x) = ptext{CTR} times ptext{CVR}$$
Notice that both $ptext{CTR}$ and $ptext{CTCVR}$ can be observed and trained directly over the entire impression space $mathcal{D}_{text{impressed}}$:
– $z=1$ indicates clicked impression; $z=0$ indicates unclicked impression.
– $y=1$ indicates converted impression; $y=0$ indicates non-converted impression.

(3) Loss Function Formulation:
ESMM constructs two neural networks: a CTR tower predicting $f_{text{CTR}}(x)$ and a CVR tower predicting $f_{text{CVR}}(x)$. Both share the same feature embedding table.
$$mathcal{L}(x, z, y) = mathcal{L}_{text{BCE}}(z, f_{text{CTR}}(x)) + mathcal{L}_{text{BCE}}big( z cdot y, , f_{text{CTR}}(x) times f_{text{CVR}}(x) big)$$
$$mathcal{L} = – z ln f_{text{CTR}}(x) – (1 – z) ln(1 – f_{text{CTR}}(x)) – (z cdot y) ln(f_{text{CTR}} f_{text{CVR}}) – (1 – z cdot y) ln(1 – f_{text{CTR}} f_{text{CVR}})$$
At test time, the model directly outputs $f_{text{CVR}}(x)$ as the predicted conversion probability.

四、工业级落地权衡与工程考量 (Industrial Trade-offs)

深度剖析与工程权衡:① ‘SSB 的本质是分布偏移’——训练分布(点击样本)≠ 推理分布(全样本);面试中能指出这一点是深度理解的标志。② ‘用 pCTCVR = pCTR × pCVR 绕过限制’是 ESMM 的核心洞察——它让 CVR 塔在全样本空间被间接训练。③ ‘共享嵌入缓解稀疏’——CTR 塔的大数据帮助 CVR 塔;这是多任务的经典价值。④ ‘SSB 是因果问题’——’点击’是’后处理变量’(post-treatment);这与因果推断的’选择偏差’同源。⑤ ‘HM3 建模多级转化’——点击→加购→购买 的层级关系;比单一 CVR 更细。⑥ 面试要点——被问’ESMM 解决什么’,应给出’SSB(CVR 只在点击样本训练 → 分布偏移)+ pCTCVR = pCTR × pCVR(在全样本上训练)+ 共享嵌入‘;能指出’SSB 本质是分布偏移’是深度理解的标志。

⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)

In-Depth Analysis & Engineering Trade-offs: ① Why not train CVR by predicting CTCVR directly and dividing by CTR?—If $ptext{CVR} = frac{ptext{CTCVR}}{ptext{CTR}}$, whenever an item has a very low predicted click probability ($ptext{CTR} to 0$), numerical division leads to extreme volatility ($ptext{CVR} > 1$ or infinity); ESMM models $ptext{CVR}$ as a bounded neural output in $[0, 1]$ and multiplies forward, maintaining complete numerical stability. ② Shared embedding representation transfer—because the CTR tower receives gradients from hundreds of millions of click events, the shared embedding representations are exceptionally rich; this directly benefits the CVR tower, solving the Data Sparsity dilemma. ③ Handling delayed conversions (Attribution Window)—conversions often occur hours or days after the initial click; closing attribution windows too early creates false negative labels in $ptext{CTCVR}$; production pipelines implement delayed feedback modeling or importance weighting on conversion logs. ④ Multi-stage conversion funnels (ESM2)—e-commerce funnels span multiple actions: Impression $to$ Click $to$ Add-to-Cart $to$ Purchase; ESM2 extends ESMM to a directed acyclic graph (DAG) of conditional probabilities, modeling intermediate conversion steps simultaneously. ⑤ Inference speed—online ad auctions require evaluating $eCPM = text{Bid} times ptext{CTR} times ptext{CVR}$; both towers evaluate concurrently on GPU in $< 2text{ ms}$. ⑥ Interview takeaway—formulate the two pathologies (SSB and DS), write down the probability identity $P(z=1, y=1 mid x) = P(z=1 mid x) P(y=1 mid z=1, x)$, detail the joint loss function over the entire space, and explain the numerical division failure mode.

五、常见面试避坑陷阱 (Common Pitfalls & Traps)

  • ⚠️ CVR 只在点击样本上训练(分布偏移)
  • ⚠️ 不用共享嵌入(CVR 数据稀疏)

English Pitfalls:
– Attempting to compute CVR at inference time by dividing pCTCVR by pCTR, resulting in catastrophic numerical instability and division-by-zero errors when pCTR is small.
– Training standalone CVR models purely on clicked samples without recognizing that online serving encounters unclicked impression distributions (SSB).
– Failing to share embedding parameters between the CTR and CVR networks, forfeiting the primary solution to the Data Sparsity problem.

六、高频深度面试追问与预测 (Follow-Up Questions)

  1. 什么是’样本选择偏差(SSB)’?
  2. Why is multiplying pCTR and pCVR during training numerically superior to estimating pCTCVR and dividing by pCTR?
  3. ESMM 如何’在全样本上训练 CVR’?
  4. How does the ESM2 model extend ESMM to multi-step sequential conversion funnels (Impression -> Click -> Cart -> Buy)?

七、知识图谱对齐 (Knowledge Graph Anchor)

  • 🔗 关联底层卡片:深度排序模型演进:Wide & Deep、DeepFM 二阶特征交叉、DCN 与 DIN 注意力 (Deep Ranking Models: Wide & Deep, DeepFM, DCN & DIN)
  • 🗺️ 知识图谱模块:工业级系统设计导图

🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)

本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。

👉 前往 TalentMe 交互式研读本题 (M7-063) →


Discover more from AirSOTA – Air School Of Thoughts AtoZ

Subscribe to get the latest posts sent to your email.