所属模块:
M7 · 检索、排序与推荐系统 (Retrieval, Ranking & RecSys)| 专题分类:深度推荐模型 (Deep Recommendation Models)| 难度等级:Hard
一、核心一句话结论 (One-Sentence Summary)
CTR 与 CVR 的样本空间不同(CVR 只在点击样本上训练);ESMM 在全样本空间建模 ‘pCTCVR = pCTR × pCVR’,用 CTCVR 的标签训练。
ESMM models post-click conversion rate (CVR) across the entire impression space by mathematically decomposing Click-Through and Conversion Rate (CTCVR) as pCTCVR = pCTR * pCVR, training on unselected impression logs to eliminate sample selection bias.
二、核心考点要义 (Key Insights)
- 📌 问题:CVR 只在’点击’样本上训练(样本空间有偏)
- 📌 SSB:训练分布(点击样本)≠ 推理分布(全样本)
- 📌 ESMM:建模 pCTCVR = pCTR × pCVR,在全样本上训练,共享嵌入
English Insights:
– Sample Selection Bias (SSB): Traditional CVR models are trained only on clicked samples, but are evaluated across all impressed candidates at serving time.
– Data Sparsity (DS): Conversion events are 100x rarer than clicks; training CVR in isolation leads to severe parameter underfitting.
– Multiplicative probability decomposition: p(Click, Conversion | x) = p(Click | x) * p(Conversion | Click, x); enables joint training over all impression logs.
– Shared embedding table: The CVR network and CTR network share input embedding layers, transferring rich click signals into the sparse conversion domain.
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$underbrace{p(text{CTCVR})}{text{entire space}}=underbrace{p(text{CTR})}$$}}timesunderbrace{p(text{CVR})}_{text{entire space}};qquad text{train on all samples
数学机理:两个问题——(1) 样本选择偏差(Sample Selection Bias,SSB)——(a) CVR(转化率) 的定义是’在点击的条件下转化的概率’:p(conversion | click);(b) 故 CVR 模型只能用’点击’样本训练(因为’未点击’的样本不知道’是否会转化’);(c) 问题——训练分布(点击样本)与推理分布(全样本)不一致;(d) 后果——(i) 模型学到的是’点击样本上的规律’(而这些样本是’有偏’的——被点击的物品与未被点击的物品不同);(ii) 推理时(对全样本预估)表现下降;(e) 本质——这是’分布偏移’问题(训练分布 ≠ 推理分布)。(2) 数据稀疏(DS)——点击样本远少于曝光样本(CTR 常 1%~10%)→ CVR 的训练数据更少 → 更易过拟合。ESMM(Entire Space Multi-task Model,Ma 等 2018,阿里) 的解法——(1) 核心恒等式——利用概率的乘法关系:p(CTCVR) = p(CTR) × p(CVR),其中 CTCVR = ‘点击且转化’(即’曝光后的转化’);(2) 在全样本空间建模——(a) CTR 塔——预测 p(CTR)(用全样本训练,因为’是否点击’对全样本已知);(b) CVR 塔——预测 p(CVR)(但不用 CVR 的标签训练!);(c) CTCVR 塔——把两个塔的输出相乘:p(CTCVR) = p(CTR) × p(CVR),用 CTCVR 的标签(全样本已知) 训练;(d) 关键——CVR 塔虽然只在’点击样本’上有意义,但它的参数通过’CTCVR 塔的损失’在’全样本’上被训练(因为 CTCVR 的标签对全样本已知);(e) 这绕过了’CVR 只能点击样本训练’的限制——CVR 塔在全样本空间上被间接训练。(3) 共享嵌入——两个塔共享特征嵌入(CTR 塔有大量数据,可帮助 CVR 塔学习更好的嵌入);效果——缓解 CVR 的数据稀疏。(4) 推理——预估 CVR 时,直接用 CVR 塔的输出(或 CTR×CVR 得到 CTCVR)。优势——(a) 消除 SSB(CVR 塔在全样本空间训练);(b) 缓解数据稀疏(共享嵌入 + 全样本训练);(c) 端到端(一个模型);(d) 简单有效。实证——ESMM 在阿里的场景显著提升 CVR 预估(AUC 提升数点);成为’CVR 预估’的标准方法。后续发展——(a) ESM2(多任务版本,加入’是否购买’等多个目标);(b) HM3(Hierarchical Multi-task)(建模’点击→加购→购买’的层级);(c) MMoE + ESMM 组合(多任务 + 全样本空间)。与其他问题的关系——(a) 与’样本选择偏差’(推荐/广告的普遍问题);(b) 与’多任务学习’(共享嵌入);(c) 与’因果推断’(SSB 是因果问题——’点击’是一个’后处理变量’)。实践建议——(a) CVR 预估用 ESMM(消除 SSB);(b) 共享嵌入(缓解稀疏);(c) 多级转化(点击→加购→购买) 用 HM3;(d) 监控’训练分布 vs 推理分布’的差异;(e) 与多任务模型组合。度量——(a) CVR 的 AUC;(b) 校准(预估 CVR vs 实际);(c) 在线 GMV/转化率。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Mathematical & Probabilistic Formulation: ESMM Architecture (Ma et al., Alibaba, 2018).
(1) The Two Fundamental Pathologies in CVR Modeling:
– Sample Selection Bias (SSB):
Conversion rate is defined mathematically as the conditional probability of conversion given click:
$$ptext{CVR} = P(y=1 mid z=1, x)$$
where $z in {0, 1}$ is click, $y in {0, 1}$ is conversion, and $x$ is candidate features.
Traditional models train exclusively on clicked samples $mathcal{D}_{text{clicked}} = {(x, y) mid z=1}$. However, during online serving, the ranking model scores all impressed candidates $mathcal{D}_{text{impressed}} = {x}$. Because $P(x mid z=1) neq P(x)$, the training distribution diverges drastically from the serving distribution.
– Data Sparsity (DS): Conversion samples account for only $0.01%text{–}0.1%$ of total traffic, causing standalone CVR embeddings to overfit.
(2) The Multiplicative Probability Decomposition:
Define Click-Through & Conversion Rate ($ptext{CTCVR}$):
$$ptext{CTCVR} = P(z=1, y=1 mid x) = P(z=1 mid x) times P(y=1 mid z=1, x) = ptext{CTR} times ptext{CVR}$$
Notice that both $ptext{CTR}$ and $ptext{CTCVR}$ can be observed and trained directly over the entire impression space $mathcal{D}_{text{impressed}}$:
– $z=1$ indicates clicked impression; $z=0$ indicates unclicked impression.
– $y=1$ indicates converted impression; $y=0$ indicates non-converted impression.
(3) Loss Function Formulation:
ESMM constructs two neural networks: a CTR tower predicting $f_{text{CTR}}(x)$ and a CVR tower predicting $f_{text{CVR}}(x)$. Both share the same feature embedding table.
$$mathcal{L}(x, z, y) = mathcal{L}_{text{BCE}}(z, f_{text{CTR}}(x)) + mathcal{L}_{text{BCE}}big( z cdot y, , f_{text{CTR}}(x) times f_{text{CVR}}(x) big)$$
$$mathcal{L} = – z ln f_{text{CTR}}(x) – (1 – z) ln(1 – f_{text{CTR}}(x)) – (z cdot y) ln(f_{text{CTR}} f_{text{CVR}}) – (1 – z cdot y) ln(1 – f_{text{CTR}} f_{text{CVR}})$$
At test time, the model directly outputs $f_{text{CVR}}(x)$ as the predicted conversion probability.
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
深度剖析与工程权衡:① ‘SSB 的本质是分布偏移’——训练分布(点击样本)≠ 推理分布(全样本);面试中能指出这一点是深度理解的标志。② ‘用 pCTCVR = pCTR × pCVR 绕过限制’是 ESMM 的核心洞察——它让 CVR 塔在全样本空间被间接训练。③ ‘共享嵌入缓解稀疏’——CTR 塔的大数据帮助 CVR 塔;这是多任务的经典价值。④ ‘SSB 是因果问题’——’点击’是’后处理变量’(post-treatment);这与因果推断的’选择偏差’同源。⑤ ‘HM3 建模多级转化’——点击→加购→购买 的层级关系;比单一 CVR 更细。⑥ 面试要点——被问’ESMM 解决什么’,应给出’SSB(CVR 只在点击样本训练 → 分布偏移)+ pCTCVR = pCTR × pCVR(在全样本上训练)+ 共享嵌入‘;能指出’SSB 本质是分布偏移’是深度理解的标志。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
In-Depth Analysis & Engineering Trade-offs: ① Why not train CVR by predicting CTCVR directly and dividing by CTR?—If $ptext{CVR} = frac{ptext{CTCVR}}{ptext{CTR}}$, whenever an item has a very low predicted click probability ($ptext{CTR} to 0$), numerical division leads to extreme volatility ($ptext{CVR} > 1$ or infinity); ESMM models $ptext{CVR}$ as a bounded neural output in $[0, 1]$ and multiplies forward, maintaining complete numerical stability. ② Shared embedding representation transfer—because the CTR tower receives gradients from hundreds of millions of click events, the shared embedding representations are exceptionally rich; this directly benefits the CVR tower, solving the Data Sparsity dilemma. ③ Handling delayed conversions (Attribution Window)—conversions often occur hours or days after the initial click; closing attribution windows too early creates false negative labels in $ptext{CTCVR}$; production pipelines implement delayed feedback modeling or importance weighting on conversion logs. ④ Multi-stage conversion funnels (ESM2)—e-commerce funnels span multiple actions: Impression $to$ Click $to$ Add-to-Cart $to$ Purchase; ESM2 extends ESMM to a directed acyclic graph (DAG) of conditional probabilities, modeling intermediate conversion steps simultaneously. ⑤ Inference speed—online ad auctions require evaluating $eCPM = text{Bid} times ptext{CTR} times ptext{CVR}$; both towers evaluate concurrently on GPU in $< 2text{ ms}$. ⑥ Interview takeaway—formulate the two pathologies (SSB and DS), write down the probability identity $P(z=1, y=1 mid x) = P(z=1 mid x) P(y=1 mid z=1, x)$, detail the joint loss function over the entire space, and explain the numerical division failure mode.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ CVR 只在点击样本上训练(分布偏移)
- ⚠️ 不用共享嵌入(CVR 数据稀疏)
English Pitfalls:
– Attempting to compute CVR at inference time by dividing pCTCVR by pCTR, resulting in catastrophic numerical instability and division-by-zero errors when pCTR is small.
– Training standalone CVR models purely on clicked samples without recognizing that online serving encounters unclicked impression distributions (SSB).
– Failing to share embedding parameters between the CTR and CVR networks, forfeiting the primary solution to the Data Sparsity problem.
六、高频深度面试追问与预测 (Follow-Up Questions)
- 什么是’样本选择偏差(SSB)’?
- Why is multiplying pCTR and pCVR during training numerically superior to estimating pCTCVR and dividing by pCTR?
- ESMM 如何’在全样本上训练 CVR’?
- How does the ESM2 model extend ESMM to multi-step sequential conversion funnels (Impression -> Click -> Cart -> Buy)?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
深度排序模型演进:Wide & Deep、DeepFM 二阶特征交叉、DCN 与 DIN 注意力(Deep Ranking Models: Wide & Deep, DeepFM, DCN & DIN) - 🗺️ 知识图谱模块:
工业级系统设计导图
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。