【AI 核心深度 M4-023】解释 attention 的 mask 机制(causal / padding / prefix-LM)(Attention Masking Mechanisms: Causal, Padding, and Prefix-LM Masks)深度数理推导与工程落地解析

所属模块:M4 · 序列与 Transformer (Sequences & Transformers) | 专题分类:Transformer 架构解剖 (Transformer Architecture Anatomy) | 难度等级:Easy

一、核心一句话结论 (One-Sentence Summary)

causal mask 屏蔽未来位置实现自回归;padding mask 屏蔽补齐位;prefix-LM 允许前缀双向、后缀因果。

ADVERTISEMENT · 赞助推荐

Masks set attention logits of invalid or prohibited tokens to $-infty$ prior to softmax, implementing causal autoregression, padding isolation, or hybrid prefix-LM dynamics.

二、核心考点要义 (Key Insights)

  • 📌 causal:下三角可看,上三角屏蔽(自回归必需)
  • 📌 padding:屏蔽补齐 token,防止注意力落到无效位置
  • 📌 prefix-LM:前缀双向 + 后缀因果,兼顾理解与生成

English Insights:
– Implementation: $A = text{softmax}left(frac{Q K^T}{sqrt{d}} + Mright) V$, where $M_{ij} = 0$ for allowed tokens and $-infty$ (or $-10^9$) for prohibited tokens
– Causal Mask: lower-triangular binary matrix ($M_{ij} = -infty$ for $j > i$), preventing tokens from looking into the future
– Padding Mask: sets $M_{ij} = -infty$ wherever key token $j$ is a <pad> padding token
– Prefix-LM Mask: fully bidirectional within the prompt prefix, strictly causal within generated suffix tokens

三、核心数学原理与机理推导 (Mathematical Principles & Derivation)

$$mathrm{mask}: text{logits}leftarrowtext{logits}+M,quad M_{ij}=begin{cases}0&text{allowed} -infty&text{masked}end{cases}$$

数学机理:mask 通过在 softmax 前给 logits 加上一个矩阵 M 实现:允许的位置加 0、禁止的位置加 −∞(实践中用 −1e9 等极大负数以避免数值问题),使 softmax 后这些位置的权重为 0。三类 mask:(1) causal(因果)mask——下三角矩阵(位置 i 只能看 j≤i),这是自回归语言模型的必需项(生成第 t 个 token 时不能看到 t 之后的内容,否则信息泄漏、训练目标失去意义)。(2) padding mask——序列补齐到统一长度时,补齐位(padding token)不应被关注;mask 把它们屏蔽。注意:不能仅靠’把补齐位设为 0 向量’来解决——因为注意力会对 0 向量计算出一个非零的 logit(0 与任何 query 的点积为 0,softmax 后会分到权重),故必须用 mask 显式屏蔽。(3) prefix-LM mask——前缀部分双向(可看整段前缀)、后缀部分因果;用于’输入需双向理解、输出需自回归生成’的任务(如指令微调中让指令部分双向、回答部分因果),是 Enc-Dec 结构的’单塔’替代。为什么用 −∞ 而非 0——mask 的目的是让权重恰好为 0;加 0 不改变 softmax 结果,必须加 −∞ 使 exp(−∞)=0。实践中用 −1e9 而非真 −inf,以避免 −inf 与 0 相乘产生 NaN,以及低精度下 inf 的运算问题。

📖 查看英文严格数学推导 (English Mathematical Derivation)

Mathematical Formulations:
① Causal Mask (Autoregressive Mask):
Enforces temporal causality: token $i$ can only attend to tokens $j le i$.
$M_{text{causal}}[i, j] = begin{cases} 0 & text{if } j le i \ -infty & text{if } j > i end{cases}$.
In floating-point hardware (FP16/BF16), $-infty$ is implemented as $-10^9$ or $-65504$. In softmax: $e^{-infty} = 0$. Prohibited future tokens receive exact zero attention weight: $P_{ij} = 0$.
② Padding Mask:
When batching sequences of variable length, shorter sequences are padded with “ tokens.
$M_{text{pad}}[i, j] = begin{cases} 0 & text{if token } j ne text{} \ -infty & text{if token } j = text{} end{cases}$.
Prevents dummy padding tokens from polluting context representations.
③ Prefix-LM Mask (Non-causal Prefix + Causal Generation):
Let prompt prefix tokens have length $P$, and generation tokens have length $G$.
$M_{text{prefix}}[i, j] = begin{cases} 0 & text{if } i le P text{ and } j le P text{ (bidirectional prompt)} \ 0 & text{if } i > P text{ and } j le i text{ (causal response)} \ -infty & text{otherwise} end{cases}$.
Combines the bidirectional comprehension quality of BERT with the autoregressive generation capacity of GPT.

四、工业级落地权衡与工程考量 (Industrial Trade-offs)

深度剖析与工程权衡:① mask 的实现细节——通常用’加性 mask’(logits + M)而非’乘性 mask’,因为 softmax 前是加性的;Flash Attention 内部用 mask 的块级判断来跳过被屏蔽的块(提升效率)。② 全屏蔽行的风险——若某行所有位置都被屏蔽(如 padding 位自己作为 query),softmax 会得 0/0 = NaN;故需保证每行至少有一个可关注位置(或加保护)。这是实现中最常见的 NaN 来源之一。③ prefix-LM 的实用价值——它在指令微调中很常见:把 system+user 部分设为双向、assistant 回复部分设为因果,使模型对指令的理解更充分(双向)而不破坏生成的自回归性。④ causal mask 与 KV cache 的关系——正因为有 causal mask,推理时才能用 KV cache 增量解码(只算新 token 的 attention,复用历史的 K/V);若无因果性(如双向编码器),则每次改动都需重算。⑤ 与’文档内注意力’的对比——检索增强生成中,若把多个文档拼进上下文,可用’块对角 mask’让各文档内部双向、文档间隔离;这是 mask 灵活性的体现。⑥ 面试要点——被问’mask 怎么用’,应给出三类 mask 的用途与公式,并强调’必须用 −∞ 而非 0‘与’全屏蔽行会 NaN‘这两个实现细节;能提到 prefix-LM 与块对角 mask 是加分。

⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)

Hardware optimization: In modern fused attention kernels (FlashAttention-2), causal masking is computed in hardware via thread block row-column index comparison ($j > i$) without ever materializing an actual $N times N$ mask matrix in GPU memory.

五、常见面试避坑陷阱 (Common Pitfalls & Traps)

  • ⚠️ 用 0 作为 mask 值(softmax 后权重不为 0)
  • ⚠️ 忽略全屏蔽行导致的 0/0 NaN

English Pitfalls:
– Using $0$ as a mask value instead of $-infty$; setting logits to $0$ results in equal non-zero attention weights ($e^0 = 1$)
– Encountering all-masked rows where an entire sequence is padding, causing $text{softmax}(-infty) = 0/0 = text{NaN}$

六、高频深度面试追问与预测 (Follow-Up Questions)

  1. 为什么 mask 要用 −∞(或极大负数)而不是 0?
  2. How does FlashAttention implement causal masking in CUDA without creating an actual $N times N$ mask tensor in memory?
  3. padding mask 为什么不能靠’补齐为 0 向量’解决?
  4. What causes all-masked row NaNs in multi-head attention when using padding masks?

七、知识图谱对齐 (Knowledge Graph Anchor)

  • 🔗 关联底层卡片:Transformer 核心架构解剖:Pre-LN vs Post-LN 与多头注意力 (Transformer Block Deep Dive: Pre-LN vs Post-LN & MHA)
  • 🗺️ 知识图谱模块:大语言模型全景图谱

🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)

本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。

👉 前往 TalentMe 交互式研读本题 (M4-023) →


Discover more from AirSOTA – Air School Of Thoughts AtoZ

Subscribe to get the latest posts sent to your email.