【AI 核心深度 M4-043】解释 cross-attention 与 self-attention 的差异与用途(Cross-Attention vs Self-Attention: Structural Differences and Multimodal Roles)深度数理推导与工程落地解析

所属模块:M4 · 序列与 Transformer (Sequences & Transformers) | 专题分类:注意力变体 (Attention Variants (MHA / MQA / GQA)) | 难度等级:Medium

一、核心一句话结论 (One-Sentence Summary)

self-attention 的 Q/K/V 同源(序列内交互);cross-attention 的 Q 来自解码器、K/V 来自编码器(跨序列对齐)。

ADVERTISEMENT · 赞助推荐

Self-attention derives Query, Key, and Value from the same sequence; Cross-attention derives Query from the target sequence and Key/Value from an external conditioning source.

二、核心考点要义 (Key Insights)

  • 📌 self:序列内 token 互相交互,长度相同
  • 📌 cross:解码器查询编码器,长度可不同
  • 📌 cross 是 Enc-Dec 架构的信息桥梁;Dec-only 模型无 cross

English Insights:
– Self-Attention: $Q, K, V = X W_Q, X W_K, X W_V$; models intra-sequence dependencies; square $N times N$ or causal triangular matrix
– Cross-Attention: $Q = Y W_Q, ; K = X W_K, ; V = X W_V$; rectangular attention matrix $N_Y times N_X$; models conditional alignment
– Applications: Encoder-Decoder translation (T5), Diffusion conditioning (Stable Diffusion text-to-image), Multimodal perception (Flamingo)

三、核心数学原理与机理推导 (Mathematical Principles & Derivation)

$$text{self}: Q,K,V=XW^{QKV};qquad text{cross}: Q=Y W^{Q}, K,V=XW^{KV}$$

数学机理:self-attention 的 Q、K、V 都来自同一个序列 X(Q=XW^Q、K=XW^K、V=XW^V),实现序列内的位置间信息交换(token mixing);序列长度前后一致(输出长度=输入长度)。cross-attention 的 Q 来自目标序列 Y(如解码器的隐状态),K/V 来自源序列 X(如编码器的输出):Attn=softmax((YW^Q)(XW^K)ᵀ/√d)(XW^V)。它实现跨序列的信息对齐——解码器每步’查询’源序列的相关部分,这是 Enc-Dec 架构(翻译/摘要/ASR)的信息桥梁。关键差异:(1) 长度不同——Q 长度=|Y|、K/V 长度=|X|,注意力矩阵为 |Y|×|X|(非方阵);(2) 方向性——cross-attention 无因果 mask(源序列已完整可见,无需屏蔽未来);(3) KV cache 复用——K/V 来自编码器,对所有解码步相同,故可只计算一次并缓存(无需随解码增长);这是 Enc-Dec 推理的一个优势。在 Dec-only 模型中的替代——Dec-only 无 cross-attention,源信息通过’拼接进上下文’(self-attention 内部交互)实现;代价是源序列也要被自回归地’读’一遍(更长的序列、更多的 KV cache),且源序列的注意力是因果的(不能双向理解源)。这就是 Enc-Dec 在’源需双向理解’任务上的结构优势所在。多模态中的 cross-attention——VLM 中常用’文本 query 查询图像特征’(如 Flamingo 的 gated cross-attention),把图像作为 K/V 注入 LLM。

📖 查看英文严格数学推导 (English Mathematical Derivation)

Structural Comparison:
Let source sequence be $X in mathbb{R}^{N_x times d_x}$ and target sequence be $Y in mathbb{R}^{N_y times d_y}$.
① Self-Attention:
$Q = X W_Q, quad K = X W_K, quad V = X W_V$.
Attention matrix: $A = text{softmax}left( frac{Q K^T}{sqrt{d}} right) in mathbb{R}^{N_x times N_x}$.
Every token routes information to and from tokens within its own sequence. Computes contextual representations.
② Cross-Attention:
Target acts as Query; Source acts as Key and Value:
$Q = Y W_Q in mathbb{R}^{N_y times d}, quad K = X W_K in mathbb{R}^{N_x times d}, quad V = X W_V in mathbb{R}^{N_x times d}$.
Attention matrix: $A = text{softmax}left( frac{Q K^T}{sqrt{d}} right) in mathbb{R}^{N_y times N_x}$.
Output: $O = A V in mathbb{R}^{N_y times d}$.
Each target token dynamically queries the entire source memory. No causal mask is applied along $N_x$ because the conditioning context is fully available.

四、工业级落地权衡与工程考量 (Industrial Trade-offs)

深度剖析与工程权衡:① KV cache 的差异——self-attention 的 KV cache 随解码增长(每步新增一个 K/V);cross-attention 的 KV cache 固定(源序列的 K/V 只算一次),故 Enc-Dec 在长源 + 短输出的任务(摘要、检索增强生成)上更省显存。② Enc-Dec 与 Dec-only 的计算对比——Dec-only 把源拼进上下文后,源部分也要参与 O(L²) 注意力且被因果读取;Enc-Dec 用编码器双向处理源(一次 O(L²))+ 解码器 cross-attention(O(|Y|·|X|)),在’长源’场景更高效。③ 交叉注意力的可解释性——cross-attention 权重可直接视为’输出对输入的对齐’(翻译中近似对角),比 self-attention 更易解读。④ 多模态融合的三种方式——(a) cross-attention 注入(Flamingo、早期的 VLM);(b) 前缀拼接(把图像特征作为前缀 token 输入 LLM,如 LLaVA);(c) 统一 tokenization(原生多模态,如 Chameleon、Qwen-VL 的部分设计)。三者在训练难度、参数量、灵活性上各有取舍。⑤ 与检索增强的关系——RAG 的’检索结果注入’也可用 cross-attention(RETRO 用 chunked cross-attention 检索邻居)或前缀拼接(更常见);前者可处理更长的检索内容。⑥ 面试要点——被问’cross vs self attention’,应给出’Q/K/V 来源(同源 vs 跨序列)+ 长度(相同 vs 不同)+ mask(因果 vs 无)+ KV cache(增长 vs 固定)‘四维对比,并说明’Dec-only 用前缀拼接替代 cross-attention 的代价’;能提到 VLM 的三种融合方式是加分。

⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)

Serving optimization in Cross-Attention: Source sequence $X$ is processed only once. Its Key and Value tensors ($K_X, V_X$) are precomputed and frozen, allowing the target decoder to autoregressively query static KV states with zero recomputation.

五、常见面试避坑陷阱 (Common Pitfalls & Traps)

  • ⚠️ 以为 cross-attention 需要因果 mask
  • ⚠️ 忽略 cross-attention 的 KV cache 可固定复用

English Pitfalls:
– Applying causal masking along the source dimension in cross-attention; conditioning context is fully visible
– Recomputing cross-attention Key and Value tensors at every autoregressive decoding step instead of caching them

六、高频深度面试追问与预测 (Follow-Up Questions)

  1. Dec-only 模型如何替代 cross-attention?
  2. How does Stable Diffusion use cross-attention to condition image latent denoising on text embeddings?
  3. cross-attention 的 KV cache 如何管理?
  4. Why is the KV-cache of cross-attention static throughout autoregressive decoding while self-attention KV-cache grows dynamically?

七、知识图谱对齐 (Knowledge Graph Anchor)

  • 🔗 关联底层卡片:注意力变体:Multi-Head (MHA)、Multi-Query (MQA) 与 Grouped-Query (GQA) (Attention Variants: MHA, MQA & Grouped-Query Attention (GQA))
  • 🗺️ 知识图谱模块:AI 基础设施工程导图

🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)

本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。

👉 前往 TalentMe 交互式研读本题 (M4-043) →


Discover more from AirSOTA – Air School Of Thoughts AtoZ

Subscribe to get the latest posts sent to your email.