所属模块:
M4 · 序列与 Transformer (Sequences & Transformers)| 专题分类:长上下文 (Long Context Extensions & Scaling)| 难度等级:Medium
一、核心一句话结论 (One-Sentence Summary)
大部分层用滑窗(局部)省算力,少量全局 token 或全注意力层保证信息能跨长距离流动。
Hybrid attention restricts the majority of tokens to attend within a local sliding window of size $W$ while allowing select global tokens to attend across the entire sequence, reducing complexity from $O(L^2)$ to $O(L cdot W)$ while preserving long-range routing.
二、核心考点要义 (Key Insights)
- 📌 纯滑窗感受野受限(L 层 × w)
- 📌 全局 token(或每 N 层一个全注意力)打通长距离
- 📌 Mistral/Gemma-2 用’SWA + 全注意力交替’
English Insights:
– Sliding Window Attention (SWA): each token only attends to its nearest $W$ neighboring tokens; stack of $K$ layers expands effective receptive field to $K times W$
– Global Token Mechanism: designated tokens (e.g., [CLS], question tokens, or initial tokens) attend to all tokens and are attended to by all tokens, acting as an information hub
– Popularized by Longformer, BigBird, and Mistral/Mixtral; slashes compute and KV cache size while retaining cross-document connectivity
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$text{mix}: text{SWA layers}+text{global tokens / full-attn layers};qquad text{receptive field}=O(L)$$
数学机理:纯滑动窗口的问题——每个 token 只看前后 w 个位置,信息要跨越超过 w 的距离需逐层传递(L 层网络最多传 L·w);对需要长距离依赖的任务(如跨章节指代、长代码的符号定义与使用),纯滑窗要么需要极深的网络、要么无法覆盖。混合设计的两种形式:(1) 全局 token(global tokens)——保留少量’全局 token’(如序列开头的若干 token、或专门的 [GLOBAL] 标记),所有位置都可关注它们、它们也可关注所有位置;这样信息可通过全局 token 作为’中转站’跨越任意距离(两步:局部 → 全局 token → 目标位置)。Longformer 的’局部窗口 + 全局 token’即此设计。(2) 层间交替(interleaved)——大部分层用滑窗(省算力),每 N 层插入一层全注意力(提供长距离通路);Mistral-7B、Gemma-2 采用此设计(如 Gemma-2 每 2 层一个滑窗层、每 2 层一个全注意力层)。效果——感受野仍为 O(L)(因为全注意力层可直达任意位置),但总计算量大幅降低(全注意力层只占 1/N)。与 attention sink 的关系——全局 token 常位于序列开头,与 attention sink 位置重合,故’保留初始 token 作为 sink’与’提供全局通路’是同一设计的两面(StreamingLLM 的’保留 sink + 滑窗’即此)。工程优势——滑窗层的 KV cache 只需保留最近 w 个(环形缓冲,显存 O(w)),故长上下文推理的显存大幅降低;全注意力层仍需全 KV(但层数少,总量可控)。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Mathematical Mechanism: 1. Sliding Window Attention: For sequence $x_1, dots, x_L$ and window size $W$, the causal attention mask $M_{ij}$ is defined as: $$M_{ij} = begin{cases} 0 & text{if } 0 le i – j le W \ -infty & text{otherwise} end{cases}$$ The attention computation for each token requires only $W$ dot products rather than $L$, reducing total complexity per layer from $O(L^2)$ to $O(L cdot W)$. Across $N$ layers, the theoretical receptive field at layer $N$ spans $N times W$ tokens. 2. Global Token Cross-Attention: Let $mathcal{G}$ be the set of global token indices (with $|mathcal{G}| = G ll L$). The mask is updated such that: $$M_{ij} = 0 quad text{if } i in mathcal{G} lor j in mathcal{G} lor (0 le i – j le W)$$ Global tokens aggregate information from the entire sequence in layer $l$ and broadcast it back to all local tokens in layer $l+1$, creating an $O(1)$ token-distance communication highway across arbitrary sequence lengths.
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
深度剖析与工程权衡:① ‘信息中转’的效率——通过全局 token 传递信息需两步(局部→全局→局部),比全注意力的单步直达效率低,但成本也低得多;这是一种’成本-效率’的折中。② KV cache 的混合管理——滑窗层的 KV 是 O(w)(环形缓冲)、全注意力层是 O(L);混合架构的显存介于两者之间,且大部分层是滑窗层,故总量显著低于全注意力。③ 与位置编码的交互——滑窗层若用 RoPE 的’远距离衰减’,则其局部性被进一步强化;部分实现只对全注意力层施加位置编码或使用不同的位置策略。④ 与 MoE/SSM 的组合——混合架构可进一步与 MoE(稀疏 FFN)与 SSM(线性递归)组合,形成’全注意力 + 滑窗 + SSM + MoE’的多层混合(如 Jamba);这是当前’效率优先’架构的主流方向。⑤ 训练的一致性——混合架构需在预训练阶段就使用(否则推理时的注意力模式与训练不一致);故它是’架构选择’而非’推理优化’。⑥ 面试要点——被问’长上下文如何省算力’,应给出’滑窗(省算力)+ 全局 token 或全注意力层(保长距离通路)‘的混合设计,并说明’感受野仍为 O(L) 但计算大幅降低’与’滑窗层 KV 是 O(w)’;能联系到 attention sink 与 MoE/SSM 组合是深度理解的标志。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
Deep Dive & Engineering Trade-offs: ① KV Cache Memory Savings: In pure sliding window decoding (Mistral), the KV cache can be maintained as a circular buffer of fixed size $W$. Once $t > W$, oldest tokens are overwritten, keeping KV cache VRAM bounded at $O(W)$ regardless of total generation length. ② The Attention Sink Requirement: Pure circular sliding window buffers without initial tokens suffer catastrophic perplexity collapse (as discovered in StreamingLLM). Preserving the first 4 initial tokens ($S=4$) as global attention sinks stabilizes attention scores completely. ③ Layer-Alternating Designs: Rather than placing global tokens inside every layer, modern architectures alternate between local sliding window layers and full global attention layers (e.g., 3 sliding window layers followed by 1 full attention layer in Gemma-2 and Qwen-2.5). ④ Task Degradation: Pure sliding window models struggle with needle-in-a-haystack retrieval if the needle lies outside the receptive field and global tokens fail to compress the relevant facts. ⑤ Interview Strategy: Contrast receptive field expansion through depth vs explicit global tokens, explain the circular buffer KV cache implementation, and highlight the necessity of initial sink tokens.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 认为纯滑窗足以处理长距离依赖
- ⚠️ 忽略混合架构必须在预训练阶段就使用
English Pitfalls:
– Assuming sliding window attention limits model receptive field to strictly $W$ (receptive field grows linearly with depth $N times W$)
– Evicting the first few initial tokens from a circular sliding window buffer (causes perplexity collapse due to attention sink loss)
– Using sliding window attention for tasks requiring exact, dense long-range cross-referencing without global layers
六、高频深度面试追问与预测 (Follow-Up Questions)
- 为什么纯滑窗不够?
- Why does alternating sliding window layers with full attention layers (like Gemma-2) outperform uniform sliding window attention?
- 全局 token 与全注意力层如何选择?
- How does StreamingLLM combine attention sinks with sliding window attention to achieve infinite context streaming?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
长上下文扩展:NTK-Aware 插值、YaRN 与大海捞针 (Needle-in-Haystack) 评估(Long Context Extension: NTK Interpolation, YaRN & Retrieval) - 🗺️ 知识图谱模块:
大语言模型全景图谱
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。