所属模块:
M3 · 深度学习基础 (Deep Learning Foundations)| 专题分类:分布式训练 (Distributed Training Basics)| 难度等级:Hard
一、核心一句话结论 (One-Sentence Summary)
3D 并行 = 数据并行 × 张量并行 × 流水线并行;序列并行把 LayerNorm/dropout 等非 matmul 部分沿序列维切分。
3D Parallelism combines Data (DP), Tensor (TP), and Pipeline (PP) parallelism; Sequence Parallelism shards activation-heavy operations (LayerNorm, Dropout) along the sequence dimension.
二、核心考点要义 (Key Insights)
- 📌 3D 并行按’数据/权重/深度’三个维度同时切分
- 📌 SP 解决 TP 后激活仍沿序列维冗余的问题
- 📌 配置原则:节点内 TP+SP、节点间 PP、最外层 DP
English Insights:
– 3D Parallelism composition: Total GPUs $= text{DP} times text{TP} times text{PP}$; maps dimensions to cluster physical network hierarchy
– Sequence Parallelism (Megatron-SP): splits sequence tokens across TP ranks in non-matmul layers, reducing activation memory by factor $text{TP}$
– Context Parallelism (Ring Attention / DeepSpeed Ulysses): shards sequence dimension globally across ranks using Ring Attention to train 1M+ token contexts
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$text{total GPUs}=text{DP}timestext{TP}timestext{PP};qquad text{SP}: text{split along sequence dim}$$
数学机理:3D 并行是三种并行维度的组合:DP(数据并行,扩吞吐、通信 ∝参数量)+ TP(张量并行,切层内权重、通信 ∝激活、需 NVLink)+ PP(流水线并行,切层、通信小但有气泡)。总 GPU 数 = DP×TP×PP。为什么需要 SP:TP 把权重矩阵切分后,matmul 的激活也被切分(如按列切分后每卡只有部分输出),但非 matmul 的部分(LayerNorm、dropout、残差相加)在 TP 中仍是每卡完整计算同一份激活——即这些算子的激活沿序列维冗余复制在 TP 组内。序列越长,这份冗余越大(激活 ∝ L×d)。序列并行(SP) 的做法是:把 LayerNorm/dropout 等沿序列维切分到 TP 组的各卡上,使每卡只处理 L/TP 个 token 的归一化;而在 matmul 前用 all-gather 把序列拼回来。关键优化:SP 的 all-gather 与 TP 的 all-reduce 可以合并为一次通信(Megatron 的 SP 实现即如此),使 SP 几乎不增加通信开销,却把激活显存降低 TP 倍。更进一步:SP 与 Flash Attention 结合(Megatron 的 SP 变体)可进一步降低注意力激活。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Mathematical Formulations:
① The 3D Mapping Principle:
– TP (Tensor Parallelism): Shards GEMMs. Communication $propto$ activations. Map strictly inside the same node (NVLink, $900text{GB/s}$).
– PP (Pipeline Parallelism): Shards layers. Peer-to-peer communication of boundary activations. Map across nodes over InfiniBand.
– DP (Data Parallelism / FSDP): Shards batches. Map across remaining nodes.
② Sequence Parallelism in Megatron-LM (Korthikanti et al., 2022):
Standard Tensor Parallelism only shards $Q, K, V$ projections and MLP projections; LayerNorm, Dropout, and residual additions remain duplicated across all TP ranks, redundantly storing activations of shape $(B, S, D)$.
– Megatron-SP Innovation: Notice that LayerNorm and Dropout are element-wise or token-independent along sequence length $S$. Split the sequence dimension $S$ across $TP$ workers: length becomes $S / text{TP}$.
– Replace the forward All-Reduce with a Reduce-Scatter (reducing partial matmul sums into sequence shards) and replace the backward All-Reduce with an All-Gather. Total communication volume is mathematically unchanged, but activation memory is slashed by factor $text{TP}$!
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
深度剖析与工程权衡:① 激活显存的主导地位——长序列训练时激活显存 ∝ L×d×层数,常超过参数显存;SP 是降低激活的关键手段之一(与检查点、Flash Attention 并列)。② 配置原则的由来——节点内 NVLink 带宽高 → 放 TP+SP(通信密集);节点间 IB 带宽低 → 放 PP(通信少);最外层 DP(可与计算重叠)。这是’按通信需求匹配带宽层次’的经典设计。③ 与 ZeRO 的组合——大模型常用 ‘TP + SP + PP + ZeRO-1’(ZeRO-1 只分片优化器状态,通信代价小);或 ‘FSDP + TP + SP’。④ 长上下文的完整方案——SP(省激活)+ GQA/MQA(省 KV cache)+ Flash Attention(省注意力激活)+ RoPE 位置插值(支持更长)+ 上下文并行(context parallel,把序列维切到多卡)。⑤ 通信复杂度的现实——3D 并行的调试与调优复杂(需平衡各维度、避免某维成为瓶颈);工业上常用 Megatron-LM + DeepSpeed 的组合模板,而非从零实现。⑥ 面试要点——被问’训练 100B 模型怎么配并行’,应给出’TP=8(节点内)+ SP + PP=若干(跨节点)+ DP(最外)+ ZeRO-1/FSDP‘的具体方案并解释每维的通信特征;这是分布式训练的高阶问题。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
Long-Context Evolution: Megatron-SP is tied to TP degree (typically $le 8$). For ultra-long contexts ($128text{K}-1text{M}$ tokens), use Context Parallelism (DeepSpeed Ulysses / Ring Attention), which shards the sequence globally across arbitrary numbers of nodes.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 认为 TP 之后就无需 SP(非 matmul 部分仍有激活冗余)
- ⚠️ 把 SP 与上下文并行(CP)混为一谈
English Pitfalls:
– Confusing Megatron Sequence Parallelism (which shards LayerNorm/Dropout within TP) with Context Parallelism (which shards entire attention across independent nodes)
– Attempting to scale Tensor Parallelism beyond 8 GPUs across slow InfiniBand switches, causing massive latency bottlenecks
六、高频深度面试追问与预测 (Follow-Up Questions)
- 为什么 TP 之后还需要 SP?
- How does Megatron Sequence Parallelism replace All-Reduce with Reduce-Scatter and All-Gather without increasing communication volume?
- SP 与 TP 的通信如何合并?
- How does DeepSpeed Ulysses use All-to-All communication to parallelize attention along the sequence dimension?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
分布式并行基础:DDP 数据并行、Ring All-Reduce 与 ZeRO 显存切分(Distributed Training: DDP, Ring All-Reduce & ZeRO Memory) - 🗺️ 知识图谱模块:
AI 基础设施工程导图
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。