所属模块:
M3 · 深度学习基础 (Deep Learning Foundations)| 专题分类:分布式训练 (Distributed Training Basics)| 难度等级:Easy
一、核心一句话结论 (One-Sentence Summary)
数据并行复制模型、切分数据;张量并行切分单层权重;流水线并行切分层(按 stage);三者常组合为 3D 并行。
Data Parallelism splits batches across workers (All-Reduce); Tensor Parallelism shards weight matrices intra-node (All-Gather/Reduce-Scatter); Pipeline Parallelism partitions layers sequentially inter-node (P2P).
二、核心考点要义 (Key Insights)
- 📌 DP 通信量 ∝ 参数量(all-reduce 梯度)
- 📌 TP 通信量 ∝ 激活(每层 all-reduce),需高带宽(NVLink)
- 📌 PP 通信量小(仅 stage 边界)但有流水线气泡
English Insights:
– Data Parallelism (DDP/FSDP): replicates or shards model parameters; scales with batch size; communication $propto$ model parameters
– Tensor Parallelism (Megatron-LM): shards individual linear/attention matrices across GPUs; requires ultra-high bandwidth (NVLink); communication $propto$ activations
– Pipeline Parallelism (GPipe): partitions layers across nodes; point-to-point communication; suffers from idle pipeline bubble overhead
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$text{DP}: text{replicate} theta, text{split} mathcal{D};quad text{TP}: text{split} W;quad text{PP}: text{split layers}$$
数学机理:数据并行(DP)——每个设备持有完整模型副本、处理不同数据分片,反向结束后用 all-reduce 同步梯度(取平均)。优点是实现简单、通信与参数量成正比(一次 all-reduce);缺点是每设备需存完整模型 + 优化器状态,模型大到单卡放不下时失效。张量并行(TP)——把单层的权重矩阵切分到多设备(如按列切 W₁、按行切 W₂),每层前向/反向需 all-reduce 或 all-gather 中间激活。优点是可切分任意大的层、显存按 TP 度线性下降;缺点是每层都要通信、通信量 ∝ 激活大小、延迟敏感,故必须在同一节点内用 NVLink/高速互联(跨节点带宽不足会严重拖慢)。流水线并行(PP)——按层切分成若干 stage,每设备持有连续几层,micro-batch 依次流过各 stage(类似工厂流水线)。优点是通信量小(仅 stage 边界传激活)、可跨节点;缺点是流水线气泡——首尾 stage 在填充/排空阶段空闲,设备利用率下降(气泡比例约 (P−1)/(M+P−1),P 为 stage 数、M 为 micro-batch 数)。3D 并行即 DP×TP×PP 组合(如 Megatron-LM 的配置)。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Taxonomy and Communication Patterns:
① Data Parallelism (DDP / FSDP):
Batch dimension $B$ is split across $N$ GPUs ($B/N$ per GPU). Each worker executes independent forward/backward passes. Gradients are synchronized via All-Reduce ($2 times |text{params}|$ bytes transferred). Scalable across nodes; bottlenecked by batch size lower bounds.
② Tensor Parallelism (TP / Megatron-LM):
For MLP: shard $W_1$ column-wise ($d times frac{4d}{TP}$) and $W_2$ row-wise ($frac{4d}{TP} times d$). Output: $Y = text{GeLU}(X W_1) W_2$. Requires exactly 2 All-Reduce operations per Transformer layer. Communication is high-frequency and latency-sensitive, strictly requiring intra-node NVLink ($900text{GB/s}$).
③ Pipeline Parallelism (PP / GPipe, 1F1B):
Partitions $L$ layers across $P$ pipeline stages. Stage $k$ receives activations from stage $k-1$ and sends outputs to stage $k+1$. Communication is lightweight peer-to-peer (P2P), suitable across slow inter-node InfiniBand networks. Introduces pipeline bubble fraction: $F_{text{bubble}} = frac{P – 1}{M + P – 1}$, where $M$ is the number of micro-batches.
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
深度剖析与工程权衡:① 三者的正交性——它们切分的是不同维度:DP 切数据、TP 切权重、PP 切深度;理论上可任意组合(受互联拓扑约束)。实践配置原则是’节点内 TP、节点间 PP、最外层 DP‘,因为 TP 需高带宽、PP 带宽需求低、DP 通信可与计算重叠。② 通信量对比——DP 每步一次 all-reduce(∝参数量);TP 每层两次通信(∝激活);PP 每 micro-batch 传一次边界激活。故 TP 对带宽最敏感。③ 流水线气泡的缓解——(a) 增大 micro-batch 数 M(气泡 ∝1/M)、(b) 用 interleaved/1F1B 调度(交错前后向)、(c) 用 zero-bubble pipeline 等新调度;但 M 增大会增加激活显存(与检查点冲突)。④ ZeRO/FSDP 的位置——ZeRO 是’数据并行的显存优化’(分片优化器状态/梯度/参数),可视为 DP 的改进版;它不与 TP/PP 冲突,常叠加使用(如 ZeRO-3 + TP + PP)。⑤ 序列并行(SP)——在 TP 基础上把 LayerNorm/dropout 等’非矩阵乘’部分沿序列维度切分,进一步降低激活显存;是长序列训练的常用补充。⑥ 面试要点——被问’如何训练一个放不下的模型’,应给出’先 ZeRO/FSDP(省显存)→ 再 TP(切层内)→ 再 PP(切层间)→ 配合 DP(扩吞吐)‘的决策顺序,并说明’TP 需 NVLink’这一硬约束。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
3D Parallelism Composition Rule: Map Tensor Parallelism intra-node across 8 GPUs via NVLink; map Pipeline Parallelism across nodes via InfiniBand; wrap outer layer in Data Parallelism (FSDP / ZeRO) to scale across thousands of GPUs.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 认为 TP 可以随意跨节点(带宽不足会严重拖慢)
- ⚠️ 忽略流水线气泡对利用率的影响
English Pitfalls:
– Running Tensor Parallelism across nodes over standard Ethernet/InfiniBand, causing extreme communication bottlenecks
– Setting the number of pipeline micro-batches $M$ too small, causing the pipeline bubble to consume $>50%$ of compute time
六、高频深度面试追问与预测 (Follow-Up Questions)
- 为什么 TP 要求高带宽而 PP 不要求?
- How does the 1F1B (One-Forward-One-Backward) schedule reduce pipeline memory footprint compared to GPipe?
- 什么是流水线气泡?如何缓解?
- Why does Megatron-LM pair column-parallel linear projection with row-parallel projection in attention and MLPs?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
分布式并行基础:DDP 数据并行、Ring All-Reduce 与 ZeRO 显存切分(Distributed Training: DDP, Ring All-Reduce & ZeRO Memory) - 🗺️ 知识图谱模块:
AI 基础设施工程导图
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。