所属模块:
M4 · 序列与 Transformer (Sequences & Transformers)| 专题分类:模型压缩与蒸馏 (Model Compression & Distillation)| 难度等级:Easy
一、核心一句话结论 (One-Sentence Summary)
非结构化(单个权重,稀疏但硬件难加速)、结构化(整通道/头/层,硬件友好)、半结构化(N:M 稀疏,硬件支持)。
Model pruning removes redundant parameters across three distinct granularities: unstructured pruning maximizes theoretical sparsity but requires sparse kernels, semi-structured 2:4 pruning matches NVIDIA Ampere Sparse Tensor Cores, and structured pruning removes entire channels or heads for immediate hardware acceleration.
二、核心考点要义 (Key Insights)
- 📌 非结构化:粒度最细、压缩率高,但需专门硬件/库才能加速
- 📌 结构化:直接减小矩阵形状,通用硬件即可加速
- 📌 半结构化(2:4):NVIDIA 稀疏张量核心支持,兼顾两者
English Insights:
– Unstructured pruning: zeroes out individual weights with lowest magnitude; highest compression ratio with zero accuracy drop, but yields zero speedup on standard dense matrix hardware
– Semi-structured 2:4 sparsity: enforces exactly 2 non-zero values for every 4 consecutive elements; directly accelerated by NVIDIA Ampere/Hopper Sparse Tensor Cores for a hardware $2times$ speedup
– Structured pruning: removes complete attention heads, FFN neurons, or transformer layers; directly shrinks dense tensor dimensions, achieving immediate speedup and memory reduction on any hardware
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$text{unstructured}: {w_i=0};qquad text{structured}: text{prune channels/heads/layers};qquad text{semi}: 2{:}4$$
数学机理:三种粒度。(1) 非结构化剪枝(unstructured)——按单个权重剪(如把绝对值最小的 90% 权重置 0);压缩率最高(可达 90%+ 稀疏),模型变成稀疏矩阵。问题——通用 GPU 对稀疏矩阵的加速有限(稠密张量核心无法直接利用稀疏性;需专门的稀疏库如 cuSPARSE,但实际加速常远低于稀疏率)。故’稀疏 90%’常只带来很小的实际加速。(2) 结构化剪枝(structured)——按结构单元剪:剪掉整个通道(channel)、注意力头(head)、甚至整层(layer);结果是更小的稠密矩阵,故通用硬件可直接加速(矩阵变小、FLOPs 真实下降)。代价——同样的’参数量减少’下,结构化剪枝对精度的损害通常大于非结构化(因为剪掉的是整个功能单元)。如何决定剪哪些——用重要性准则:(a) 幅度(权重的 L1/L2 范数);(b) 梯度/二阶信息(如 Taylor 展开、OBD/OBS 用 Hessian);(c) 激活(如通道激活的方差、BN 的缩放因子 γ,Network Slimming 用 γ 作为重要性);(d) 可学习(用门控/掩码,训练时学出哪些该剪)。(3) 半结构化(semi-structured / N:M 稀疏)——要求每 M 个连续权重中恰好有 N 个非零(如 2:4,即每 4 个保留 2 个);NVIDIA Ampere 及以后的稀疏张量核心原生支持,可带来约 2 倍的理论加速。这是’兼顾压缩率与硬件加速’的方案,已成为工业界主流(如 2:4 稀疏)。其他——(a) 非结构化 + 硬件(未来若有稀疏加速硬件);(b) 块稀疏(按块剪,硬件友好)。实践——’先剪枝再微调‘(prune then finetune)是标准流程;也有’训练时稀疏’(如 sparse from scratch、Lottery Ticket)。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Mathematical Mechanism: 1. Unstructured Pruning: For weight tensor $W in mathbb{R}^{d_1 times d_2}$ and binary mask $M in {0, 1}^{d_1 times d_2}$: $$W_{text{sparse}} = W odot M, quad M_{ij} = mathbb{I}(|W_{ij}| ge tau)$$ Non-zero elements are scattered randomly across the matrix. Standard dense hardware (Tensor Cores) cannot skip zero multiplications efficiently without custom sparse matrix formats (CSR/CSC), resulting in slower execution due to memory index lookup overhead. 2. Semi-Structured 2:4 Sparsity (NVIDIA Sparse Tensor Cores): In every group of 4 contiguous values in a row, exactly 2 values must be zero: $$forall k: sum_{i=4k}^{4k+3} M_{i} = 2$$ NVIDIA Ampere architecture compresses the 4 values into 2 FP16 values along with a 2-bit metadata index. The hardware Sparse Tensor Core fetches compressed weights from memory, doubling memory bandwidth efficiency and doubling math execution throughput ($2times$ FLOPs). 3. Structured Pruning: Removes entire blocks or dimensions: $$W_{text{pruned}} = W_{[:d_1′, :d_2′]} in mathbb{R}^{d_1′ times d_2′} quad (d_1′ < d_1, d_2' < d_2)$$ Reduces dense GEMM matrix dimensions directly, requiring no custom kernels or hardware support.
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
深度剖析与工程权衡:① ‘压缩率 ≠ 加速比’是核心教训——非结构化剪枝的稀疏率与加速比严重脱节(因访存不规则、需索引开销);故工业界更看重结构化/半结构化(有真实硬件支持)。这与稀疏注意力、MoE 的’FLOPs 降低不等于加速’是同一类问题。② N:M 稀疏的工业地位——2:4 稀疏在 NVIDIA 硬件上有原生支持(稀疏张量核心),且精度损失可通过’稀疏感知训练’控制;故它是当前’稀疏加速’最实用的方案。③ 剪枝与量化的正交性——剪枝减少’参数个数’、量化减少’每参数字节’;两者可叠加(稀疏 + 低比特)。④ Lottery Ticket 假说——Frankle & Carbin 发现’随机初始化的稠密网络中存在一个稀疏子网络(中奖彩票),单独训练它能达到原网络精度’;这暗示’稠密网络大部分参数是冗余的’,但也引发了’稀疏训练是否真能省算力’的争论(因为找到子网络的成本高)。⑤ LLM 剪枝的特殊性——LLM 对剪枝敏感(尤其结构化剪枝会显著掉点),因为其知识分布式存储;故 LLM 压缩更依赖量化与蒸馏(而非剪枝);近年有 SparseGPT/Wanda 等’一次性剪枝’方法(用校准数据估计重要性,可剪 50% 而损失较小)。⑥ 面试要点——被问’剪枝’,应给出’非结构化(高压缩、难加速)/ 结构化(易加速、损精度)/ 半结构化 2:4(硬件支持,主流)‘三类与取舍,并强调’压缩率 ≠ 加速比‘;能提到’SparseGPT/Wanda 的一次性剪枝’与’Lottery Ticket’是深度理解的标志。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
Deep Dive & Engineering Trade-offs: ① Accuracy vs Hardware Accelerability Trade-off: – Unstructured: Sparsity $80text{–}90%$ with minimal loss, but effective speedup $approx 1.0times$. – Semi-Structured 2:4: Fixed $50%$ sparsity, guaranteed $1.5text{–}1.8times$ real-world speedup on A100/H100, requires fine-tuning to recover small accuracy drops. – Structured: Sparsity $20text{–}40%$, guaranteed dense speedup across any CPU/GPU/mobile device, but drops accuracy rapidly beyond $30%$ pruning. ② Pruning Criteria: Magnitude-based pruning ($|W|$) is simple; second-order Hessian-based pruning (Optimal Brain Damage / SparseGPT) measures loss curvature $Delta w = -frac{H^{-1} w}{[H^{-1}]_{ii}}$, preserving accuracy at much higher sparsity levels. ③ Lottery Ticket Hypothesis: Dense randomly-initialized networks contain sparse sub-networks (‘winning tickets’) that, when trained in isolation from scratch with original initializations, match the accuracy of the full dense model. ④ Pruning LLMs: In 70B+ models, structured pruning of layers or intermediate dimensions often causes severe disruption to emergent capabilities; 2:4 semi-structured pruning with SparseGPT has become the primary production route. ⑤ Interview Strategy: Draw the 2:4 pattern diagram, contrast theoretical sparsity vs physical wall-clock speedup on modern GPUs, and explain how Sparse Tensor Cores achieve hardware acceleration.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 只看稀疏率不看实际加速比
- ⚠️ 以为非结构化剪枝能在通用 GPU 上线性加速
English Pitfalls:
– Assuming 70% unstructured pruning automatically provides a 70% speedup on GPUs (it usually runs slower due to irregular memory indexing)
– Applying structured pruning aggressively to LLMs without extensive post-training fine-tuning
– Failing to recognize that 2:4 sparsity requires NVIDIA Ampere (A100) or newer GPUs for hardware acceleration
六、高频深度面试追问与预测 (Follow-Up Questions)
- 为什么非结构化剪枝难以加速?
- How does SparseGPT apply second-order Hessian compensation to prune 100B+ models in hours?
- 结构化剪枝如何决定剪哪些通道?
- How do NVIDIA Sparse Tensor Cores physically compress 2:4 sparse weights in memory?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
知识蒸馏 (Knowledge Distillation):温度超参、软标签损失与学生网络(Knowledge Distillation: Temperature Scaling & Soft Targets) - 🗺️ 知识图谱模块:
大语言模型全景图谱
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。