所属模块:
M5 · NLP 与大语言模型 (NLP & Large Language Models)| 专题分类:Scaling Laws (Scaling Laws & Compute Allocation)| 难度等级:Easy
一、核心一句话结论 (One-Sentence Summary)
Kaplan:损失随参数量/数据/算力幂律下降,参数优先;Chinchilla:参数量与数据量应同比例增长(约 20 token/参数)。
Kaplan et al. concluded that model parameters should scale faster than training tokens ($N propto C^{0.73}, D propto C^{0.27}$), while Chinchilla proved that parameters and tokens should scale in equal proportion ($N propto C^{0.5}, D propto C^{0.5}$, $approx 20$ tokens per parameter).
二、核心考点要义 (Key Insights)
- 📌 幂律:损失随 N、D 幂律下降(对数坐标下线性)
- 📌 Kaplan:早期结论偏向’增大参数’
- 📌 Chinchilla:N 与 D 应同比例,约 20:1(修正了 Kaplan)
English Insights:
– Kaplan Scaling Law (OpenAI 2020): concluded that model size matters far more than data; recommended scaling parameters aggressively while scaling data slowly, leading to severely under-trained models like GPT-3 175B (300B tokens)
– Chinchilla Correction (DeepMind Hoffmann et al. 2022): fixed sub-optimal learning rate schedules in Kaplan’s setup; proved parameters $N$ and tokens $D$ should scale symmetrically ($1:1$ ratio in compute exponents)
– Compute-optimal rule: for compute-optimal training, a model should be trained on approximately 20 tokens per parameter (e.g., a 70B model requires $approx 1.4text{T}$ tokens for compute optimality)
三、核心数学原理与机理推导 (Mathematical Principles & Derivation)
$$mathcal{L}(N,D)=mathcal{L}_infty+AN^{-alpha}+BD^{-beta};qquad text{Chinchilla}: D^{star}approx20N^{star}$$
数学机理:Kaplan 等(2020) 发现语言模型损失随参数量 N、数据量 D、训练算力 C 呈幂律下降:L(N,D)=L_∞+A·N^{−α}+B·D^{−β}(在对数坐标下近似线性)。基于此,他们给出’给定算力下的最优配置’,结论是优先增大参数量(参数比数据更’划算’),并给出了’参数量与数据量不必同比例增长’的建议。Chinchilla(Hoffmann 等 2022) 用更多训练运行修正了结论:发现 Kaplan 的实验设置(学习率调度、参数量与数据的耦合)导致低估了数据的作用;修正后的结论是——参数量 N 与训练 token 数 D 应大致同比例增长,最优比约 D≈20N(即每个参数配 20 个 token)。核心差异与影响——(a) 按 Chinchilla,当时的大模型(如 GPT-3 175B 用 300B token,比约 1.7:1)训练不足;(b) 遵循 Chinchilla 的模型(如 Chinchilla 70B 用 1.4T token)在同等算力下优于更大的模型;(c) 这直接催生了’用更多数据训练更小模型’的实践(LLaMA 系列即此路线的代表)。公式的其他形式——算力与损失:C≈6ND(前向 2 + 反向 4 FLOPs/参数/token),故’给定算力 C 下最优化 N 与 D’即最小化 L(N, C/(6N)),可得 N∝C^{…}、D∝C^{…}(两者同阶)。注意——Chinchilla 是’计算最优(compute-optimal)‘结论(在给定训练算力下最小化损失);而推理成本会改变这个结论(见下一题)。
📖 查看英文严格数学推导 (English Mathematical Derivation)
Mathematical Mechanism: 1. The Compute Budget Constraint: Pre-training compute FLOPs are approximately: $$C approx 6 N D$$ 2. Power-Law Loss Formulation: Loss is modeled as a function of parameters $N$ and dataset size $D$: $$L(N, D) = E + frac{A}{N^alpha} + frac{B}{D^beta}$$ Where $E$ is irreducible loss. To find the compute-optimal allocation of $N(C)$ and $D(C)$ that minimizes $L(N, D)$ subject to $6 N D = C$: Set up the Lagrangian: $mathcal{L}(N, D, lambda) = E + A N^{-alpha} + B D^{-beta} + lambda (6 N D – C)$. Solving $frac{partial mathcal{L}}{partial N} = 0$ and $frac{partial mathcal{L}}{partial D} = 0$ yields: $$N propto C^{frac{beta}{alpha + beta}}, quad D propto C^{frac{alpha}{alpha + beta}}$$ 3. Kaplan vs Chinchilla Parameter Values: – Kaplan (2020): $alpha = 0.076, beta = 0.057 implies N propto C^{0.73}, D propto C^{0.27}$. Recommended $N$ scale $3times$ faster than $D$. – Chinchilla (2022): Evaluated $>400$ models where LR schedule cosine duration matched actual token count. Fitted: $alpha approx 0.34, beta approx 0.28 implies N propto C^{0.50}, D propto C^{0.50}$. Thus, both parameters and tokens should scale in equal proportion: $N_{text{opt}} propto sqrt{C}, D_{text{opt}} propto sqrt{C}$, with ratio $D / N approx 20$.
四、工业级落地权衡与工程考量 (Industrial Trade-offs)
深度剖析与工程权衡:① ‘计算最优 ≠ 部署最优’——Chinchilla 优化的是’训练算力’;但若考虑推理成本(服务大量用户),则’用更多数据训练更小的模型’(over-training)更划算——因为小模型的推理成本低。这就是 LLaMA 系列’用远超 20:1 的比例训练小模型’的原因(如 LLaMA-2 7B 用 2T token,比约 285:1)。② 幂律的实用性——(a) 预测:可用小规模实验拟合幂律、外推大模型的损失(指导是否值得扩大);(b) 资源规划:给定预算算最优 N/D;(c) 限制认知:幂律无法无限延续(数据有限、损失有下界 L_∞)。③ Kaplan 为何出错——主要是实验设计问题(学习率调度未随规模调优、以及’固定数据重复’的影响);这提醒我们’scaling law 的结论依赖实验设置‘,需谨慎解读。④ 数据受限时的 scaling——当高质量数据耗尽时,幂律失效(数据项饱和);此时需 (a) 数据高效方法(更好的架构/优化)、(b) 合成数据、(c) 多轮 epoch(收益递减)。⑤ 与架构的关系——scaling law 的常数(A、B、α、β)依赖架构与优化;故’某架构的 scaling law’不能直接用于另一架构(需重新拟合)。⑥ 面试要点——被问’scaling law’,应给出’L=L_∞+AN^{−α}+BD^{−β} 幂律‘与’Chinchilla 修正为 D≈20N 同比例增长‘,并说明’计算最优 ≠ 部署最优(推理成本改变结论)‘;能指出’Kaplan 的偏差源于实验设置’是深度理解的标志。
⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)
Deep Dive & Engineering Trade-offs: ① Kaplan’s Flaw: Kaplan kept the cosine learning rate schedule fixed to a large token count while evaluating models on early checkpoints. Because models at early steps of a long cosine schedule have higher learning rates than models trained with a cosine schedule tuned specifically for that step count, Kaplan systematically underestimated the power of training on more data. ② Why Chinchilla 70B Beat Gopher 280B: DeepMind’s Chinchilla (70B parameters trained on 1.4T tokens) used the exact same compute budget as Gopher (280B trained on 300B tokens), but decisively outperformed Gopher across all tasks while being $4times$ smaller and faster at inference. ③ Inference Over-Training Trend: Chinchilla defines training compute optimality. In production, models are used for billions of inference requests; over-training smaller models (e.g., LLaMA-3 8B trained on 15T tokens $approx 1875$ tokens/param) vastly outperforms Chinchilla optimality in total life-cycle cost. ④ Downstream Benchmark Correlation: Pre-training cross-entropy loss follows strict power-law scaling across 6 orders of magnitude; downstream task performance generally scales monotonically with cross-entropy loss. ⑤ Interview Strategy: Write down $C approx 6ND$, set up the Lagrangian optimization, contrast the Kaplan ($C^{0.73}$ vs $C^{0.27}$) and Chinchilla ($C^{0.5}$ vs $C^{0.5}$) exponents, and explain Kaplan’s learning rate schedule flaw.
五、常见面试避坑陷阱 (Common Pitfalls & Traps)
- ⚠️ 以为 Chinchilla 的 20:1 就是’部署最优’
- ⚠️ 把某架构的 scaling law 直接套用到另一架构
English Pitfalls:
– Assuming Chinchilla’s 20 tokens/param is a hard upper bound (it is the training compute-optimal point, not the inference cost-optimal point)
– Believing Kaplan’s scaling law is still valid (it was superseded by Chinchilla due to experimental methodology flaws)
– Forgetting the factor of 6 in $C = 6ND$ (2 FLOPs per parameter forward + 4 FLOPs backward)
六、高频深度面试追问与预测 (Follow-Up Questions)
- 为什么 Kaplan 与 Chinchilla 结论不同?
- Why did Kaplan’s fixed cosine learning rate schedule systematically underestimate the benefit of training tokens?
- 如何用 Chinchilla 指导预算分配?
- How does factoring in inference query volume alter the optimal token-to-parameter ratio beyond Chinchilla’s 20:1?
七、知识图谱对齐 (Knowledge Graph Anchor)
- 🔗 关联底层卡片:
缩放法则 (Scaling Laws):Chinchilla 计算最优配比与涌现能力(Scaling Laws: Kaplan, Chinchilla Optimal Compute & Emergence) - 🗺️ 知识图谱模块:
大语言模型全景图谱
🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)
本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。