【AI 核心深度 M5-019】解释后训练阶段的 scaling(RL/推理算力)与预训练 scaling 的差异。(Post-Training Scaling (RL and Test-Time Compute) vs. Pre-Training Scaling)深度数理推导与工程落地解析

所属模块:M5 · NLP 与大语言模型 (NLP & Large Language Models) | 专题分类:Scaling Laws (Scaling Laws & Compute Allocation) | 难度等级:Hard

一、核心一句话结论 (One-Sentence Summary)

预训练 scaling 平滑可预测;RL/后训练的 scaling 更不稳定、更依赖任务与数据,且常呈’先快后慢’。

ADVERTISEMENT · 赞助推荐

Pre-training scaling expands general world knowledge via next-token prediction over static text, whereas post-training scaling expands complex reasoning through reinforcement learning exploration and test-time search computation.

二、核心考点要义 (Key Insights)

  • 📌 预训练:平滑幂律、可跨规模外推
  • 📌 RL/后训练:曲线不规则、依赖任务与奖励设计
  • 📌 推理时算力(test-time)是第三条 scaling 轴

English Insights:
– Pre-training scaling (System 1): predictable power laws over compute, parameters, and tokens; builds broad world knowledge, lexical fluency, and pattern recognition
– Post-training RL scaling (System 2): reinforcement learning with verifiable rewards (math, code) enables self-discovery of extended Chain-of-Thought, error correction, and backtracking (OpenAI o1, DeepSeek-R1)
– Test-Time compute scaling: spending additional inference compute per query (rejection sampling, best-of-$N$, Monte Carlo tree search) trades inference latency for mathematical accuracy

三、核心数学原理与机理推导 (Mathematical Principles & Derivation)

$$mathcal{L}_{text{pretrain}} text{smooth power law};qquad text{RL}: text{task-dependent}, text{less predictable}$$

数学机理:三条 scaling 轴。(1) 预训练 scaling(参数 × 数据 × 训练算力)——平滑幂律、跨规模可外推(见前述)。(2) 后训练/RL scaling(RL 算力)——用强化学习(RLHF/RLVR)提升能力时,性能随 RL 算力提升,但曲线更不规则:(a) 依赖任务与奖励设计(不同任务的曲线差异大);(b) 常呈’先快后慢’(早期快速提升、后期平台或退化);(c) 不稳定(奖励黑客、KL 漂移、崩溃);(d) 难以跨规模外推(RL 的超参与数据依赖强)。原因:RL 优化的是非平稳目标(策略变化 → 数据分布变化 → 奖励估计变化),且奖励是代理(proxy)而非真实目标。(3) 推理时算力 scaling(test-time compute)——在推理时投入更多计算(长 CoT、多次采样、搜索、验证器)可提升表现;研究表明其效果可与模型规模互补(某些任务上’小模型 + 更多推理算力’可匹配’大模型’)。三条轴的差异与配合——(a) 可预测性:预训练 >> 推理时 >> RL;(b) 成本结构:预训练是一次性、推理时是持续的、RL 介于两者;(c) 能力类型:预训练提供’知识与基础能力’、RL 提供’对齐与推理习惯’、推理时提供’按需的深度思考’。实践——(i) 先用预训练 scaling 达到基础能力;(ii) 用 SFT + RLHF/RLVR 对齐与提升推理;(iii) 用推理时算力按需增强(如难题用长 CoT)。三者互补而非替代。

📖 查看英文严格数学推导 (English Mathematical Derivation)

Mathematical Mechanism: 1. Pre-training Power Law: Loss follows passive unsupervised next-token prediction: $$mathcal{L}(C_{text{pretrain}}) = E + left(frac{C_0}{C_{text{pretrain}}}right)^alpha$$ Diminishing returns require $10times$ more compute for incremental loss reductions. 2. Test-Time Compute Scaling Formulations: For complex problem $x$, spending additional inference compute $C_{text{test}}$ scales accuracy along two complementary axes: – Parallel Sampling (Best-of-$N$ with Verifier): Sample $N$ independent solutions ${y_1, dots, y_N} sim P_theta(cdot mid x)$ and select the highest-scoring candidate via verifier $V(x, y)$: $$text{Pass@}N = 1 – prod_{i=1}^N (1 – P(y_i text{ is correct}))$$ As $N$ expands, accuracy increases logarithmically with compute $C_{text{test}} = N times C_{text{single}}$. – Sequential Thinking (Extended CoT / DeepSeek-R1): The model generates an internal monologue of reasoning tokens $T_{text{reasoning}}$. Compute scales as $C_{text{test}} = (L_{text{prompt}} + T_{text{reasoning}}) times text{FLOPs}$. RL rewards longer reasoning trajectories when they successfully resolve edge cases and backtrack from errors.

四、工业级落地权衡与工程考量 (Industrial Trade-offs)

深度剖析与工程权衡:① ‘RL scaling 更难’的根源是’非平稳性’——RL 的优化目标随策略变化(数据分布由当前策略产生),故不像监督学习那样有稳定的损失曲线;这使’预测与调参’更难。② 奖励设计的杠杆作用——RL 的收益高度依赖奖励是否’与真实目标对齐’;若奖励可被钻空子(reward hacking),算力越多反而越差。故 RL scaling 的前提是可靠的奖励(可验证任务天然满足)。③ 推理时算力的经济性——推理时算力是持续成本(每次请求都付费),而预训练是一次性成本;故需按’推理量’权衡(类似 over-training 的经济学)。对高频服务,’小模型 + 少量推理算力’可能比’大模型’更经济。④ ‘推理算力与模型规模的等价性’——有研究(如’Scaling LLM Test-Time Compute’)表明:在某些任务上,用更多推理算力的小模型可匹配更大模型;但不是所有任务都如此(依赖任务是否可分解为可验证的步骤)。⑤ 与’能力上限’的关系——推理时算力无法突破模型的知识上限(不知道的事实无法通过多想获得);故它主要提升’推理/组合’能力,而非’知识’。⑥ 面试要点——被问’后训练 scaling 与预训练 scaling 有何不同’,应给出’可预测性(预训练平滑、RL 不规则)+ 非平稳性(RL 目标随策略变化)+ 三条轴的互补配合‘;能指出’RL scaling 的前提是可靠奖励’与’推理算力无法突破知识上限’是深度理解的标志。

⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)

Deep Dive & Engineering Trade-offs: ① The Shift in Compute Economics: Instead of spending $$100text{M}$ to pre-train a 500B parameter model, labs can pre-train an efficient 70B model for $$10text{M}$ and invest remaining compute into extensive post-training RL and dynamic test-time search, achieving superior reasoning at lower overall cost. ② Verifiable vs Subjective Domains: Post-training RL scaling works exceptionally well in domains with verifiable ground truth (coding tests, Olympiad mathematics, formal logic). In open-ended creative writing, lack of precise verifiers leads to reward hacking and sycophancy. ③ Inference Latency SLA: Test-time compute scaling directly inflates user latency (generating 5,000 reasoning tokens takes 20-40 seconds). Systems must route queries dynamically: use fast direct generation for simple questions and engage test-time reasoning only for complex multi-step problems. ④ Self-Correction Emergence: DeepSeek-R1 showed that purely through RL reward on answer correctness, models organically discover cognitive behaviors like backtracking (‘Wait, let me double check this equation…’), self-questioning, and alternative exploration. ⑤ Interview Strategy: Contrast System 1 (intuitive pre-training next-token prediction) vs System 2 (deliberate test-time search and reasoning), explain the Best-of-$N$ and extended CoT scaling mechanics, and cite OpenAI o1 / DeepSeek-R1.

五、常见面试避坑陷阱 (Common Pitfalls & Traps)

  • ⚠️ 假设 RL 性能随算力平滑可预测地提升
  • ⚠️ 认为推理时算力可替代知识学习

English Pitfalls:
– Assuming post-training RL can replace pre-training (RL refines reasoning and search, but cannot inject factual world knowledge missing from pre-training)
– Applying test-time search to open-ended creative tasks without objective outcome verifiers
– Ignoring the user-facing latency and financial cost of generating thousands of internal thinking tokens per query

六、高频深度面试追问与预测 (Follow-Up Questions)

  1. 为什么 RL 的 scaling 更难预测?
  2. How does DeepSeek-R1 incentivize the spontaneous emergence of self-correction behaviors without supervised human demonstrations?
  3. 三条 scaling 轴如何配合?
  4. What is the mathematical trade-off between parallel Best-of-$N$ sampling and sequential extended reasoning tokens?

七、知识图谱对齐 (Knowledge Graph Anchor)

  • 🔗 关联底层卡片:缩放法则 (Scaling Laws):Chinchilla 计算最优配比与涌现能力 (Scaling Laws: Kaplan, Chinchilla Optimal Compute & Emergence)
  • 🗺️ 知识图谱模块:大语言模型全景图谱

🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)

本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。

👉 前往 TalentMe 交互式研读本题 (M5-019) →


Discover more from AirSOTA – Air School Of Thoughts AtoZ

Subscribe to get the latest posts sent to your email.