【AI 核心深度 M5-061】解释零 RL / RLVR 的 cold start 问题。(Cold Start Problem in Zero-RL / RLVR and Mitigation Strategies)深度数理推导与工程落地解析

所属模块:M5 · NLP 与大语言模型 (NLP & Large Language Models) | 专题分类:GRPO 与推理模型 (GRPO & Reasoning Models) | 难度等级:Hard

一、核心一句话结论 (One-Sentence Summary)

从纯基座做 RL 可能输出格式混乱、语言混杂、可读性差;冷启动 SFT 或格式奖励可稳定起始阶段。

ADVERTISEMENT · 赞助推荐

The cold start problem occurs when a base model has a near-zero success rate on hard reasoning tasks, producing zero positive rewards and stalling RLVR exploration, solved through curated CoT seed demonstrations, prompt scaffolding, and difficulty curricula.

二、核心考点要义 (Key Insights)

  • 📌 纯 RL 从基座开始:输出格式混乱、语言混杂、可读性差
  • 📌 冷启动 SFT(少量长 CoT 数据)可稳定起始阶段
  • 📌 或加’格式奖励’(要求特定输出格式)

English Insights:
– Root mechanism: in binary verification ($R in {0, 1}$), if policy $pi_theta$ has pass rate $p=0$ on problem $x$, all $G$ rollouts fail; advantages evaluate to zero, and policy gradients vanish
– DeepSeek-R1 solution: cold-start SFT stage fine-tuning the base model on a few thousand high-quality long-CoT examples before initiating GRPO
– Alternative mitigations: difficulty curriculum learning (starting with grade-school math), synthetic hints and sub-goal scaffolding, and rejection sampling initialization

三、核心数学原理与机理推导 (Mathematical Principles & Derivation)

$$text{cold start}: text{base model}totext{RL} Rightarrow text{format collapse}, text{language mixing};qquad text{fix}: text{SFT warmup}$$

数学机理:cold start 问题——若直接从纯基座模型(未经 SFT)开始 RLVR,会观察到若干退化现象:(a) 格式混乱——输出没有清晰结构(不分步骤、不区分思考与答案);(b) 语言混杂——中英文夹杂(因为基座模型的多语言能力未被约束);(c) 可读性差——推理过程难以理解(跳跃、省略);(d) 奖励提取困难——因为答案没有明确标记,验证器难以提取答案(导致奖励信号错误)。成因——基座模型只学过’续写’,没有’按指令组织回答’的能力;RL 虽有正确性奖励,但格式与可读性没有奖励信号(除非显式设计),故模型会’用最省力的方式’追求正确性(可能输出混乱但答案对)。解决方案:(1) 冷启动 SFT(cold-start SFT)——用少量(数百到数千条)高质量的长 CoT 数据做 SFT,教模型’如何组织推理’(格式、步骤、语言);这为 RL 提供了’格式正确的起点’(DeepSeek-R1 用了约数千条冷启动数据)。(2) 格式奖励——在奖励中加入’是否遵循格式’(如’把答案放在 oxed{} 中’、’用与问题相同的语言’);这使 RL 能自己学会格式。(3) 混合奖励——正确性奖励 + 格式奖励 + 语言一致性奖励。实证——DeepSeek-R1 报告:纯 RL(无冷启动) 也能提升推理能力,但输出可读性差、语言混杂;加冷启动 SFT 后(R1 的正式版本)输出清晰可读且性能更好。故’冷启动 + RL’是推荐配方。与’涌现’的关系——cold start 的存在不否定’长 CoT 可涌现’(纯 RL 也能学到推理),而是说明’可读性与格式需要额外引导‘。冷启动数据量——通常很少(数千条);关键是格式示范而非知识注入。

📖 查看英文严格数学推导 (English Mathematical Derivation)

Mathematical Mechanism: 1. The Zero-Gradient Exploration Trap: In GRPO, the advantage for sample $i$ in group $G$ is: $$hat{A}_i = frac{R_i – bar{R}}{sigma_R + epsilon}$$ If problem $x$ is too difficult for current policy $pi_theta$, all sampled candidates fail: $R_1 = R_2 = dots = R_G = 0$. Consequently: $$bar{R} = 0, quad sigma_R = 0 implies hat{A}_i = frac{0 – 0}{0 + epsilon} = 0 quad forall i in {1, dots, G}$$ The policy gradient evaluates to identically zero: $$nabla_theta mathcal{L}_{text{GRPO}} = mathbf{0}$$ The model receives zero directional feedback. Without an initial spark of success, RL exploration cannot bootstrap itself. 2. Cold-Start Initialization Bridge: Pre-train policy $theta$ on a curated seed demonstration set $mathcal{D}_{text{cold_start}}$: $$theta_{text{init}} = argmax_theta sum_{(x, R, y) in mathcal{D}_{text{cold_start}}} log P_theta(R, y mid x)$$ Raising initial pass rate $p(x)$ from $0.00$ to $ge 0.05$ ensures that across $G=16$ samples, the probability of sampling at least one correct path jumps from $0%$ to $1 – (0.95)^{16} approx 56%$, kickstarting the RL engine.

四、工业级落地权衡与工程考量 (Industrial Trade-offs)

深度剖析与工程权衡:① ‘冷启动只需少量数据’是关键——它的作用是’示范格式’而非’教推理’(推理由 RL 学);故数千条足够。这与’SFT 数据要少而精’的原则一致。② ‘格式奖励’的工程价值——它是’让 RL 自己学会格式’的替代方案(无需 SFT);但对复杂格式(多步骤、结构化输出)可能不如 SFT 示范有效。③ ‘语言一致性’的实际问题——多语言模型在 RL 后可能出现’语言混杂’(中英夹杂),影响用户体验;故常加’语言一致性奖励’(要求用 prompt 的语言回答)。④ ‘奖励提取’的工程细节——若输出格式不固定,验证器难以可靠提取答案(可能把’思考中的中间结果’当成答案);故格式约束(如 oxed{})是 RLVR 的工程前提(与约束解码相关)。⑤ 与’蒸馏’的关系——冷启动数据常来自’更强的模型’(如用 R1 的输出去教小模型);故冷启动与蒸馏在数据层面相通。⑥ 面试要点——被问’纯 RL 从基座开始行不行’,应给出’能提升推理但输出可读性差、语言混杂 → 需冷启动 SFT 或格式奖励‘,并说明’冷启动只需少量数据(教格式而非教推理)‘与’格式约束是奖励提取的前提‘;能指出’纯 RL 不否定涌现、只说明格式需引导’是深度理解的标志。

⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)

Deep Dive & Engineering Trade-offs: ① DeepSeek-R1 Cold-Start Design: DeepSeek curated a few thousand high-quality reasoning examples with explicit readability constraints: (a) reasoning steps formatted inside ` … `; (b) markdown headers; and (c) reflection checks. This eliminated the linguistic chaos and multi-language mixing observed in R1-Zero while providing the initial non-zero success rate needed for RLVR. ② Curriculum Learning Schedule: Start RLVR on high-success problems (GSM8K, simple LeetCode with $p > 0.3$). As the model masters intermediate problem-solving patterns, gradually inject harder problems (MATH, Codeforces, AIME). The model transfers learned decomposition skills to maintain $p > 0$ on difficult tiers. ③ Scaffolding with Partial Prompts: For intractable problems, provide the first 2-3 reasoning steps as prompt context; the model learns to complete the remaining steps, receiving reward. Gradually truncate the scaffolding over successive epochs until the model solves the problem from scratch. ④ Risk of Oversized Cold-Start Data: Training on too much SFT demonstration data ($>50text{k}$ examples) biases the model toward rigid memorization of human reasoning styles, restricting its ability to discover novel, super-human search paths during RL. Keep cold-start data small ($sim 1text{k}text{–}5text{k}$ samples). ⑤ Interview Strategy: Formulate the zero-gradient mathematical trap when $R_i = 0$ for all $i$, explain how cold-start SFT raises pass rate past the $1 – (1-p)^G$ threshold, and describe curriculum scheduling.

五、常见面试避坑陷阱 (Common Pitfalls & Traps)

  • ⚠️ 以为纯 RL 从基座开始就能得到可读的输出
  • ⚠️ 冷启动用大量数据(应少量精标)

English Pitfalls:
– Attempting to launch RLVR directly on Olympiad math without cold-start seeding or curriculum (training stalls at step 0)
– Using massive SFT datasets for cold start, which suppresses the model’s ability to discover creative self-correction strategies
– Failing to filter cold-start demonstrations for strict single-language and markdown formatting standards

六、高频深度面试追问与预测 (Follow-Up Questions)

  1. 为什么纯 RL 会出现’语言混杂’?
  2. Why did DeepSeek-R1 use only a few thousand cold-start demonstrations rather than a traditional massive 500k-sample SFT corpus?
  3. 冷启动 SFT 数据需要多少?
  4. How does prompt scaffolding mathematically bootstrap policy exploration on zero-pass-rate problems?

七、知识图谱对齐 (Knowledge Graph Anchor)

  • 🔗 关联底层卡片:GRPO 组相对策略优化:DeepSeek-R1 纯强化学习推理与长思考链涌现 (GRPO: Group Relative Policy Optimization & DeepSeek-R1)
  • 🗺️ 知识图谱模块:大语言模型全景图谱

🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)

本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。

👉 前往 TalentMe 交互式研读本题 (M5-061) →


Discover more from AirSOTA – Air School Of Thoughts AtoZ

Subscribe to get the latest posts sent to your email.