【AI 核心深度 M5-028】解释拒绝采样微调(RFT / STaR)的机制与作用。(Rejection Sampling Fine-Tuning (RFT / STaR) Mechanisms)深度数理推导与工程落地解析

所属模块:M5 · NLP 与大语言模型 (NLP & Large Language Models) | 专题分类:指令微调与 SFT (Instruction Tuning & Supervised Fine-Tuning) | 难度等级:Hard

一、核心一句话结论 (One-Sentence Summary)

用模型自己生成多个答案,按正确性/奖励筛选后做 SFT;迭代进行可自我提升,是可验证任务的标准做法。

ADVERTISEMENT · 赞助推荐

Rejection Sampling Fine-Tuning (RFT / STaR) samples multiple candidate reasoning paths from a model, filters them using automated verifiers, and fine-tunes the model on its own successful generations to bootstrap reasoning capabilities.

二、核心考点要义 (Key Insights)

  • 📌 采样多个答案 → 筛选正确/高奖励 → SFT → 迭代
  • 📌 对可验证任务(数学/代码)质量可控
  • 📌 与 CoT 蒸馏、RLVR 同属’自提升’范式

English Insights:
– Core mechanism: for each problem $x$, sample $N$ candidate solutions from policy $pi_theta$; evaluate candidates with deterministic verifiers (unit tests, math ground truth); retain only correct paths
– Self-Taught Reasoner (STaR): pairs rejection sampling with rationalization (asking the model to generate a chain of thought when an answer is initially incorrect), creating a self-improving loop
– Bridge between SFT and RL: RFT acts as an offline, supervised surrogate for reinforcement learning, avoiding the training instability of PPO while harvesting diverse correct reasoning traces

三、核心数学原理与机理推导 (Mathematical Principles & Derivation)

$$mathcal{D}{text{new}}={ysimpitheta: text{verify}(y)=1};qquad thetaleftarrowtext{SFT}(theta,mathcal{D}_{text{new}})$$

数学机理:RFT(Rejection sampling Fine-Tuning) 又名 STaR(Self-Taught Reasoner)、RAFT 等,流程为:(1) 采样——用当前模型对每个问题生成多个回答(如 16~64 个,含 CoT 步骤);(2) 筛选(rejection)——按验证器(对可验证任务:答案是否与标准答案一致;或单元测试是否通过)或奖励模型打分,只保留正确/高奖励的样本;(3) SFT——用筛选后的样本微调模型;(4) 迭代——用新模型重新采样,重复 (1)~(3)。为什么有效:(a) 扩大有效训练集——把’模型能解出但采样才得到’的问题变成训练数据(把’偶尔成功’转化为’稳定成功’);(b) 过滤错误——只学正确轨迹,避免继承错误(这是相对 CoT 蒸馏的优势);(c) 自我提升——模型在自己能解的问题上变得更强,从而能解更多问题(迭代扩展能力边界);(d) 无需外部强教师(自蒸馏)。与 RLVR 的关系——RFT 用二值筛选(正确/错误)做监督学习(模仿正确轨迹);RLVR(GRPO 等)用连续奖励做策略梯度(奖励越高概率越大)。RFT 更简单稳定(像监督学习),RLVR 能利用’部分正确’的信号且可优化未通过验证的样本。实践中两者常组合——RFT 做冷启动(拿到第一批正确轨迹),RLVR 做后续优化。局限——(a) 只学’已能解出’的(无法超越当前能力边界,除非有更强的验证器或教师);(b) 多样性下降(反复筛选正确轨迹会使输出趋同);(c) 依赖可靠的验证器(对开放任务不可用)。

📖 查看英文严格数学推导 (English Mathematical Derivation)

Mathematical Mechanism: 1. Rejection Sampling Generation Phase: For input problem $x_i in mathcal{D}$ and current policy $pi_theta$, sample $K$ candidate solutions with high temperature ($T in [0.7, 1.0]$): $$y_{i, 1}, y_{i, 2}, dots, y_{i, K} sim pi_theta(cdot mid x_i)$$ 2. Automated Verification Filtering: Evaluate candidate correctness using an oracle verifier $V(x, y) in {0, 1}$ (e.g., Python execution or symbolic math check). Construct the filtered dataset $mathcal{D}^*$: $$mathcal{D}^* = bigcup_{i} left{ (x_i, y_{i, k}) mid V(x_i, y_{i, k}) = 1 right}$$ 3. Fine-Tuning Update Phase: Fine-tune policy $theta$ on $mathcal{D}^*$ via standard supervised cross-entropy: $$theta^* = argmax_theta sum_{(x, y) in mathcal{D}^*} sum_{t=1}^{|y|} log pi_theta(y_t mid x, y_{<t})$$ Iterating this process ($t=1, 2, dots, M$) progressively bootstraps the model's reasoning accuracy.

四、工业级落地权衡与工程考量 (Industrial Trade-offs)

深度剖析与工程权衡:① ‘把偶尔成功变成稳定成功’是关键机制——采样 64 次可能得到几个正确答案,这些样本正是’模型能力边界附近’的宝贵监督信号(pass@64 高但 pass@1 低的问题);RFT 把它们转化为训练数据,直接提升 pass@1。② RFT vs RLVR 的选择——RFT 更简单(监督学习、稳定、无需策略梯度);RLVR 更强(能优化未通过验证的样本、利用部分奖励)但更复杂(需处理 KL、方差、不稳定)。实践中常’RFT 冷启动 + RLVR 精调’。③ ‘验证器’是前提——RFT/RLVR 都依赖可靠验证;对数学(对答案)、代码(跑测试)可行;对写作、开放推理不可行(需奖励模型,可靠性下降)。④ 迭代次数的收益递减——每次迭代都提升,但提升幅度递减(因为容易的问题已被’消化’);且过度迭代会导致多样性崩塌(输出趋同)。⑤ 与’蒸馏’的关系——RFT 是’自蒸馏’(用自己筛选的样本);若有强教师,则用教师的正确轨迹(CoT 蒸馏)——两者可结合(先用教师数据冷启动,再用 RFT 迭代)。⑥ 面试要点——被问’如何让模型自己变强’,应给出’RFT/STaR:采样 → 筛选 → SFT → 迭代‘与’把偶尔成功转化为稳定成功‘这一机制,并区分’RFT(监督、二值筛选)vs RLVR(策略梯度、连续奖励)’;能指出’RFT 无法超越验证器/教师的上限’是深度理解的标志。

⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)

Deep Dive & Engineering Trade-offs: ① Advantages over PPO / Online RL: RFT uses standard supervised learning pipelines (SFT), which are completely stable, highly parallelizable, and free from PPO hyperparameter volatility (value network training, clipping epsilon, advantage estimation). ② Solution Diversity vs Mode Collapse: High sampling temperature ($T=0.8$) encourages the model to explore disparate reasoning strategies. Fine-tuning on diverse correct paths teaches the model multiple valid problem-solving routes. ③ The Exploration Frontier Limit: RFT cannot improve on problems where the model’s pass@$K$ probability is exactly zero ($P(text{success}) = 0$). To expand the frontier, problems must be paired with hints or seeded with teacher demonstrations. ④ Iterative Refinement (STaR / V-STaR): In each iteration, the updated model $pi_{theta_{t+1}}$ can solve harder problems that $pi_{theta_t}$ failed, expanding the training corpus $mathcal{D}^*$ over successive rounds. ⑤ Interview Strategy: Detail the 3-step cycle (Sample $to$ Verify $to$ Fine-tune), contrast RFT with PPO, explain why deterministic verifiers are essential, and address the pass@$K=0$ exploration bottleneck.

五、常见面试避坑陷阱 (Common Pitfalls & Traps)

  • ⚠️ 以为 RFT 能突破当前能力上限(需更强验证器/教师)
  • ⚠️ 无限迭代导致输出多样性崩塌

English Pitfalls:
– Using RFT in subjective domains without reliable ground-truth verifiers (leads to reinforcement of plausible hallucinations)
– Sampling with temperature $T=0$ (yields $K$ identical outputs and zero exploration diversity)
– Ignoring the exploration limit when a model fails to generate even a single correct trace across all $K$ samples

六、高频深度面试追问与预测 (Follow-Up Questions)

  1. RFT 与 RLVR 的区别?
  2. How does STaR handle problems where the model fails to generate a correct solution during initial sampling?
  3. 为什么 RFT 需要迭代?
  4. Why is Rejection Sampling Fine-Tuning significantly more computationally stable than PPO in mathematical domains?

七、知识图谱对齐 (Knowledge Graph Anchor)

  • 🔗 关联底层卡片:指令微调 (SFT):Loss Masking 掩码、Data Packing 样本打包与灾难性遗忘 (Supervised Fine-Tuning: Loss Masking & Data Packing)
  • 🗺️ 知识图谱模块:大语言模型全景图谱

🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)

本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。

👉 前往 TalentMe 交互式研读本题 (M5-028) →


Discover more from AirSOTA – Air School Of Thoughts AtoZ

Subscribe to get the latest posts sent to your email.