【AI 核心深度 M5-024】解释 SFT 数据配比与多任务平衡。(SFT Data Mixture and Multi-Task Balancing)深度数理推导与工程落地解析

所属模块:M5 · NLP 与大语言模型 (NLP & Large Language Models) | 专题分类:指令微调与 SFT (Instruction Tuning & Supervised Fine-Tuning) | 难度等级:Hard

一、核心一句话结论 (One-Sentence Summary)

多任务混合需平衡各任务的样本量与难度,避免’强任务主导’或’弱任务被淹没’;用温度采样/权重调整。

ADVERTISEMENT · 赞助推荐

SFT multi-task balancing allocates data proportions across dialogue, coding, mathematics, reasoning, and safety to prevent over-fitting to single modalities while maintaining generalist conversational capabilities.

二、核心考点要义 (Key Insights)

  • 📌 问题:任务样本量差异大 → 大任务主导梯度
  • 📌 对策:温度采样(n^α)、显式权重、按任务均衡
  • 📌 同时需平衡难度与长度(避免长样本主导 token 预算)

English Insights:
– Core objective: prevent task competition and cross-domain interference while maximizing general capabilities across diverse user intents
– Typical industrial mixture: general chat (30-40%), coding and technical tasks (20-30%), mathematical reasoning (15-20%), structured extraction/tools (10-15%), safety & refusal (5%)
– Temperature-based rebalancing: prevents dominant high-volume datasets from drowning out small, critical datasets (like safety refusals or multi-turn tool calling)

三、核心数学原理与机理推导 (Mathematical Principles & Derivation)

$$p(text{task }i)propto n_i^{alpha} (alpha<1 text{rebalances});qquad text{or explicit weights }w_i$$

数学机理:多任务 SFT 的不平衡问题——真实数据集中各任务的样本量差异极大(如通用问答 10 万条、某专业任务 500 条)。若按样本数自然采样,大任务会主导梯度(每个 batch 中大任务样本占多数),导致 (a) 小任务学不好(欠训练);(b) 模型偏向大任务的风格;(c) 小任务的能力被’淹没’。对策一:温度采样(temperature sampling)——把采样概率从 p_i∝n_i 改为 p_i∝n_i^α(α∈(0,1),常取 0.5~0.7);这压低大任务的权重、抬高小任务的权重,使各任务的’被训练程度’更均衡。α=1 是自然比例、α=0 是完全等权。对策二:显式权重——直接为各任务设权重 w_i(按业务重要性或’能力缺口’);更可控但需调参。对策三:按’有效 token 数’均衡——因为不同任务的回答长度差异大(长推理 vs 短分类),按’样本数’均衡仍可能让长样本主导 token 预算;故应按token 数均衡。其他维度:(a) 难度平衡——混合简单与困难样本(全简单学不到能力、全难学不稳);(b) 能力覆盖——确保关键能力(指令遵循、推理、代码、安全)都有足够样本;(c) 冲突处理——某些任务的风格冲突(如’简洁’与’详尽’),需权衡或加任务标识。评估——需逐任务评估(而非只看总体平均);并检查是否出现’某任务提升但另一任务下降’(跷跷板)。实践——工业界常用’温度采样 + 显式权重微调’,并通过消融实验确定配比。

📖 查看英文严格数学推导 (English Mathematical Derivation)

Mathematical Mechanism: 1. Multi-Task SFT Loss: Given $M$ distinct task datasets $mathcal{D}_1, dots, mathcal{D}_M$ with dataset sizes $N_1, dots, N_M$: $$mathcal{L}(theta) = sum_{m=1}^M w_m mathbb{E}_{(x, y) sim mathcal{D}_m} left[ -sum_{t=1}^{|y|} log P_theta(y_t mid x, y_{<t}) right]$$ 2. Temperature Sampling Rebalancing: If data is sampled proportional to raw size $N_m$, a massive chat dataset ($100text{k}$ examples) completely starves a small safety dataset ($2text{k}$ examples). Temperature rebalancing scales sampling probability by temperature $tau ge 1$: $$p_m = frac{N_m^{1 / tau}}{sum_{j=1}^M N_j^{1 / tau}}$$ – When $tau = 1$: sampling is proportional to raw dataset size. – When $tau to infty$: sampling is completely uniform across all tasks ($p_m = 1/M$). – Standard practice uses $tau in [2, 5]$ to boost the presence of high-value, low-volume datasets without extreme duplication.

四、工业级落地权衡与工程考量 (Industrial Trade-offs)

深度剖析与工程权衡:① ‘按 token 数均衡’比’按样本数均衡’更合理——因为梯度贡献与 token 数相关(每个 token 一个损失项);长回答的任务在 token 预算中占比更高。这是易被忽视的细节。② 温度采样的 α 选择——α 越小则越均衡(小任务权重越高);但 α 过小会让大任务(通常是通用能力)训练不足。故需在’均衡’与’不损害主力任务’间权衡(常取 0.5~0.7)。③ ‘跷跷板’现象——多任务学习中常见’此消彼长’(某任务提升伴随另一任务下降);成因是任务间的梯度冲突(见 M3 的梯度冲突题)或容量竞争。缓解:(a) 调整配比;(b) 梯度投影(PCGrad);(c) 用 MoE/适配器分离任务。④ 与’数据质量’的交互——低质量任务的样本即使量少也会损害(因为模型会学到其错误);故先保证每个任务的数据质量,再谈配比。⑤ 与 RLHF 的衔接——SFT 的配比决定’基础行为’;RLHF 阶段可进一步用偏好数据调整(如对某任务给更高奖励权重)。故 SFT 配比不必完美,RLHF 可补偿。⑥ 面试要点——被问’多任务 SFT 怎么配比’,应给出’温度采样(n^α)+ 显式权重 + 按 token 而非样本数均衡 + 难度/能力覆盖‘,并强调’逐任务评估、警惕跷跷板‘;能指出’按 token 均衡更合理’是深度理解的标志。

⚙️ 查看英文落地权衡分析 (English Systems & Trade-offs)

Deep Dive & Engineering Trade-offs: ① Task Conflict Dynamics: Certain tasks exhibit negative gradient transfer: aggressive safety training degrades coding and math performance (causing false refusal of benign code exploits); aggressive short-answer training degrades detailed multi-step reasoning. Domain weights must be tuned to keep refusal rates bounded. ② Maximum Sample Capping: A simple, robust alternative to temperature sampling is hard capping: set a maximum threshold $N_{max} = 20,000$ samples per task category, ensuring no single dataset dominates the training run. ③ Role of System Prompts: Prepend task-specific system prompts (e.g., ‘You are an expert mathematician’, ‘You are a helpful coding assistant’) to help the Transformer’s attention heads condition on distinct behavioral modes, minimizing cross-task interference. ④ Multi-Turn vs Single-Turn Balancing: Maintain at least $30text{–}40%$ multi-turn conversational data; models trained exclusively on single-turn Q&A fail to track context across extended dialogues. ⑤ Interview Strategy: Present a realistic industrial data mixture breakdown, derive the temperature sampling formula $N_m^{1/tau}$, and discuss how system prompts and sample caps mitigate negative task transfer.

五、常见面试避坑陷阱 (Common Pitfalls & Traps)

  • ⚠️ 按样本数等权采样(忽略 token 数差异)
  • ⚠️ 只看总体平均指标(掩盖任务间失衡)

English Pitfalls:
– Training on raw dataset proportions where a large chat dataset starves critical small datasets like tool-calling or safety
– Over-indexing on safety data, leading to severe false-refusal rates on programming and security queries
– Omitting multi-turn conversation data during SFT (leads to context amnesia in chat applications)

六、高频深度面试追问与预测 (Follow-Up Questions)

  1. 为什么按样本数等权采样不够?
  2. How does temperature parameter $tau$ balance sample diversity against overfitting on small datasets in multi-task SFT?
  3. 如何评估多任务平衡效果?
  4. What causes negative transfer between safety refusal training and code vulnerability analysis?

七、知识图谱对齐 (Knowledge Graph Anchor)

  • 🔗 关联底层卡片:指令微调 (SFT):Loss Masking 掩码、Data Packing 样本打包与灾难性遗忘 (Supervised Fine-Tuning: Loss Masking & Data Packing)
  • 🗺️ 知识图谱模块:大语言模型全景图谱

🔬 算法科学家与机器学习深度考察全量题库 (Science Depth)

本题收录于 TalentMe 算法科学家深度考察真题库 (Science Depth)。全库共 856 道硬核考点,深度覆盖数学统计、经典ML、深度学习、Transformer、大语言模型、多模态、推荐系统与 MLOps。支持 Jev 面经智能匹配、一键离线单文件 HTML 手册导出并直连 Obsidian 本地记忆。

👉 前往 TalentMe 交互式研读本题 (M5-024) →


Discover more from AirSOTA – Air School Of Thoughts AtoZ

Subscribe to get the latest posts sent to your email.